Papers reviewed and determined not to be word norm studies. Use the flag icon to report errors or suggest re-inclusion.
16504 papers
It is no secret that people often use taboo words when speaking about persons and objects in their environment. Taboo words are charged with emotion and have observable impact on the listener as well as the speaker. The purpose of this study was to determine whether taboo words were quantitatively more offensive when used in combination with a proper name versus being used with a non-human object. We found that using taboo words to describe proper names does not cause a significant effect; however, we found that participants rated certain categories of taboo words as more offensive than other categories. In a second experiment, taboo words did affect ratings and memory for proper names and non-human objects.
The statistical parsing of morphologically rich languages is hindered by the inability of parsers to collect solid statistics because of the large number of word types in such languages. There are however two separate but connected problems, reducing data sparsity of known words and handling rare and unknown words. Methods for tackling one problem may inadvertently negatively impact methods to handle the other. We perform a tightly controlled set of experiments to reduce data sparsity through class-based representations in combination with unknown word signatures with two PCFG-LA parsers that handle rare and unknown words differently on the German TiGer treebank. We demonstrate that methods that have improved results for other languages do not transfer directly to German, and that we can obtain better results using a simplistic model rather than a more generalized model for rare and unknown word handling.
The focus of the current study was on idiom comprehension in younger and older adults. Due to inconsistent results in previous studies, it is unclear whether older adults may have problems understanding idioms. For the current study, I used a sentence-to-word matching task presented on an iPad with software that recorded participants’ response time and accuracy. Participants also completed a familiarity task where they rated idioms on how frequently these phrases were encountered. I predicted that older adults would have more difficulty comprehending idioms because of the context in which the idioms were embedded and the timed nature of the task. I also predicted that both age groups would rate the idioms as highly familiar because we purposefully selected these types of expressions. With respect to the sentence-to-word matching task, results showed that although older adults were slower overall, both younger and older adults showed faster response times and greater accuracy for idiomatic targets following idiomatically-biased contexts than for literal targets following literally-biased contexts. With respect to the familiarity ratings task, results showed that both age groups were very familiar with the idioms. These findings suggest that older adults are able to successfully use context to understand familiar ambiguous idioms and that they do not have difficulty comprehending idioms in a cognitively demanding timed task.
Abstract People remember events and materials better when these are congruent with their mood at retrieval; this is known as the mood-congruent memory bias. This effect is largest when the materials are self-referential and this is known as the self-reference effect. We present two word rating studies, to create a list of self-referential valenced words that may be used as stimuli to investigate the influence of valence on cognitive processing in depressive ruminators. Words selected from the Affective Norms for English Words pool were rated by an unselected sample for self-referentiality (Study 1) and validated with ratings provided by depressive ruminators. As hypothesized, depressive ruminators rated negative words as more self-referential than an unselected sample. Using this list, valence differentiated performance between depressive ruminators and healthy controls in a working memory updating task. We thus created a list of self-referential valenced words matched on factors that influence word processing.
This study aims at exploring new norms as to the textual additions in parentheses (=TAiPs) in the translation of a Quranic text as writer-oriented devices of textuality. Coding for this sort of information could be useful in establishing an impact on any decision-making process on the TL version; such TAiPs can give a translated text of the Quran unity and purpose and distinguish it from a disconnected sequence of sentences. Six small-sized chapters of the Quran were selected as a research sample including a number of four handred forty two (442) TAiPs. Two writer-oriented kinds of textuality were found: cohesivity at the levels of grammar and lexis to be in form of recurrence, reference, substitution, ellipsis and conjunction; and relationality by coherence and intentionality to be in form of reiteration, collocation, connotation, evocation and interpretation. The study is a detailed analysis of such a severely criticized yet officially approved English interpretation of the Quran as the Hilali and Khan Translation (=HKT) against a predetermined set of text-linguistic norms. The strength or weakness of TAiPs as to how they might alleviate or aggravate the TL version is eventually identified for sake of improvement.
The first edition of one of the most important and mysterious novels of the 20th century appeared more than fifty years ago. Despite the passage of time The Master and Margarita still enjoys popularity; it also intrigues and inspires. Until now five Polish translations of Bulgakov’s novel have appeared. It is known that the interpretation of the original might be expressed in the form of many potential texts that are communicatively equivalent. There is no doubt that it is the translator who plays a vital role in any translation; her/his personality, life experience, knowledge, skills, and also the times s/he lives in regulate the target text. That is why, no matter how many times a text is translated, the final product will always be different. Taking this into consideration, the author will compare the three Polish translations of Bulgakov’s Master and Margarita, paying attention to the diachronic perspective as far as linguistic norms are concerned, the modernity of language, and the way the anthroponyms are expressed.
The repertoire of forms of address can be considered as one of the determinants of the discourse genre, which makes it possible to capture its evolution and cultural variations. From such comparative, intra- and intercultural perspective, adopting an interactive approach in the analysis of political discourse, we will look at the practice of addressing one another in the French and Polish politicalmedia discourse. While in both languages the linguistic norm recommends the use of the polite forms of address in official situations, the cases of the use of the familiar pronoun tu / ty in media interactions between politicians are not rare at all. Whether it is an informal talk of politicians caught by the media, a television pre-election debate, or a meeting of the heads of state, addressing the other person by the familiar forms is a manifestation of a deliberate blurring of the boundaries between the front-stage and backstage in political discourse in order to create the impression of intimacy andequality between the interlocutors.
This study examines how the acoustic input (the surface form) and the abstract linguistic representation (the underlying representation) interact during spoken word recognition by investigating left-dominant tone sandhi, a tonal alternation in which the underlying tone of the first syllable spreads to the sandhi domain. We conducted an auditory-auditory priming lexical decision experiment on Shanghai left-dominant sandhi words, in which each disyllabic target ([tɕi55 dɛ31] “egg”) was preceded by monosyllabic primes either sharing the same underlying tone ([tɕi55]), surface tone ([tɕi53] “machine”), or being unrelated to the tone of the first syllable of the sandhi targets ([tɕi24] “to remember”). Results showed a surface priming effect, but not an underlying priming effect. Moreover, the surface priming did not interact with speakers’ familiarity ratings to the sandhi targets. The results are discussed in the context of how phonological opacity, productivity, and the directionality of tone sandhi patterns influence the representation of tone sandhi words as well as how the lexicality of the primes and the participants’ usage pattern of Shanghai may have influenced the results.
The Universal Dependencies project is currently comprised of 71 languages and 122 treebanks, and aims to find morphological and syntactic characteristics that can be applied to multiple languages for parallel language processing. In this paper, we introduce Universal POS, which is a morphological tagset for UD, and propose a method to automatically convert existing Korean morphological tagset into UPOS. In order to apply the UPOS tagset, which is based on refraction words such as English, to the Korean language, it is necessary to try a one-to-many mapping between the UPOS individual tag and the 21st century Sejong tag combination. (Yonsei University)
Because the most common transition systems are projective, training a transition-based dependency parser often implies to either ignore or rewrite the non-projective training examples, which has anadverse impact on accuracy. In this work, we propose a simple modification of dynamic oracles, which enables the use of non-projective data when training projective parsers. Evaluation on 73~treebanks shows that our method achieves significant gains (+2 to +7 UAS for the most non-projective languages) and consistently outperforms traditional projectivization and pseudo-projectivizationapproaches.
Polycentric Spanish Norm Towards the Polish‑Spanish Legal Translation The Spanish, being the official language of Spain and many other countries, is characterized by an important dialectal diversity that is reflected in the differences at all linguistic levels: phonetic, morphological, syntactic and lexico‑semantic, etc. All these differences raise controversies and discussions about the existence of a linguistic norm depending on the perspective that can have a monocentric or polycentric character. In this contribution we present some arguments for the second one. To this end, we rely on translations, starting simultaneously from the semasiological and onomasiological perspective, of some Polish‑Spanish legal terms in which it is essential to take into account, the diatopic variation as well as the norm whose character is polycentric.
The short note describes the chart parser for multimodal type-logical grammars which has been developed in conjunction with the type-logical treebank for French. The chart parser presents an incomplete but fast implementation of proof search for multimodal type-logical grammars using the "deductive parsing" framework. Proofs found can be transformed to natural deduction proofs.
The authors try to answer two questions: 1. How the Polish philology students understand the concept of linguistic norm? and 2. When, according to those students, people should follow it? The article presents the results of the survey conducted among 200 respondents. It turns out that the students understand the concept of norm well, usually as a set of rulles, which are established by linguists or/and accepted by society. They also think that the respect of rules takes effect in varying degrees in different communication situations.
To reinforce sports informatization management and the sports service quality on college campuses, this paper researches combined multi-agent technology, and structures undergraduate sports service system framework. It has elaborated functions of every feature and workflow of the system, and puts forward agent design procedure based on JADE. What is more, it also adopts FIPAACl linguistic norms between Agents communication and proposes some advices which based on the fundamental of Agent.
Emotional imagery is a common induction technique used in the laboratory and also employed in various exposure therapy treatments across the anxiety spectrum (e.g., specific and social phobias). Despite its clinical uses, there is a surprising dearth of literature regarding the basic central neural processes underlying emotional imagery, though other peripheral physiological processes have been investigated extensively using heart rate, skin conductance, and startle-blink responses. One imagery study that used a central nervous system psychophysiological measure -event specific brainwave or the event-related potential (ERP) technique- suggests the late positive potential (LPP) of the ERP is larger for unpleasant versus neutral stimuli, implying this ERP may index emotional engagement during imagery. This effect is consistent with the visual perception literature of emotion; however, the visual perception literature also indicates that the LPP is larger for pleasant stimuli versus neutral stimuli and positively correlated with subjective emotional arousal ratings. Using script-driven emotional imagery, we will extend research on the LPP to establish whether 1) the LPP is larger for both pleasant and unpleasant scripts relative to neutral ones and 2) this LPP effect is positively correlated with emotional arousal ratings. Fifty-five participants will make subjective ratings of the scripts and then imagine the scripts while electroencephalographic data are, recorded. Upon demonstrating the LPP is larger for emotional (both pleasant and unpleasant) scripts than neutral ones, this study will lay the foundation for future work aimed at determining whether LPP effects are hyper- or hypo-active for socially anxious participants.
Currently, the biaffine classifier has been attracting attention as a method to introduce an attention mechanism into the modeling of binary relations. For instance, in the field of dependency parsing, the Deep Biaffine Parser by Dozat and Manning has achieved state-of-the-art performance as a graph-based dependency parser on the English Penn Treebank and CoNLL 2017 shared task. On the other hand, it is reported that parameter redundancy in the weight matrix in biaffine classifiers, which has O(n^2) parameters, results in overfitting (n is the number of dimensions). In this paper, we attempted to reduce the parameter redundancy by assuming either symmetry or circularity of weight matrices. In our experiments on the CoNLL 2017 shared task dataset, our model achieved better or comparable accuracy on most of the treebanks with more than 16% parameter reduction.
Annotation Guidelines for Text Analytics in Social Media A person's language use reveals much about their profile, however, research in author profiling has always been constrained by the limited availability of training data, since collecting textual data with the appropriate meta-data requires a large collection and annotation effort (Maamouri et al. 2010; Diab et al. 2008; Hawwari et al. 2013).For every text, the characteristics of the author have to be known in order to successfully profile the author. Moreover, when the text is written in a dialectal variety such as the Arabic text found online in social media a representative dataset need to be available for each dialectal variety (Zaghouani et al. 2012; Zaghouani et al. 2016).The existing Arabic dialects are historically related to the classical Arabic and they co-exist with the Modern Standard Arabic in a diglossic relation. While the standard Arabic, has a clearly defined set of orthographic standards, the various Arabic dialects have no official orthographies and a given word could be written in multiple ways in different Arabic dialects (Maamouri et al. 2012; Jeblee et al. 2014).This abstract presents the guidelines and annotation work carried out within the framework of the Arabic Author profiling project (ARAP), a project that aims at developing author profiling resources and tools for a set of 12 regional Arabic dialects. We harvested our data from social media which reflect a natural and spontaneous writing style in dialectal Arabic from users in different regions of the Arabworld.For the Arabic language and its dialectal varieties as foundin social media, to the best of our knowledge, there is nocorpus available for the detection of age, gender, nativelanguage and dialectal variety. Most of the existingresources are available for English or other Europeanlanguages. Having a large amount of annotated data remains the key to reliable results in the taskof author profiling. In order to start the annotation process, we createdguidelines for the annotation of the Tweets according totheir dialectal variety, their native language, the gender of the user and the age. Before starting theannotation process, we hired and trained a group of annotators and we implemented a smooth annotation pipeline to optimize the annotation task. Finally, we followed a consistent annotation evaluation protocol to ensure a high inter-annotator agreement.The Annotations were done by carefully analyzing each ofthe user's profiles, their tweets, and when possible, weinstructed the annotators to use external resources such asLinkedIn or Facebook. We created a general profilesvalidation guidelines and task-specific guidelines toannotate the users according to their gender, age, dialectand their native language. For some accounts, the annotators were not able to identifythe gender as this was based in most of the cases on thename of the person or his profile photo and in some casesby their biography or profile description. In case thisinformation is not available, we instructed the annotators toread the user posts and find linguistic indicators of thegender of the user.Like many other languages, Arabic conjugates verbsthrough numerous prefixes and suffixes and the gender issometimes clearly marked such as in the case of the verbsending in taa marbuTa which is usually of femininegender.In order to annotate the users for their age, we used threecategories: under 20 years, between 20 years and 40 years,and 40 years and up.In our guidelines, we asked our annotators to try their bestto annotate the exact age, for example, they can check theeducation history of the users in LinkedIn and Facebookprofile and find when the graduated from high school forexample in order to guess the age of the users. As the dialect and the regions are known in advance to theannotators, we instructed them to double check and markthe cases when the profile appears to be from a differentdialect group. This is possible despite our initial filteringbased on distinctive regional keywords. We noticed that inmore than 90% the profiles selected belong to the specifieddialect group. Moreover, we asked the annotators to mark and identifyTwitter profiles with a native language other than Arabic,so they are considered as Arabic L2 speakers. In order tohelp the annotators identify those, we instructed them tolook for various cues such as the writing style, the sentence structure, the word order and the spelling errors.AcknowledgementsThis publication was made possible by NPRP grant #9-175-1-033 from the Qatar National Research Fund (a member ofQatar Foundation). The statements made herein are solelythe responsibility of the authors. ReferencesDiab Mona, Aous Mansouri, Martha Palmer, Olga Babko-Malaya, Wajdi Zaghouani, Ann Bies, Mohammed Maamouri. A Pilot Arabic Propbank; LREC 2008, Marrakech, Morocco, May 28-30, 2008.Hawwari, A.; Zaghouani, W.; O»Gorman, T.; Badran, A.; Diab, M., «Building a Lexical Semantic Resource for Arabic Morphological Patterns,» Communications, Signal Processing, and their Applications (ICCSPA), 2013, vol., no., pp.1,6, 12-14 Feb. 2013. Jeblee Serena; Houda Bouamor; Wajdi Zaghouani; Kemal Oflazer. CMUQ@QALB-2014: An SMT-based System for Automatic Arabic Error Correction. In Proceedings of the EMNLP 2014 Workshop on Arabic Natural Language Processing (ANLP), Doha, Qatar, October 2014.Maamouri Mohamed, Ann Bies, Seth Kulick, Wajdi Zaghouani, Dave Graff and Mike Ciul. 2010. From Speech to Trees: Applying Treebank Annotation to Arabic Broadcast News. In Proceedings of LREC 2010, Valetta, Malta, May 17-23, 2010.Maamouri Mohammed, Wajdi Zaghouani, Violetta Cavalli-Sforza, Dave Graff and Mike Ciul. 2012. Developing ARET: An NLP-based Educational Tool Set for Arabic Reading Enhancement. In Proceedings of The 7th Workshop on Innovative Use of NLP for Building Educational Applications, NAACL-HLT 2012, Montreal, Canada.Obeid Ossama, Wajdi Zaghouani, Behrang Mohit, Nizar Habash, Kemal Oflazer and Nadi Tomeh. A Web-based Annotation Framework For Large- Scale Text Correction. In Proceedings of IJCNLP'2013, Nagoya, Japan.Zaghouani Wajdi, Nizar Habash, Ossama Obeid, Behrang Mohit, Houda Bouamor, Kemal Oflazer. 2016. Building an arabic machine translation post-edited corpus: Guidelines and annotation. In Proceedings of the International Conference on Language Resources and Evaluation (LREC»2016).Zaghouani Wajdi, Abdelati Hawwari and Mona Diab. 2012. A Pilot PropBank Annotation for Quranic Arabic. In Proceedings of the first workshop on Computational Linguistics for Literature, NAACL-HLT 2012, Montreal, Canada.
Prestige and dominance are thought to be two evolutionarily distinct routes to gaining status and influence in human social hierarchies. Prestige is attained by having specialist knowledge or skills that others wish to learn, whereas dominant individuals use threat or fear to gain influence over others. Previous studies with groups of unacquainted students have found prestige and dominance to be two independent avenues of gaining influence within groups. We tested whether this result extends to naturally-occurring social groups. We ran an experiment with 30 groups of 5 people from Cornwall, UK (n=150). Participants answered general knowledge questions individually and as a group, and subsequently nominated a team representative to answer bonus questions to win money on behalf of the team. Participants then rated all other team-mates anonymously on scales of prestige, dominance, likeability and influence on the task. Using a model comparison approach with Bayesian multi-level models, we found that prestige and dominance ratings were predicted by influence ratings on the task, replicating previous studies. However, prestige and dominance ratings did not predict who was nominated as group representative. Instead, participants nominated team members with the highest individual quiz scores, despite this information being unavailable to them. Interestingly, team members who were initially rated as being high status in the group, such as a team captain or group administrator, had higher ratings of both dominance and prestige than other group members. In contrast, those who were initially rated as someone from whom group members would like to learn had higher prestige ratings, but not higher dominance ratings, supporting the claim that prestige reflects social learning opportunities. Our results suggest that prestige and dominance hierarchies do become established in naturally occurring human social groups, but that these hierarchies may be more domain-specific and less flexible than we anticipated.
The article deals with the issue of translation as an important means of communication between individuals who speak different languages and belong to different cultures.The article analyzes the translation as interlingual communicative phenomenon.The translation process is determined by the linguistic norms, communicative situations, functional parameters of the original text and translation norms.The role and tasks of an interpreter in the process of intercultural communication of individuals are defined.
This chapter discusses the new and changing conditions for linguistic norms in literary fiction of the post-Soviet era. In particular, it looks at how the interrelationship between the language of literature (<italic>iazyk literatury</italic>) and the standard language (<italic>literaturnyi iazyk</italic>) has been challenged by several processes of sociolinguistic change, including initiatives in language policy.
In this paper, a contrastive analysis of normative issues concerning the grammatical categories of determinants and pronouns included in the Spanish Royal Academy Grammars and Dictionaries is carried out. The corpus comprises the different editions of its grammatical work (1771, 1796, 1854, 1870, 1883, 1911, 1917, the 1973 Sketch and the NGLE of 2009), as well as the twenty-three editions of its lexicographical work (from 1780 to 2014). Issues related to the linguistic norm, which have been extracted from a comparison between the different editions throughout history, have been examined, classified and described. Results from this study reveal the following data: on the one hand, Grammars pay closer attention to prescriptive issues; on the other hand, a lack of coherence between Grammars and Dictionaries both in the follow-up of the linguistic norm and in the correction criteria used can be observed. In addition to that and with regard to dictionaries, the manual editions along with the 23rd edition are exceptional provided their interest in collecting a greater number of allusions to proper linguistic use.
The article discusses the terms that nominate the language of written artifacts documented by the Cyrillic on the Ukrainian-Byelorussian lands in the XIV-XVI centuries; the expediency of using the notion “literary language” as to the Ukrainian literary written tradition of the XIV–XVI centuries is clarified; the content of the term “linguistic norm” is outlined, its characteristics in the investigated period are determined.
The linguistic database is also positioned as an actual way of formalizing and organizing phraseological units, terms for designating types of phraseological units. The main principle of systematization of the latter in the study is the thesaurus principle, that is the filling of the paradigm «terminological system – terminological microsystem – terminological subsystem – term», represented by a linguistic database.
Abstract The aim of the contribution is to introduce a database of linguistic forms and their functions built with the use of the multi-layer annotated corpora of Czech, the Prague Dependency Treebanks. The purpose of the Prague Database of Forms and Functions (ForFun) is to help the linguists to study the form-function relation, which we assume to be one of the principal tasks of both theoretical linguistics and natural language processing. We demonstrate possibilities of the exploitation of the ForFun database. This article is largely based on a paper presented at the 16th International Workshop on Treebanks and Linguistic Theories in Prague (Bejček et al., 2017).
Functional alterations of the default mode network (DMN) are frequently reported in psychotic disorders, but the functional role of these alterations remains poorly known. In addition to previous studies that have applied different types of tasks or recorded resting-state neuroimaging data, there has recently been more interest in the use of movie stimuli in studying brain functioning in patient populations, because this could provide a more naturalistic account of brain functioning in real life-like situations. Seventy-one first-episode psychosis (FEP) patients (mean age = 26.0 yrs, 47 (66%) males) and 57 controls (mean age = 26.86 yrs, 24 (42%) males) from the Helsinki Early Psychosis Study watched scenes from the movie Alice in Wonderland (Tim Burton, 2010) during 3 T fMRI-BOLD imaging. We used intersubject correlation (ISC) analysis, in which the correlation between voxel-wise BOLD time series in every within-group pair of subjects is calculated. In this study, time-windowed ISC was calculated with a 10-TR (time of repetition, 1.8 s) window with 1-TR steps over the fMRI time series. In each ISC window, a two-sample t test was performed to obtain a t-statistic time series of differences between the groups. An independent group of control subjects (n = 17, 10 males, mean age 26.5 yrs) rated how emotionally arousing the currently seen events of the stimulus are, producing a time-varying rating used as a regressor. General linear model was used to identify brain regions where the t-statistic time series covaries with the arousal rating. To make the interpretation of results less ambiguous, the arousal rating was divided into high and low arousal regressor by z scoring the rating and taking only the positive and negative values, respectively. Nonparametric clusterwise permutation test was used for statistical inference (cluster-defining threshold of p = 0.05, familywise error corrected threshold of p = 0.05, number of permutations = 5000). Furthermore, by using an experience-sampling setup during the same brain-scanning session, a partially overlapping sample of participants reported how emotionally aroused they were feeling during scanning. The results show significant correlation between the t-statistic time series and low arousal regressor, especially in the DMN including the anterior and posterior cingulate cortex, medial prefrontal cortex, precuneus, and bilateral lateral temporoparietal regions. Closer inspection reveals that during moments of low arousal in the movie stimulus, the ISC of healthy controls goes up but the ISC of patients does not. In the experience-sampling portion of the study, the patients reported more arousal than the control subjects. Intersubject correlation in the DMN depended differentially on arousal in FEP patients and control subjects. More specifically, during moments when the stimulus was rated less emotionally arousing, control subjects’ DMN functioning synchronized more while the patients’ did not. In connection with the difference in reported arousal during the same imaging session, our findings provide preliminary evidence for a contribution of arousal on the functional alterations of the DMN and suggest that this may be related to higher baseline arousal in the patients. Higher arousal and the related distortion of high order integrative functioning that characterizes DMN could contribute to the pathogenesis of psychosis.
A puzzling fact about linguistic norms is that they are mainly stable, but the conventional variant sometimes changes. These transitions seem to be mostly S-shaped and, therefore, directed. Previous models have suggested possible mechanisms to explain these directed changes, mainly based on a bias favoring the innovative variant. What is still debated is what is the mechanism that leads to such a bias. In this paper we propose a refined taxonomy of mechanisms of language change and identify a family a mechanisms explaining self-actuated language changes. We exemplify this type of mechanism with the preference-based selection mechanism that relies on agents having dynamical preferences for different variants of the linguistic norm. The key point is that if these preferences can align through social interactions, then new changes can be actuated. We present results of a multi-agent model and demonstrate that the model produces trajectories that are typical of language change.
This chapter takes up the issue of authenticity in language pedagogy. Traditional views of authenticity take the native speaker to be the primary authority for linguistic norms. Written standard language is especially highly valued here. It is argued herein that TELL environments are equally valid as learning environments, and that students can use the freedom they provide to develop their own locally negotiated cultural and linguistic norms. Evidence is provided that students on a net-based MA program develop their own norms for reducing language, and use them and other means to mark membership of a local TELL community. Thus, TELL is a rich and authentic environment for learners of English to become what is referred to as “language practitioners.”
Abstract In the context of the Index Thomisticus Treebank project, we have enhanced the full text of Bellum Catilinae by Sallust with semantic annotation. The annotation style resembles the one used for the so called “tectogrammatical” layer of the Prague Dependency Treebank. By exploiting the results of semantic role labeling, ellipsis resolution and coreference analysis, this paper presents a network-based study of the main Actors and Actions (and their relations) in Bellum Catilinae.
並列構造解析の主たるタスクは並列する句の範囲を同定することである.並列構造は文の構文・意味の解析において有用な特徴となるが,これまで決定的な解析手法が確立されておらず,現在の最高精度の構文解析器においても誤りを生じさせる主たる要因となっている.既存の並列句範囲の曖昧性解消手法は並列構造の類似性のみの特性や構文解析器の結果に強く依存しているという問題があった.本研究では,近年自然言語解析に広く使用されているリカレントニューラルネットワークを用いて,構文解析の結果を用いずに単語の表層形と品詞情報のみから並列句の類似性と可換性の特徴ベクトルを計算し,並列構造の範囲を予測する手法を提案する.Penn Treebank と GENIA コーパスを用いた実験の結果,提案手法によって先行研究を上回る解析精度を得た.
Accurate natural language processing systems rely heavily on annotated datasets. In the absence of such datasets, transfer methods can help to develop a model by transferring annotations from one or more rich-resource languages to the target language of interest. These methods are generally divided into two approaches: 1) annotation projection from translation data, aka parallel data, using supervised models in rich-resource languages, and 2) direct model transfer from annotated datasets in rich-resource languages. In this thesis, we demonstrate different methods for transfer of dependency parsers and sentiment analysis systems. We propose an annotation projection method that performs well in the scenarios for which a large amount of in-domain parallel data is available. We also propose a method which is a combination of annotation projection and direct transfer that can leverage a minimal amount of information from a small out-of-domain parallel dataset to develop highly accurate transfer models. Furthermore, we propose an unsupervised syntactic reordering model to improve the accuracy of dependency parser transfer for non-European languages. Finally, we conduct a diverse set of experiments for the transfer of sentiment analysis systems in different data settings. A summary of our contributions are as follows: * We develop accurate dependency parsers using parallel text in an annotation projection framework. We make use of the fact that the density of word alignments is a valuable indicator of reliability in annotation projection. * We develop accurate dependency parsers in the absence of a large amount of parallel data. We use the Bible data, which is in orders of magnitude smaller than a conventional parallel dataset, to provide minimal cues for creating cross-lingual word representations. Our model is also capable of boosting the performance of annotation projection with a large amount of parallel data. Our model develops cross-lingual word representations for going beyond the traditional delexicalized direct transfer methods. Moreover, we propose a simple but effective word translation approach that brings in explicit lexical features from the target language in our direct transfer method. * We develop different syntactic reordering models that can change the source treebanks in rich-resource languages, thus preventing learning a wrong model for a non-related language. Our experimental results show substantial improvements over non-European languages. * We develop transfer methods for sentiment analysis in different data availability scenarios. We show that we can leverage cross-lingual word embeddings to create accurate sentiment analysis systems in the absence of annotated data in the target language of interest. We believe that the novelties that we introduce in this thesis indicate the usefulness of transfer methods. This is appealing in practice, especially since we suggest eliminating the requirement for annotating new datasets for low-resource languages which is expensive, if not impossible, to obtain.
The contents and structure of semantic memory have been the focus of much recent research, with major advances in the development of distributional models, which use word co-occurrence information as a window into the semantics of language. In parallel, connectionist modeling has extended our knowledge of the processes engaged in semantic activation. However, these two lines of investigation have rarely been brought together. Here, we describe a processing model based on distributional semantics in which activation spreads throughout a semantic network, as dictated by the patterns of semantic similarity between words. We show that the activation profile of the network, measured at various time points, can successfully account for response times in lexical and semantic decision tasks, as well as for subjective concreteness and imageability ratings. We also show that the dynamics of the network is predictive of performance in relational semantic tasks, such as similarity/relatedness rating. Our results indicate that bringing together distributional semantic networks and spreading of activation provides a good fit to both automatic lexical processing (as indexed by lexical and semantic decisions) as well as more deliberate processing (as indexed by ratings), above and beyond what has been reported for previous models that take into account only similarity resulting from network structure.
Abstract Facial expressions are fundamental to interpersonal communication, including social interaction, and allow people of different ages, cultures, and languages to quickly and reliably convey emotional information. Historically, facial expression research has followed from discrete emotion theories, which posit a limited number of distinct affective states that are represented with specific patterns of facial action. Much less work has focused on dimensional features of emotion, particularly positive and negative affect intensity. This is likely, in part, because achieving inter-rater reliability for facial action and affect intensity ratings is painstaking and labor-intensive. We use computer-vision and machine learning (CVML) to identify patterns of facial actions in 4,648 video recordings of 125 human participants, which show strong correspondences to positive and negative affect intensity ratings obtained from highly trained coders. Our results show that CVML can both (1) determine the importance of different facial actions that human coders use to derive positive and negative affective ratings, and (2) efficiently automate positive and negative affect intensity coding on large facial expression databases. Further, we show that CVML can be applied to individual human judges to infer which facial actions they use to generate perceptual emotion ratings from facial expressions.
This paper carries out an empirical analysis of various dropout techniques for language modelling, such as Bernoulli dropout, Gaussian dropout, Curriculum Dropout, Variational Dropout and Concrete Dropout. Moreover, we propose an extension of variational dropout to concrete dropout and curriculum dropout with varying schedules. We find these extensions to perform well when compared to standard dropout approaches, particularly variational curriculum dropout with a linear schedule. Largest performance increases are made when applying dropout on the decoder layer. Lastly, we analyze where most of the errors occur at test time as a post-analysis step to determine if the well-known problem of compounding errors is apparent and to what end do the proposed methods mitigate this issue for each dataset. We report results on a 2-hidden layer LSTM, GRU and Highway network with embedding dropout, dropout on the gated hidden layers and the output projection layer for each model. We report our results on Penn-TreeBank and WikiText-2 word-level language modelling datasets, where the former reduces the long-tail distribution through preprocessing and one which preserves rare words in the training and test set.
Abstract In this article we provide a practical demonstration of how syntactically annotated corpora (treebanks), particularly the English Historical Parsed Corpora Series, can be used to investigate research questions with a diachronic depth and synchronic breadth that would not otherwise be possible. The phenomenon under investigation is split coordination, in which two parts of a conjoined constituent appear separated in the clause (e.g., and this is where my aunt lives and my uncle ). It affects every type of coordinated constituent (subject/object DPs, predicate and attributive ADJPs, ADVPs, PPs and DP objects of P) in Old English (OE); and it, or a superficially similar construction, occurs continuously throughout the attested period from approximately 800 to the present day. Despite its synchronic range and diachronic persistence, split coordination has received surprisingly little attention in the diachronic literature, with the exception of Perez Lorido’s (2009) limited study of split subjects in eight OE texts. Its modern counterpart is most frequently analysed as Bare Argument Ellipsis (BAE). Although the OE and Present-Day English constructions appear superficially similar, we show that not all of the OE data is amenable to a BAE analysis. We bring to bear different types of evidence (structural, discourse/performance effects, rate of change, etc.) to argue that split coordination in fact represents two different constructions, one of which remains stable over time while the other is lost in the post-Middle English period.
NLTK toolkit is an API platform built with Python language to interact with humans through natural language. The very first version of NLTK was released in 2005 (1.4.3), which was compatible with Python 2.4. The latest version was in September 2017 NLTK (3.2.5), which incorporated features like Arabic stemmers, NIST evaluation, MOSES tokenizer, Stanford segmenter, treebank detokenizer, verbnet, and vader, etc. NLTK was created in 2001 as a part of Computational Linguistic Department at the University of Pennsylvania. Since then it has been tested and developed. The important packages of this system are 1) corpus builder, 2) tokenizer, 3) collocation, 4) tagging, 5) parsing, 6) metrics, and 7) probability distribution system. Toolbox NLTK was built to meet four primary requirements: 1) Simplicity: An substantive framework for building blocks; 2) Consistency: Consistent interface; 3) Extensibility: Which can be easily scaled; and 4) Modularity: All modules are independent of each other.
Language models, which are used in various tasks including speech recognition and sentence completion, are usually used with texts covering various domains. Therefore, domain adaptation has been a long-ongoing challenge in language model research. Conventional methods mainly work by the addition of a domain dependent bias. In this paper, we propose a novel way to adapt neural network-based language models. Our proposed approach relies on a linear combination of factorised hidden layers, which are learnt by the network. For domain adaptation, we use topic features from latent Dirichlet allocation. These features are input into an auxiliary network, and the output of this network is used to calculate the hidden layer weights. Both the auxiliary network and the main network can be trained jointly by error backpropagation. This makes our proposed approach completely unsupervised. To evaluate our method, we show results for the well-known Penn Treebank and the TED-LIUM dataset.
Easy-first parsing relies on subtree re-ranking to build the complete parse tree. Whereas the intermediate state of parsing processing is represented by various subtrees, whose internal structural information is the key lead for later parsing action decisions, we explore a better representation for such subtrees. In detail, this work introduces a bottom-up subtree encoding method based on the child-sum tree-LSTM. Starting from an easy-first dependency parser without other handcraft features, we show that the effective subtree encoder does promote the parsing process, and can make a greedy search easy-first parser achieve promising results on benchmark treebanks compared to state-of-the-art baselines. Furthermore, with the help of the current pre-training language model, we further improve the state-of-the-art results of the easy-first approach.
espanolEn este articulo se lleva a cabo un analisis contrastivo de las cuestiones normativas, referentes a las categorias gramaticales de los determinantes y de los pronombres, incluidas en las gramaticas y en los diccionarios de la Real Academia Espanola. El corpus lo integran las diferentes ediciones de su obra gramatical (1771, 1796, 1854, 1870, 1883, 1911, 1917, el Esbozo de 1973 y la NGLE de 2009), asi como las veintitres ediciones de su obra lexicografica (desde 1780 hasta 2014). Se han examinado, clasificado y descrito los asuntos referentes a la norma, extraidos tras el cotejo de ambas obras a lo largo de la historia. Las principales conclusiones de este estudio revelan los siguientes datos: por una parte, son las gramaticas las que dedican una mayor atencion a los temas prescriptivos; por otra parte, en ciertas ocasiones, se observa una falta de coherencia entre las gramaticas y los diccionarios tanto en el seguimiento de la norma como en los criterios de correccion empleados; finalmente, con respecto a los diccionarios, sobresalen las ediciones manuales y la vigesima tercera edicion, dado su interes por recoger un mayor numero de alusiones al buen uso linguistico. EnglishIn this paper, a contrastive analysis of normative issues concerning the grammatical categories of determinants and pronouns included in the Spanish Royal Academy Grammars and Dictionaries is carried out. The corpus comprises the different editions of its grammatical work (1771, 1796, 1854, 1870, 1883, 1911, 1917, the 1973 Sketch and the NGLE of 2009), as well as the twenty-three editions of its lexicographical work (from 1780 to 2014). Issues related to the linguistic norm, which have been extracted from a comparison between the different editions throughout history, have been examined, classified and described. Results from this study reveal the following data: on the one hand, Grammars pay closer attention to prescriptive issues; on the other hand, a lack of coherence between Grammars and Dictionaries both in the follow-up of the linguistic norm and in the correction criteria used can be observed. In addition to that and with regard to dictionaries, the manual editions along with the 23rd edition are exceptional provided their interest in collecting a greater number of allusions to proper linguistic use.
Combination of customer needs and quantitative data is an idea to produce a various value. Therefore, this research proposed method for multidisciplinary design optimization of hub airport which optimized customer needs and quantitative data simultaneously. For customer needs, target of tourism and season were assigned into design variables and optimal weight was calculated by using SGD method. Design variables of quantitative data were number of transit, transportation fee and time from airport to World Heritage, and calculated value by using AHP which Monte Carlo simulation was applied for updating optimal value. To obtain optimal value for customer needs, review of tourism was extracted by text mining and count number of review in each World Heritage, word rating point was decided to calculate the weight of each World Heritage. Next, results of customer needs and quantitative data were normalized for collaborative optimization. Finally, the results showed a various optimal solution obtained for each design variable.
Cross-entropy loss is a common choice when it comes to multiclass classification tasks and language modeling in particular. Minimizing this loss results in language models of very good quality. We show that it is possible to fine-tune these models and make them perform even better if they are fine-tuned with sum of cross-entropy loss and reverse Kullback-Leibler divergence. The latter is estimated using discriminator network that we train in advance. During fine-tuning probabilities of rare words that are usually underestimated by language models become bigger. The novel approach that we propose allows us to reach state-of-the-art quality on Penn Treebank: perplexity decreases from 52.4 to 52.1. Our fine-tuning algorithm is rather fast, scales well to different architectures and datasets and requires almost no hyperparameter tuning: the only hyperparameter that needs to be tuned is learning rate.
Ratings are important in attracting foreign capital so they play a great role in the financial system of a country. The aim of the study is to investigate the impact of macroeconomic indicators on sovereign credit ratings assigned by Fitch. For this aim Panel ordered probit model was applied to the annual data from 2000 to 2011. The analysis rests on panel of 44 countries. According to the results obtained it can be concluded that gross domestic product growth rate, per capita gross domestic product, unemployment, export, default history and the level of economic development significantly affect ratings.
Shi, Huang, and Lee (2017) obtained state-of-the-art results for English and Chinese dependency parsing by combining dynamic-programming implementations of transition-based dependency parsers with a minimal set of bidirectional LSTM features. However, their results were limited to projective parsing. In this paper, we extend their approach to support non-projectivity by providing the first practical implementation of the MH_4 algorithm, an $O(n^4)$ mildly nonprojective dynamic-programming parser with very high coverage on non-projective treebanks. To make MH_4 compatible with minimal transition-based feature sets, we introduce a transition-based interpretation of it in which parser items are mapped to sequences of transitions. We thus obtain the first implementation of global decoding for non-projective transition-based parsing, and demonstrate empirically that it is more effective than its projective counterpart in parsing a number of highly non-projective languages
In this work we describe the system built for the three English subtasks of\nthe SemEval 2016 Task 3 by the Department of Computer Science of the University\nof Houston (UH) and the Pattern Recognition and Human Language Technology\n(PRHLT) research center - Universitat Polit`ecnica de Val`encia: UH-PRHLT. Our\nsystem represents instances by using both lexical and semantic-based similarity\nmeasures between text pairs. Our semantic features include the use of\ndistributed representations of words, knowledge graphs generated with the\nBabelNet multilingual semantic network, and the FrameNet lexical database.\nExperimental results outperform the random and Google search engine baselines\nin the three English subtasks. Our approach obtained the highest results of\nsubtask B compared to the other task participants.\n
Abstract. Patterns of facial reactivity and attentional allocation to emotional facial expressions, and how these are moderated by gaze direction, are not clearly established. Among a sample of undergraduate university students, aged between 17 and 22 years (76% female), corrugator and zygomatic reactivity, as measured by facial electromyography, and attention allocation, as measured by the startle reflex and startle-elicited N100, was examined while viewing happy, neutral, angry and fearful facial expressions, which were presented at either 0- or 30-degree gaze. Results indicated typically observed facial mimicry to happy faces but, unexpectedly, “smiling” facial responses to fearful, and to a lesser extent, angry faces. This facial reactivity was not influenced by gaze direction. Furthermore, emotional facial expressions did not elicit increased attentional allocation. Likewise, matched facial expressions did not elicit increased attentional allocation. Rather, happy and fearful faces with direct (0°) gaze elicited increased controlled attentional allocation, and averted (30°) gaze faces, regardless of emotional expression, elicited preferential, early cortical processing. These findings suggest typical facial mimicry to happy faces, but unexpected facial reactivity to angry and fearful faces, perhaps due to an attempt to regulate social bonds during threat perception. Findings also suggest a divergence in controlled versus preferential, early cortical attentional processing for direct compared to averted gaze faces. These findings relate to young, mostly female, adults attending university. The experiment should be repeated with a larger sample drawn from the general community, with a broader age range and gender balance, and with a stimulus set with validated subjective valence and arousal ratings. This can reduce Type II error and establish normative patterns of facial reactivity and attentional processing of emotional facial expressions with different gaze directions.
Requirement is a formal expression of user’s need. It is the main foundation of any software development project. Natural language (NL) is often used to express and write system requirements specifications as well as user requirements. However, there is a very high probability that more than half natural language requirements can be ambiguous, incomplete and inaccurate. A software engineer can miss-interpret the natural language requirements and can generate an erroneous software model, which finally will lead to project failure. Earlier, we have introduced a prototype tool that provides natural language requirements authoring facilities and consistency checking to assist requirement engineers when working with informal and semi-formal requirements. However, the tool has pattern limitation to support the extraction of the essential requirements from the NL requirements. Therefore this study is aimed to enhance the accuracy and scalability of the tool to capture the essential requirements from the NL requirements. Our approach is to implement lexical analysis and embed an English lexical database where it will serve as a thesaurus in the tool. This tool is expected to be able to find the synonym of the extracted phrases (essential requirements) in the database to match it to the essential interaction pattern (phrases and expressions) in the library. Our future work will focus on the next phase of requirements engineering, which is requirements validation.