Papers reviewed and determined not to be word norm studies. Use the flag icon to report errors or suggest re-inclusion.
16504 papers
The peculiarities of the system of teaching Chinese students Russian as a foreign language are brought about by the specificity of the Chinese language. The article substantiates the system of methods and approaches facilitating efficient acquisition of a foreign language and the culture of its speakers; the author believes that the systemic approach (allowing appropriation of linguistic norms as a system embracing deferent levels) and the process-focused approach (allowing formation of a variable communicative field of work with the text) are the leading ones. The author pays special attention to teaching reading and working with the text. The author also discusses some of the didactic techniques that can help to train students in text comprehension, and the techniques for checking understanding of the text. The author believes that the main accent in the process of teaching Russian and speech culture to Chinese students should be oriented towards enrichment of vocabulary, consolidation of morphology and syntactical constructions characteristic of certain situations of communication, as well as towards development of speech culture and acquisition of the rules of speech etiquette. The author offers a short plan of a lesson as an example of such work.
The lexicon of emotion words is fundamental to interpersonal communication. To examine how emotion word acquisition interacts with societal context, the present study investigated emotion word development in three groups of child Korean users aged 4–13: those who use Korean primarily outside the home as a majority language (MajKCs) or inside the home as a minority language (MinKCs), and those who use Korean both inside and outside the home (KCs). These groups, along with a group of L1 Korean adults, rated the emotional valence of 61 Korean emotion words varying in frequency, valence, and age of acquisition. Results showed KCs, MajKCs, and MinKCs all converging toward adult-like valence ratings by ages 11–13; unlike KCs and MajKCs, however, MinKCs did not show age-graded development and continued to diverge from adults in emotion word knowledge by these later ages. These findings support the view that societal context plays a major role in emotion word development, offering one reason for the intergenerational communication difficulties reported by immigrant families.
In this paper we focus on collocations, which have been studied in computational linguistics since they constitute a key factor when processing natural languages. For instance, they usually represent a challenge in automatic translation because the association of two terms is not easily computed. We proposed that the parser should be provided with a lexical database in order to make more effective the identification of collocations during the parsing process. We assessed this claim by using a corpus of 6’000 sentences retrieved from the British magazine The Economist Espresso. The corpus was parsed twice, first with the collocation detection component turned on and then with it turned off, and to make the comparison the Fips tagger was used. The results showed an improvement of the quality when the parser has access to collocation knowledge.
Norms are essential to the human condition. Whether in the guise of tradition, culture, canon or rules, norms are therefore central to studies in the humanities. This book focuses on Russian language culture of the post-revolutionary and post-Soviet periods, times when norms — linguistic and otherwise — have been eagerly debated, challenged, broken and redefined. Exploring the intersections between linguistic authority and creative response, an international team of scholars examines different realms of linguistic practice (literary fiction, internet slang, literary criticism and aesthetics, writers’ blogs, linguistic play) and various arenas for “talk about talk” (the classroom, blogs, the media, or the courtroom). By combining various approaches and disciplines — linguistics, literary criticism, new media studies — the book as a whole explores the multiplicity of meanings that are accorded to the notion of linguistic norms in the Russian community. The result is both a broad and a detailed picture of important trends in modern Russian language culture.
This chapter concerns the largely ignored phenomenon of inner speech within religious groups. In three major parts, the first considers some linguistic, philosophical and theological approaches used to frame the phenomenon in the past, the second exemplifies and compares inner speech in, Sikh devotional focus on divine words, Zen Buddhist philosophical concern with silence and human being and, one contemporary Christian context concerning prayer. The third part asks how inner speech might relate to the interpersonal dynamics of religious group membership. One contextual background to the phenomenology of inner speech concerns the cultural shift from spoken to silent reading. Sikhism strongly advocates inner speech for devotional goals. Sikhism is both a radically corporate endeavour of a religious community and a profoundly individual pursuit of union with the divine. The sociological interest of this Zen case lies precisely in the attempt at a serious deconstruction of linguistic norms and 'structures'.
This paper discusses the use of ‘by’-phrases in Impersonal Passives in Icelandic. It has been claimed in the literature on Icelandic syntax that ‘by’-phrases that express the agent are not very good or even ungrammatical in Impersonal Passives. The pa-per shows that this point of view oversimplifies the facts because various examples of this pattern can be found in natural data and these examples do not seem to reflect mistakes in linguistic performance. We discuss examples from the Icelandic treebank and from the web and we suggest that ‘by’-phrases are more likely to be used in Impersonal Passives if they involve new information and/or if they are heavy. One of the conclusions of the article is that large and well annotated corpora are important for linguistic research that focuses on rare constructions.
Previous research has shown that good looks, particularly being deemed as attractive or competent-looking, can provide an electoral advantage. There is also evidence to support the notion that more dominant looks are associated with military success as cadets are more likely to rise in the ranks early in their career if they are more dominant-looking. To date, there has been little research into the effect of looks on political leadership success in a non democratic setting. This project explores the effect of facial attractiveness and dominance on the political success of leaders after leading a successful coup d’état. We examine a comprehensive set of coup d'états from 1946 to 2013. Attractiveness and dominance ratings are created via surveys, as in previous research, but with a novel way to control for the potential bias arising from respondent characteristics. Defining political success as taking executive power, longer time-to-office exit, and avoiding constraint on executive power, we find that both dominant and attractive facial features provide distinct advantages for leaders.
The article describes the process of retaining the recessive units in the correctness publications which have been the sources of the codified norm for the last hundred years. This includes the following forms: interesa, czochrze, w Prusiech. Such elements marked i.a. as rare, former, out of use or outdated belonged in the given period to the linguistic norm of some speakers of Polish. Therefore linguists made attempts to retain them in the codified norm at least for some time so that they could serve as evidence of their former correctness. They were used with qualifiers which provided information about potential question concerning the topicality of the recessive elements. Such units met different fates. They were no longer provided or perceived as wrong in the succeeding dictionaries that were examined. They often functioned as recessive until contemporaneity or even returned as equal variants. A detailed typology of these relations was discussed in the article.
The paper deals with the influence of emotionality on ecology of communication. Relations between positive / negative emotions and ecology of communication are studied. Ecology of communication lies in compliance with communicative ethic and emotive linguistic norms where the speaker feels comfortably. On the other hand, negative emotions of communicants form conflict communicative situations. Speech and emotions are viewed in the ecolinguistic paradigm as a specific sphere of person’s activity. It affects consciousness, reflects dominant strategies in the world comprehension, and defines peculiarities of verbal and nonverbal behavior. Сommunicating persons affect each other by emotions. Communicative and ecological function of language is closely linked to emotive function. That is why it is necessary to pay attention to the formation of emotive competence of а lingual personality in addition to the development of its common cultural and speech literacy. Human emotional sphere depends on communicative and ecological competence, linguistic and environmental identity of linguistic personality.
ForFun is a database of linguistic forms and their syntactic functions built with the use of the multi-layer annotated corpora of Czech, the Prague Dependency Treebanks. The purpose of the Prague Database of Forms and Functions (ForFun) is to help the linguists to study the form-function relation, which we assume to be one of the principal tasks of both theoretical linguistics and natural language processing. A prototypical question to be asked is What purposes does a preposition 'po' serve for or What are the linguistic means in the sentence that can express the meaning 'a destination of an action'?. There are almost 1500 distinct forms (besides the 'po' preposition) and 65 distinct functions (besides the 'destination').
Dependency parsing is considered as the state of the art technology for a better information extraction methodology in Natural Language Processing. With the ever-growing need for linguistic analysis for different languages, the demand for multilingual dependency parsing has increased dramatically. In this research work we studied a novel collection of treebanks with homogeneous syntactic dependency annotation [1] for six languages along with other recent techniques in this area. We investigated the possibility of adding new languages in this module and successfully added universal Bangla dependency annotation. Additionally, we combined simple and complex feature representations to improve parsing output.
This chapter provides an overview of language policy and the media by reviewing the state of the art, both in terms of literature and in terms of research. It outlines key terms and their uses, and explains the types of language policy and media. The chapter also provides an overview of disciplinary perspectives on language policy and the media, with a particular focus on the evolution from traditional news to new and social media. It reviews research within a range of national and globalized contexts, and discusses the core areas of status, corpus and acquisition planning. The chapter examines relationship between language policy and the media in two overarching areas: in the chronological transition from nation-states to globalization, and in status, corpus, and acquisition planning. Language policy concerns the production and enforcement of linguistic norms; the reality of policymaking and its implementation is much more complex than such a simplistic label implies.
In this paper, we introduce the novel concept of densely connected layers\ninto recurrent neural networks. We evaluate our proposed architecture on the\nPenn Treebank language modeling task. We show that we can obtain similar\nperplexity scores with six times fewer parameters compared to a standard\nstacked 2-layer LSTM model trained with dropout (Zaremba et al. 2014). In\ncontrast with the current usage of skip connections, we show that densely\nconnecting only a few stacked layers with skip connections already yields\nsignificant perplexity reductions.\n
Sleep disturbance is a core symptom of Pediatric post-traumatic stress disorder (PPTSD). Given the link between sleep and affective processing, disruptions in sleep-dependent processing of emotional material have been suggested to contribute to symptom maintenance. In this study we sought to asses the relationship between sleep and emotion processing in youth with PPTSD relative to healthy control subjects. Seven participants with PTSD (aged 14.5 ± 2.9; CAPS-CA score 64.8 ± 17.8) and three age- and sex- matched controls (aged 12.0 ± 0.8) completed two overnight high-density EEG (256-channel) polysomnography sleep studies. Prior to sleep on night 2, participants rated 70 neutral and 70 negative scenes with respect to level of arousal on a scale of 1–9 using the Self-Assessment Manikin rating system (SAM). The following morning, recall was tested with 100 previously viewed images and 100 new images. Participants again rated images on level of arousal. Sleep: No significant differences were observed between groups for any PSG variable. However, differences in the percent of NREM stage 2 decreased for both groups during sleep on the task night relative to baseline night, with a larger % decrease in the PTSD group (control %2.29 ± 9.04 vs. PTSD %6.34 ± 0.63). Behavior: Recall accuracy (control 78.3 ± 9.8 vs. PTSD 81.6 ± 15.3) did not differ between groups, nor did recall accuracy for emotional images (control 81.3 ± 10.8 vs. PTSD 77.6 ± 14.8). Ratings of arousal did not distinguish groups at encoding. However, ratings of arousal during the recall session for previously viewed negative images was markedly different in each group. As a group, PPTSD children subtly increased arousal ratings of negative remembered-images from evening to morning (mean night, 6.18 ± 2.31, morning 6.34 ± 2.66) while the control group robustly decreased arousal ratings (7.22 ± 2.67, 2.18 ± 1.21). Given our sample size, firm conclusions are impossible. However, our results suggest that sleep quality may adversely impact the process of emotional evaluation in children with PTSD. none
FicTree is a dependency treebank of Czech fiction manually annotated in the format of the analytical layer of the Prague Dependency Trebank. The treebank consists of 12,760 sentences (166,432 tokens). The texts come from eight literary works published in the Czech Republic between 1991 and 2007. The syntactic annotation of the treebank was first performed by two distinct parsers (MSTParser and MaltParser) trained on the PDT training data, then manually corrected. Any differences between the two versions were resolved manually (by another annotator). The corpus is provided in a vertical format, where sentence boundaries are marked with a blank line. Every word form is written on a separate line, followed by five tab-separated attributes: lemma, tag, ID (word index in the sentence), head and deprel (analytical function, afun in the PDT formalism). The texts are shuffled in random chunks of maximum 100 words (respecting sentence boundaries). Each chunk is provided as a separate file, with the suggested division into train, dev and test sets written as file prefix.
This paper provides a new method to correct annotation errors in a treebank. The previous error correction method constructs a pseudo parallel corpus where incorrect partial parse trees are paired with correct ones, and extracts error correction rules from the parallel corpus. By applying these rules to a treebank, the method corrects errors. However, this method does not achieve wide coverage of error correction. To achieve wide coverage, our method adopts a different approach. In our method, we consider that if an infrequent pattern can be transformed to a frequent one, then it is an annotation error pattern. Based on a tree mining technique, our method seeks such infrequent tree patterns, and constructs error correction rules each of which consists of an infrequent pattern and a corresponding frequent pattern. We conducted an experiment using the Penn Treebank. We obtained 1,987 rules which are not constructed by the previous method, and the rules achieved good precision.
This paper describes the submission from the University of Helsinki to the\nshared task on cross-lingual dependency parsing at VarDial 2017. We present\nwork on annotation projection and treebank translation that gave good results\nfor all three target languages in the test set. In particular, Slovak seems to\nwork well with information coming from the Czech treebank, which is in line\nwith related work. The attachment scores for cross-lingual models even surpass\nthe fully supervised models trained on the target language treebank. Croatian\nis the most difficult language in the test set and the improvements over the\nbaseline are rather modest. Norwegian works best with information coming from\nSwedish whereas Danish contributes surprisingly little.\n
This article reflects on the notions of educated linguistic norm, standard linguistic norm and formal written mode in order to establish a productive dialogue with Portuguese teachers, especially those teaching and assessing their students’ written productions. We questioned some common procedures involved in the grading of written productions, such as focusing on grammar aspects of the texts based on obsolete grammar rules. In order to do so, we discussed the often overlooked and controversial senses associated with terms such as “educated” and “standard”, used interchangeably, for instance, by the graders of Brazil’s National High School Exam between 1998 and 2014 when assessing the written part of the exam. Later on, we performed a qualitative analysis of a school essay that presents a few particular morphosyntactic aspects that are in fact used by “educated” Brazilians in their daily and monitored writing practices. We sought to demonstrate that grammar tradition itself already often validates these aspects in their prescription. This way, possible interdictions in formal written contexts, such as in the National High School Exam, arise from an ideology that preaches unattainable standard linguistic rules, resulting in an obsolete, “pure” and overcorrect linguistic standard.Keywords: standard linguistic norm, grading, written production.
This poster will describe a collaboration between nursing facility staff and university researchers to establish individualized music interventions to improve targeted behavioral outcomes for nursing home residents with dementia. We used single-case design to assess intervention effectiveness for each resident. We will present data from 4–5 patients to illustrate how this type of design can be used to inform intervention targeting. We used independent sample t-tests to compare mean percentages of observed affect and behaviors from the Philadelphia Geriatric Center Affect Rating Scale and a modified Passivity in Dementia Scale before and during the music intervention and with a control period of attention only. Preliminary analyses show that for one resident, the music intervention did not produce significant differences in affect or behavior, but for a second resident positive affect (t(21) = -3.784, p =.001) was significantly higher during the intervention, and negative affect (t(19.674) = 2.595, p =.017) and negative behaviors (t(15.453) = 3.242, p =.005) were significantly lower during the intervention. We illustrate how reports from these simple analyses can be used to inform nursing home staff about how to use individualized music with residents who will most benefit from the intervention.
This study explores the question of whether native and non-native listeners, i.e. natives familiar with the language they are judging and non-natives who are not, manage to distinguish a foreign accent from a native accent in the speech of native speakers (NSs) and nonnative speakers (NNSs). Participants included 21 speakers (11 NSs and 10 NNSs who were native Turkish speakers) as well as two listener groups that consisted of 61 Finnish listeners (FLs), and 10 Turkish listeners (TLs) without Finnish experience. This study compares accent ratings by these two listener groups that evaluated the 21 spontaneous speech samples for foreign accent using a 9-point scale. The results showed a very significant difference between the listener groups for the NSs but no significant difference for the NNSs. The difference between the FL and the TL groups was because the FLs managed to distinguish the NSs from the NNSs, but otherwise these two listener groups exercised statistically similar ratings. Therefore, these results demonstrate that the listeners' familiarity with Finnish, the target language, hence listeners' native speaker status strongly affect ratings of foreign accents, since native listeners could distinguish the NSs, whereas non-native listeners could not. The results suggest that listeners' familiarity with the target language plays a much moreprofound role in accent detection than their familiarity with the accent language. Moreover, the results show that contrary to previous research, in the absence of listeners' familiarity with the target language, it is much more challenging to detect a foreign accent. The results also showed that speech rate correlated with the judgments provided by the TLs but not with the judgments provided by the FLs. This result raises the possibility that there are salient universal features of non-native speech such as speech rate that even non-native listeners unfamiliar with the language they are judging utilize while judging a foreign accent.
Study objective: The competitive sports environment can enhance social and cultural pressure towards having ideal body weight in weight-sensitive sports. The close relationship between body image and performance makes the elite athletes vulnerable to eating disorders. Thus, the purpose of this research was to study eating disorders and body image among weight-class elite athletes. Methods: A cross-sectional study was carried out with elite martial arts athletes (Karate, Taekwondo, and Judo) who were considered to be of higher risk for eating disorders. 63 elite martial arts male athletes (18.59 ± 5.29 yrs), and 63 non-athlete persons (17.3 ± 3.4 yrs) were recruited. Body Mass Index (BMI), Waist Hip Ratio (WHR), and Percent Body Fat (PBF) were measured using caliper and meter. Eating Disorder Diagnosis Scale (EDDS) and Body Image Rating Scale (BIRS) were used to study eating disorders and body image among elite martial arts athletes. Results: no sign of clinical EDDS were found among the investigated athletes, and non-athletes. There were significant differences in total score of EDDS (p=0.001), eating disorder and weight concern subscales (respectively p=0.012, p=0.001) in athletes and non-athletes. Furthermore, compared with the non-athlete group, elite athlete group with middle, good, and great body images scored higher on total score and all subscales of EDDS (p ≤ 0.05). Conclusion: The results from our study show the presence of worriment about eating disorder especially body weight and eating concern in elite athletes and the early detection of it may prevent progression to severe eating disorders.
Some attempts have been made in the academic community to carry out an automatic morphological analysis of the Qur'anic text. Among the well-known endeavors in this regard is the morphological annotation of the Quranic Arabic Corpus (QAC) which was carried out in Leeds University, UK. In addition, researchers in the University of Haifa had previously implemented a computational system for the morphological analysis of the Qur'an. More recently, a new Quranic corpus has been built in Mohammed I University in Morocco. To the best of our knowledge, these are the only three studies to produce a morphologically analyzed part-of-speech tagged Qur'an encoded as a structured linguistic database. This paper surveys the morphological analysis in the above-mentioned annotation projects and compares between them to test the quality of their analysis using five criteria related to display of the text in the corpus, word segmentation, morphological disambiguation, part of speech (POS) tag set and manual verification. The paper concludes that the QAC of Leeds and the Quranic corpus of Morocco surpass the Quranic corpus of Haifa with regard to most of these criteria. Furthermore, some additional POS tags for derivative nouns are suggested in a step to reach a more fine-grained tag set that could be proposed for POS tagging of Qur'anic Arabic.
<h3>Introduction</h3><br> Abstract Meaning Representation (AMR) Annotation Release 2.0 was developed by the Linguistic Data Consortium (LDC), <a href="http://www.sdl.com/">SDL/Language Weaver, Inc.</a>, the University of Colorado's <a href="http://clear.colorado.edu/start/index.html">Computational Language and Educational Research</a> group and the <a href="http://www.isi.edu/home">Information Sciences Institute</a> at the University of Southern California. It contains a sembank (semantic treebank) of over 39,260 English natural language sentences from broadcast conversations, newswire, weblogs and web discussion forums. <br> AMR captures “who is doing what to whom” in a sentence. Each sentence is paired with a graph that represents its whole-sentence meaning in a tree-structure. AMR utilizes PropBank frames, non-core semantic roles, within-sentence coreference, named entity annotation, modality, negation, questions, quantities, and so on to represent the semantic structure of a sentence largely independent of its syntax. <br> LDC also released Abstract Meaning Representation (AMR) Annotation Release 1.0 (<a href="../../../LDC2014T12">LDC2014T12</a>). <br> <h3>Data</h3><br> The source data includes discussion forums collected for the DARPA BOLT and DEFT programs, transcripts and English translations of Mandarin Chinese broadcast news programming from China Central TV, Wall Street Journal text, translated Xinhua news texts, various newswire data from NIST OpenMT evaluations and weblog data used in the DARPA GALE program. The following table summarizes the number of training, dev, and test AMRs for each dataset in the release. Totals are also provided by partition and dataset: <br> <table border="1" cellpadding="2"><br> <tbody><br> <tr><br> <td>Dataset</td><br> <td>Training</td><br> <td>Dev</td><br> <td>Test</td><br> <td>Totals</td><br> </tr><br> <tr><br> <td>BOLT DF MT</td><br> <td>1061</td><br> <td>133</td><br> <td>133</td><br> <td>1327</td><br> </tr><br> <tr><br> <td>Broadcast conversation</td><br> <td>214</td><br> <td>0</td><br> <td>0</td><br> <td>214</td><br> </tr><br> <tr><br> <td>Weblog and WSJ</td><br> <td>0</td><br> <td>100</td><br> <td>100</td><br> <td>200</td><br> </tr><br> <tr><br> <td>BOLT DF English</td><br> <td>6455</td><br> <td>210</td><br> <td>229</td><br> <td>6894</td><br> </tr><br> <tr><br> <td>DEFT DF English</td><br> <td>19558</td><br> <td>0</td><br> <td>0</td><br> <td>19558</td><br> </tr><br> <tr><br> <td>Guidelines AMRs</td><br> <td>819</td><br> <td>0</td><br> <td>0</td><br> <td>819</td><br> </tr><br> <tr><br> <td>2009 Open MT</td><br> <td>204</td><br> <td>0</td><br> <td>0</td><br> <td>204</td><br> </tr><br> <tr><br> <td>Proxy reports</td><br> <td>6603</td><br> <td>826</td><br> <td>823</td><br> <td>8252</td><br> </tr><br> <tr><br> <td>Weblog</td><br> <td>866</td><br> <td>0</td><br> <td>0</td><br> <td>866</td><br> </tr><br> <tr><br> <td>Xinhua MT</td><br> <td>741</td><br> <td>99</td><br> <td>86</td><br> <td>926</td><br> </tr><br> <tr><br> <td>Totals</td><br> <td>36521</td><br> <td>1368</td><br> <td>1371</td><br> <td>39260</td><br> </tr><br> </tbody><br> </table><br> <br> For those interested in utilizing a standard/community partition for AMR research (for instance in development of semantic parsers), data in the "split" directory contains 39,260 AMRs split roughly 93%/3.5%/3.5% into training/dev/test partitions, with most smaller datasets assigned to one of the splits as a whole. Note that splits observe document boundaries. The "unsplit" directory contains the same 39,260 AMRs with no train/dev/test partition. <br> <h3>Samples</h3><br> Please view this <a href="desc/addenda/LDC2017T10.xml">sample</a>. <br> <h3>Updates</h3><br> None at this time. <br> <h3>Acknowledgements</h3><br> From University of Colorado <br> We gratefully acknowledge the support of the National Science Foundation Grant NSF: 0910992 IIS:RI: Large: Collaborative Research: Richer Representations for Machine Translation and the support of Darpa BOLT - HR0011-11-C-0145 and DEFT - FA-8750-13-2-0045 via a subcontract from LDC. Any opinions, findings, and conclusions or recommendations expressed in this material are those of the authors and do not necessarily reflect the views of the National Science Foundation, DARPA or the US government. <br> From Information Sciences Institute (ISI) <br> Thanks to NSF (IIS-0908532) for funding the initial design of AMR, and to DARPA MRP (FA-8750-09-C-0179) for supporting a group to construct consensus annotations and the AMR Editor. The initial AMR bank was built under DARPA DEFT FA-8750-13-2-0045 (PI: Stephanie Strassel; co-PIs: Kevin Knight, Daniel Marcu, and Martha Palmer) and DARPA BOLT HR0011-12-C-0014 (PI: Kevin Knight). <br> From Linguistic Data Consortium (LDC) <br> This material is based on research sponsored by Air Force Research Laboratory and Defense Advance Research Projects Agency under agreement number FA8750-13-2-0045. The U.S. Government is authorized to reproduce and distribute reprints for Governmental purposes notwithstanding any copyright notation thereon. The views and conclusions contained herein are those of the authors and should not be interpreted as necessarily representing the official policies or endorsements, either expressed or implied, of Air Force Research Laboratory and Defense Advanced Research Projects Agency or the U.S. Government. <br> We gratefully acknowledge the support of Defense Advanced Research Projects Agency (DARPA) Machine Reading Program under Air Force Research Laboratory (AFRL) prime contract no. FA8750-09-C-0184 Subcontract 4400165821. Any opinions, findings, and conclusion or recommendations expressed in this material are those of the author(s) and do not necessarily reflect the view of the DARPA, AFRL, or the US government. <br> From Language Weaver (SDL) <br> This work was partially sponsored by DARPA contract HR0011-11-C-0150 to LanguageWeaver Inc. Any opinions, findings, and conclusion or recommendations expressed in this material are those of the author(s) and do not necessarily reflect the view of the DARPA or the US government. </br> Portions © 2002-2005, 2007-2008 Agence France Presse, © 2007 Al Ahram, © 2007 Al Hayat, © 2007 Al-Quds Al-Arabi, © 2007 Asharq Al-Awsat, © 2007 An Nahar, © 2007 Assabah, © 2002-2008 The Associated Press, © 2003-2004, 2007-2008 Central News Agency (Taiwan), © 1997, 2004-2007 China Central TV, © 2007 China Military Online, © 2007 Chinanews.com, © 1987-1989 Dow Jones & Company, Inc., © 2007 Guangming Daily, © 1995, 2003, 2007-2008 Los Angeles Times-Washington Post News Service, Inc., © 2002, 2004-2005, 2007-2008 New York Times, © 1994-1998, 2001-2008 Xinhua News Agency, © 2014, 2017 Language Weaver, Inc., © 2014, 2017 University of Colorado, © 2014, 2017 University of Southern California, © 2003, 2005, 2006, 2007, 2009, 2011, 2013, 2014, 2017 Trustees of the University of Pennsylvania
Precisely understanding the value and perception of consumers has long been recognized as essential elements of every market-oriented company's core business strategy. For this reason, customers' affection, as the basis for the formation of human values and judgment, should be considered carefully to strengthen the product quality and competitiveness. However, conventional product design places more attention to functional attributes and requires survey process to collect customers' evaluations, neglecting the in-depth study of the underlying associations between design properties and consumers' emotions based on the abundant online consumer response resources. To improve the deficiency, this study was proposed to develop a product affective properties identification approach. Particularly, data mining techniques (e.g. web mining, text mining) are applied to capture online product review resources. Considering the characteristics of user/consumer responses and evaluations, ontology is utilized to assist in the semantic analysis. With the help of product knowledge hierarchy and electronic lexical database, product properties, which can evoke consumers' affect, can be identified. Furthermore, the identified product affective properties are prioritized to provide designers with important reference for future improvement on the product. To illustrate the proposed approach, a pilot study based on iPhone 7 was conducted, in which the influential affective properties have been identified, and a ranking of them has been mapped out.
Makary Kotoko, a Chadic language spoken in the flood plain directly south of Lake Chad in Cameroon, has an estimated 16,000 speakers. An analysis of a lexical database for the language shows that of the 3000 or so distinct lexical entries in the database, almost 1/3 (916 items) have been identified as borrowed from other languages in the region. The majority of the borrowings come from Kanuri, a Nilo-Saharan language of Nigeria, with an estimated number of speakers ranging from 1 to 4 million. In this article I first present the number of borrowings specifically from Kanuri relative to the total number of borrowed items in Makary Kotoko, and the lexical/grammatical categories in Makary Kotoko that have incorporated Kanuri borrowings. I follow this by presenting the linguistic evidence which not only suggests a possible time frame for when the borrowings from Kanuri came into Makary Kotoko, but also supports the idea that this is essentially a case of completed language contact. After discussing the lexical and grammatical borrowings from Kanuri into Makary Kotoko in detail, I explore the limited evidence in Makary Kotoko for lexical and grammatical ‘calquing’ from Kanuri, resulting in almost no structural diffusion from Kanuri into Makary Kotoko. I finish with a few proposals as to why this is the case in this instance of language contact in the Lake Chad basin.
This paper makes an explorative study of part-of-speech (POS) motif and dependency motif using the treebanks of deaf students' writing in three learning stages. Firstly, the fitting to the POS motif data with Zip-Mandelbrot distribution yields a good result. However, the dependency motif distribution fits well by the Popescu-Altmann-Köhler function instead. Then, a further observation of the top-ranking motifs finds that the most frequent motifs are formed among the most frequently-used word classes and dependencies. Lastly, the hapax-type ratio for both motif types shows an extremely heavy proportion of hapax, which may indicate the discreteness of sentence structure. Through the three analyses, some changes of sentence structure and the development of syntax are witnessed across the language learning course.
Quantized Neural Networks (QNNs), which use low bitwidth numbers for representing parameters and performing computations, have been proposed to reduce the computation complexity, storage size and memory usage. In QNNs, parameters and activations are uniformly quantized, such that the multiplications and additions can be accelerated by bitwise operations. However, distributions of parameters in Neural Networks are often imbalanced, such that the uniform quantization determined from extremal values may under utilize available bitwidth. In this paper, we propose a novel quantization method that can ensure the balance of distributions of quantized values. Our method first recursively partitions the parameters by percentiles into balanced bins, and then applies uniform quantization. We also introduce computationally cheaper approximations of percentiles to reduce the computation overhead introduced. Overall, our method improves the prediction accuracies of QNNs without introducing extra computation during inference, has negligible impact on training speed, and is applicable to both Convolutional Neural Networks and Recurrent Neural Networks. Experiments on standard datasets including ImageNet and Penn Treebank confirm the effectiveness of our method. On ImageNet, the top-5 error rate of our 4-bit quantized GoogLeNet model is 12.7\%, which is superior to the state-of-the-arts of QNNs.
This paper intends to investigate Greek influence on the Latin sound change [b] > [β] suggested occasionally in the literature by surveying not only the relevant linguistic data of Latin/Romance and Koine/Modern Greek but also the relevant literature and by involving and analyzing data sets recorded from 18 Roman provinces and the city of Rome in the Computerized Historical Linguistic Database of the Latin Inscriptions of the Imperial Age (cf. http://lldb.elte.hu/) by a more differentiated phonological approach considering external sandhi rules and in a chronological distribution more detailed than any applied before. In the end, the influence of Greek has been evidenced at least for some areas and especially for the early period (1 st –3 rd century AD), which is more important in this respect than the late period (4 th –6 th century AD), since then the merger can also be explained by developments in Latin itself beside a supposed external influence.
Dependency parses are an effective way to inject linguistic knowledge into many downstream tasks, and many practitioners wish to efficiently parse sentences at scale. Recent advances in GPU hardware have enabled neural networks to achieve significant gains over the previous best models, these models still fail to leverage GPUs' capability for massive parallelism due to their requirement of sequential processing of the sentence. In response, we propose Dilated Iterated Graph Convolutional Neural Networks (DIG-CNNs) for graphbased dependency parsing, a graph convolutional architecture that allows for efficient end-to-end GPU parsing. In experiments on the English Penn TreeBank benchmark, we show that DIG-CNNs perform on par with some of the best neural network parsers.
GF (Grammatical Framework) and UD (Universal Dependencies) are two different approaches using shared syntactic descriptions for multiple languages.GF is a categorial grammar approach using abstract syntax trees and hand-written grammars, which define both generation and parsing.UD is a dependency approach driven by annotated treebanks and statistical parsers.In closer study, the grammatical descriptions in these two approaches have turned out to be very similar, so that it is possible to map between them, to the benefit of both.The demo presents a recent addition to the GF web demo, which enables the construction and visualization of UD trees in 32 languages.The demo exploits another new functionality, also usable as a command-line tool, which converts dependency trees in the CoNLL format to high-resolution L A T E Xand SVG graphics.
Modulations of the voice convey affect, and the precise mapping of voice-to-affect may vary for different languages. However, affect-related modulations occur relative to the baseline affect-neutral voice, which tends to differ from language to language. Little is known about the characteristic long-term voice settings for different languages, and how they influence the use of voice quality to signal affect. In this paper, data from a voice-to-affect perception test involving Russian, English, Spanish and Japanese subjects is re-examined to glean insights concerning likely baseline settings in these languages. The test used synthetic stimuli with different voice qualities (modelled on a male voice), with or without extreme f0 contours as might be associated with affect. Cross-language differences in affect ratings for modal and tense voice suggest that the baseline in Spanish and Japanese is inherently tenser than in Russian and English, and that as a corollary, tense voice serves as a more potent cue to high-activation affects in the latter languages. A relatively tenser baseline in Japanese and Spanish is further suggested by the fact that tense voice can be associated with intimate, a low activation state, just as readily as with the high-activation state interested.
Past research has shown that an individual's feelings at any given moment reflect currently experienced stimuli as well as internal representations of similar past experiences. However, anxious individuals' affective reactions to streams of interrelated valenced information (vs. reactions to static stimuli that are arguably less ecologically valid) are rarely tracked. The present study provided a first examination of the newly developed Tracking Affect Ratings Over Time (TAROT) task to continuously assess anxious individuals' affective reactions to streams of information that systematically change valence. Undergraduate participants (N = 141) completed the TAROT task in which they listened to narratives containing positive, negative, and neutral physically- or socially-relevant events, and indicated how positive or negative they felt about the information they heard as each narrative unfolded. The present study provided preliminary evidence for the validity and reliability of the task. Within scenarios, participants higher (vs. lower) in anxiety showed many expected negative biases, reporting more negative mean ratings and overall summary ratings, changing their pattern of responding more quickly to negative events, and responding more negatively to neutral events. Furthermore, individuals higher (vs. lower) in anxiety tended to report more negative minimums during and after positive events, and less positive maximums after negative events. Together, findings indicate that positive events were less impactful for anxious individuals, whereas negative experiences had a particularly lasting impact on future affective responses. The TAROT task is able to efficiently capture a number of different cognitive biases, and may help clarify the mechanisms that underlie anxious individuals' biased negative processing. (PsycINFO Database Record
Norms are essential to the human condition. Whether in the guise of tradition, culture, canon or rules, norms are therefore central to studies in the humanities. This book focuses on Russian language culture of the post-revolutionary and post-Soviet periods, times when norms — linguistic and otherwise — have been eagerly debated, challenged, broken and redefined. Exploring the intersections between linguistic authority and creative response, an international team of scholars examines different realms of linguistic practice (literary fiction, internet slang, literary criticism and aesthetics, writers’ blogs, linguistic play) and various arenas for “talk about talk” (the classroom, blogs, the media, or the courtroom). By combining various approaches and disciplines — linguistics, literary criticism, new media studies — the book as a whole explores the multiplicity of meanings that are accorded to the notion of linguistic norms in the Russian community. The result is both a broad and a detailed picture of important trends in modern Russian language culture.
This article deals with the recursive compounding of Old English nouns, adjectives, verbs and adverbs. It addresses the question of the textual occurrences of the compounds of Old English by means of a corpus analysis based on the Dictionary of Old English Corpus. The data of qualitative analysis have been retrieved from the lexical database of Old English Nerthus. The analysis shows that the nominal, adjectival and adverbial compounds of Old English can be recursive. Nominal compounding allows double recursivity, whereas adjectival and adverbial compounding do not. The conclusion is reached that both the type and token frequencies of recursive compounds are very low; and recursive compounds from the adjectival class are more exocentric as regards categorisation.
Abstract Marketing and advertising texts are important sites for signaling and also contributing to sociolinguistic change. For example, advertising language can maintain, challenge, or create new linguistic norms as well as document usage and practices. Linguistic fetish and styling in advertisements both rely on and contribute to enregisterment processes. This chapter explores one particular example of brand styling using ‘fake French’ and the French linguistic fetish in advertising, namely an advertisement for a cider produced by European lager brand Stella Artois. The analysis highlights how the brand is in fact making use of a metaparodic frame, simultaneously asserting and parodying its stylized ‘fake French’ style, and how the consumer is complicit in co-constructing this metaparody.
In education, essay is considered as the best tool to evaluate student’s high order thinking and understanding. In the other hand, manual processing and grading essay answers by a teacher need much time and tending to subjectivity grading. Meanwhile automatic essay grading in e-learning system find the difficulties in comparing model or key answer to student’s answer because student’s can answer the question with so various way. That means a right answer also can be so various, for they have same semantic meaning. This paper proposed automatic essay grading using Latent Semantic Analysis. But before the texts being scored, they will be pre-processed using stop words removal and synonyms checking. Calibration process implemented for dealing with the various possible right answer and help to simplify the term matrix. Implementation of this approach using Java Programming Language and WordNet as lexical database for searching the synonyms of every given words. The accuracy obtained by this method is 54.9289%.
The Web has become a tremendously huge data source hidden under linked documents. A significant number of Web documents include HTML tables generated dynamically from relational databases. Often, there is no direct public access to the databases themselves. On the other hand, RDF (Resource Description Framework) gives an efficient mechanism to represent directly data on the Web based on a Web-scalable architecture for identification and interpretation of terms. This leads to the concept of Linked Data on the Web. To allow direct access to data on the Web as Linked Data, we propose in this paper an approach to transform HTML tables into RDF triples. It consists of three main phases: refining, pre-treatment and mapping. The whole process is assisted by a domain ontology and the WordNet lexical database. A tool called Htab2RDF has been implemented. Experiments have been carried out to evaluate and show efficiency of the proposed approach.
In a paper entitled “Against markedness (and what to replace it with)”, Haspelmath argues “that the term ‘markedness’ is superfluous”, and that frequency asymmetries often explain structural (un)markedness asymmetries (Haspelmath 2006). We investigate whether this argument applies to Object and Verb orders in main (VO, marked) and subordinate (OV, unmarked) clauses of spoken and written German and Dutch, using English (without VO/OV alternation) as control. Frequency counts from six treebanks (three languages, two output modalities) do not support Haspelmath’s proposal. However, they reveal an unexpected phenomenon, most prominently in spoken Dutch and German: a small set of extremely high-frequent finite verbs with unspecific meanings populates main clauses much more densely than subordinate clauses. We suggest these verbs accelerate the start-up of grammatical encoding, thus facilitating sentence-initial output fluency.
Being less resource languages, Indian-Indian and English-Indian language MT system developments faces the difficulty to translate various lexical phenomena. In this paper, we present our work on a comparative study of 440 phrase-based statistical trained models for 110 language pairs across 11 Indian languages. We have developed 110 baseline Statistical Machine Translation systems. Then we have augmented the training corpus with Indowordnet synset word entries of lexical database and further trained 110 models on top of the baseline system. We have done a detailed performance comparison using various evaluation metrics such as BLEU score, METEOR and TER. We observed significant improvement in evaluations of translation quality across all the 440 models after using the Indowordnet. These experiments give a detailed insight in two ways: (1) usage of lexical database with synset mapping for resource poor languages (2) efficient usage of Indowordnet sysnset mapping. More over, synset mapped lexical entries helped the SMT system to handle the ambiguity to a great extent during the translation.
Lexicon-based methods using syntactic rules for polarity classification rely on parsers that are dependent on the language and on treebank guidelines. Thus, rules are also dependent and require adaptation, especially in multilingual scenarios. We tackle this challenge in the context of the Iberian Peninsula, releasing the first symbolic syntax-based Iberian system with rules shared across five official languages: Basque, Catalan, Galician, Portuguese and Spanish. The model is made available. 1
BACKGROUND: Diffusion-weighted MRI has been proposed as a new technique for imaging synovitis without intravenous contrast application. We investigated diagnostic utility of multi-shot readout-segmented diffusion-weighted MRI (multi-shot DWI) for synovial imaging of the knee joint in patients with juvenile idiopathic arthritis (JIA). METHODS: were separately rated by three independent blinded readers at different levels of expertise for the presence and the degree of synovitis on a modified 5-item Likert scale along with the level of subjective diagnostic confidence. RESULTS: Fourteen (44%) patients had active synovitis and joint effusion, nine (28%) patients showed mild synovial enhancement not qualifying for arthritis and another nine (28%) patients had no synovial signal alterations on contrast-enhanced imaging. Ratings by the 1st reader on contrast-enhanced MRI and on DWI showed substantial agreement (κ = 0.74). Inter-observer-agreement was high for diagnosing, or ruling out, active arthritis of the knee joint on contrast-enhanced MRI and on DWI, showing full agreement between 1st and 2nd reader and disagreement in one case (3%) between 1st and 3rd reader. In contrast, ratings in cases of absent vs. little synovial inflammation were markedly inconsistent on DWI. Diagnostic confidence was lower on DWI, compared to contrast-enhanced imaging. CONCLUSION: Multi-shot DWI of the knee joint is feasible in routine imaging and reliably diagnoses, or rules out, active arthritis of the knee joint in paediatric patients without the need of gadolinium-based i.v. contrast injection. Possibly due to "T2w shine-through" artifacts, DWI does not reliably differentiate non-inflamed joints from knee joints with mild synovial irritation.
Evolutionary theory was applied to Reeder and Brewer's schematic theory and Trafimow's affect theory to extend this area of research with five new predictions involving affect and ability attributions, comparing morality and ability attributions, gender differences, and reaction times for affect and attribution ratings. The design included a 2 (Trait Dimension Type: HR, PR) × 2 (Behavior Type: morality, ability) × 2 (Valence: positive, negative) × 2 (Replication: original, replication) × 2 (Sex: female or male actor) × 2 (Gender: female or male participant) × 2 (Order: attribution portion first, affect portion first) mixed design. All factors were within participants except the order and participant gender. Participants were presented with 32 different scenarios in which an actor engaged in a concrete behavior after which they made attributions and rated their affect in response to the behavior. Reaction times were measured during attribution and affect ratings. In general, the findings from the experiment supported the new predictions. Affect was related to attributions for both morality and ability related behaviors. Morality related behaviors received more extreme attribution and affect ratings than ability related behaviors. Female actors received stronger attribution and affect ratings for diagnostic morality behaviors compared to male actors. Male and female actors received similar attribution and affect ratings for diagnostic ability behaviors. Diagnostic behaviors were associated with lower reaction times than non-diagnostic behaviors. These findings demonstrate the utility of evolutionary theory in creating new hypotheses and empirical findings in the domain of attribution.
Semantic Role Labeling (SRL) is a Natural Language Processing task that enables the detection of events described in sentences and the participants of these events. For Brazilian Portuguese (BP), there are two studies recently concluded that perform SRL in journalistic texts. [1] obtained F1-measure scores of 79.6, using the PropBank.Br corpus, which has syntactic trees manually revised, [8], without using a treebank for training, obtained F1-measure scores of 68.0 for the same corpus. However, the use of manually revised syntactic trees for this task does not represent a real scenario of application. The goal of this paper is to evaluate the performance of SRL on revised and non-revised syntactic trees using a larger and balanced corpus of BP journalistic texts. First, we have shown that [1]'s system also performs better than [8]'s system on the larger corpus. Second, the SRL system trained on non-revised syntactic trees performs better over non-revised trees than a system trained on gold-standard data.
We propose a shared task on multilingual SurfaceRealization, i.e., on mapping unorderedand uninflected universal dependency trees tocorrectly ordered and inflected sentences in anumber of languages. A second deeper inputwill be available in which, in addition,functional words, fine-grained PoS and morphologicalinformation will be removed fromthe input trees. The first shared task on SurfaceRealization was carried out in 2011 witha similar setup, with a focus on English. Wethink that it is time for relaunching such ashared task effort in view of the arrival of UniversalDependencies annotated treebanks fora large number of languages on the one hand,and the increasing dominance of Deep Learning,which proved to be a game changer forNLP, on the other hand.
This paper presents a novel neural machine translation model which jointly learns translation and source-side latent graph representations of sentences. Unlike existing pipelined approaches using syntactic parsers, our end-to-end model learns a latent graph parser as part of the encoder of an attention-based neural machine translation model, and thus the parser is optimized according to the translation objective. In experiments, we first show that our model compares favorably with state-of-the-art sequential and pipelined syntax-based NMT models. We also show that the performance of our model can be further improved by pre-training it with a small amount of treebank annotations. Our final ensemble model significantly outperforms the previous best models on the standard English-to-Japanese translation dataset.
A novel character-level neural language model is proposed in this paper. The proposed model incorporates a biologically inspired temporal hierarchy in the architecture for representing multiple compositions of language in order to handle longer sequences for the character-level language model. The temporal hierarchy is introduced in the language model by utilizing a Gated Recurrent Neural Network with multiple timescales. The proposed model incorporates a timescale adaptation mechanism for enhancing the performance of the language model. We evaluate our proposed model using the popular Penn Treebank and Text8 corpora. The experiments show that the use of multiple timescales in a Neural Language Model (NLM) enables improved performance despite having fewer parameters and with no additional computation requirements. Our experiments also demonstrate the ability of the adaptive temporal hierarchies to represent multiple compositonality without the help of complex hierarchical architectures and shows that better representation of the longer sequences lead to enhanced performance of the probabilistic language model.