Papers reviewed and determined not to be word norm studies. Use the flag icon to report errors or suggest re-inclusion.
18265 papers
Recurrent neural networks (RNNs), such as long short-term memory networks (LSTMs), serve as a fundamental building block for many sequence learning tasks, including machine translation, language modeling, and question answering. In this paper, we consider the specific problem of word-level language modeling and investigate strategies for regularizing and optimizing LSTM-based models. We propose the weight-dropped LSTM which uses DropConnect on hidden-to-hidden weights as a form of recurrent regularization. Further, we introduce NT-ASGD, a variant of the averaged stochastic gradient method, wherein the averaging trigger is determined using a non-monotonic condition as opposed to being tuned by the user. Using these and other regularization strategies, we achieve state-of-the-art word level perplexities on two data sets: 57.3 on Penn Treebank and 65.8 on WikiText-2. In exploring the effectiveness of a neural cache in conjunction with our proposed model, we achieve an even lower state-of-the-art perplexity of 52.8 on Penn Treebank and 52.0 on WikiText-2.
Because of the importance of the information conveyed by the clinical documents and owing to the large quantity of raw texts produced in the healthcare system, it became a determinant challenge, in the NLP research field, to arrange the extraction and the management of meaningful data, starting from real text occurrences. In this paper we approach a corpus of 5000 medical diagnoses with sophisticated linguistic and computational devices, which are able to access the semantic dimension of words and sentences contained in it. Our morphosemantic method is grounded on a list of neoclassical formative elements pertaining to the medical domain which has been used for the automatic creation and population of medical lexical resources. The outcomes of this work are automatically built electronic dictionaries and thesauri and an annotated corpus for the NLP in the medical domain.
Because of the importance of the information conveyed by the clinical documents and owing to the large quantity of raw texts produced in the healthcare system, it became a determinant challenge, in the NLP research field, to arrange the extraction and the management of meaningful data, starting from real text occurrences. In this paper we approach a corpus of 5000 medical diagnoses with sophisticated linguistic and computational devices, which are able to access the semantic dimension of words and sentences contained in it. Our morphosemantic method is grounded on a list of neoclassical formative elements pertaining to the medical domain which has been used for the automatic creation and population of medical lexical resources. The outcomes of this work are automatically built electronic dictionaries and thesauri and an annotated corpus for the NLP in the medical domain.
This paper presents Autobank, a prototype tool for constructing a widecoverage Minimalist Grammar (MG) The front end of the tool is a graphical user interface which facilitates the rapid development of a seed set of MG trees via manual reannotation of PTB preterminals with MG lexical categories. The system then extracts various dependency mappings between the source and target trees, and uses these in concert with a non-statistical MG parser to automatically reannotate the rest of the corpus. Autobank thus enables deep treebank conversions (and subsequent modifications) without the need for complex transduction algorithms accompanied by cascades of ad hoc rules; instead, the locus of human effort falls directly on the task of grammar construction itself.
For the data-driven approach in Natural Language Processing (NLP) applications, good quality linguistic resources considered as a main factor to obtain good results. Although Arabic language is one of the main languages in the world, it is considered as low-resourced language in term of good quality and free linguistic resources. This work presents the first stage of building a new open source dependency treebank for Arabic language. It describes the prototype of the new dependency treebank that are inspired by Lexical Functional Grammar (LFG). This paper shows a main approach of developing a newly treebank and put lines the future work needed to complete this novel linguistic resource.
The Universal Dependencies Project1\n(Nivre, [9]; Nivre et al., [10]) is an\nongoing effort towards creating a set of harmonised dependency treebanks\nthat are annotated and structured according to universal guidelines. This paper reports on the addition of morphological features to the Irish Universal\nDependencies Treebank (IUDT). Our feature set subscribes to the feature inventory of the UD Project and has been mapped from Irish morpho-syntactic\ntags – the output of a Finite State Morphological Analyser for Irish (Uí Dhonnchadha and van Genabith [16]). Irish, a Celtic language, has some relatively unusual morphological features that require language-specific labels\nnot covered by the universal feature set. In this paper, we summarise the\nIrish-specific features that we have added to this set by explaining the linguistic properties that they each describe. We also report on the first parsing\nexperiments using the IUDT by assessing the effect that the inclusion of morphological features has on parsing accuracy.
The amount of data that is available for research grows rapidly, yet technology to efficiently interpret and excavate these data lags behind. For instance, when using large treebanks for linguistic research, the speed of a query leaves much to be desired. GrETEL Indexing, or GrInding, tackles this issue. The idea behind GrInding is to make the search space as small as possible before actually starting the treebank search, by pre-processing the treebank at hand. We recursively divide the treebank into smaller parts, called subtree-banks, which are then converted into database files. All subtree-banks are organized according to their linguistic dependency pattern, and labeled as such. Additionally, general patterns are linked to more specific ones. By doing so, we create millions of databases, and given a linguistic structure we know in which databases that structure can occur, leading up to a significant efficiency boost. We present the results of a benchmark experiment, testing the effect of the GrInding procedure on the SoNaR-500 treebank.
Accurate sentiment analysis models encode the sentiment of words and their combinations to predict the overall sentiment of a sentence. This task becomes challenging when applied to morphologically rich languages (MRL). In this article, we evaluate the use of deep learning advances, namely the Recursive Neural Tensor Networks (RNTN), for sentiment analysis in Arabic as a case study of MRLs. While Arabic may not be considered the only representative of all MRLs, the challenges faced and proposed solutions in Arabic are common to many other MRLs. We identify, illustrate, and address MRL-related challenges and show how RNTN is affected by the morphological richness and orthographic ambiguity of the Arabic language. To address the challenges with sentiment extraction from text in MRL, we propose to explore different orthographic features as well as different morphological features at multiple levels of abstraction ranging from raw words to roots. A key requirement for RNTN is the availability of a sentiment treebank; a collection of syntactic parse trees annotated for sentiment at all levels of constituency and that currently only exists in English. Therefore, our contribution also includes the creation of the first Arabic Sentiment Treebank (A r S en TB) that is morphologically and orthographically enriched. Experimental results show that, compared to the basic RNTN proposed for English, our solution achieves significant improvements up to 8% absolute at the phrase level and 10.8% absolute at the sentence level, measured by average F1 score. It also outperforms well-known classifiers including Support Vector Machines, Recursive Auto Encoders, and Long Short-Term Memory by 7.6%, 3.2%, and 1.6% absolute respectively, all models being trained with similar morphological considerations.
Multiword expressions (MWEs) are linguistic objects containing two or more words and showing idiosyncratic behavior at different levels. Treebanks with annotated MWEs enable studies of such properties, as well as training and evaluation of MWE-aware parsers. However, few treebanks contain full-fledged MWE annotations. We show how this gap can be bridged in Polish by projecting 3 MWE resources on a constituency treebank.
International audience
In this paper, we describe a method for mapping the phonological feature location of Swedish Sign Language (SSL) signs to the meanings in the Swedish semantic dictionary SALDO.By doing so, we observe clear differences in the distribution of meanings associated with different locations on the body.The prominence of certain locations for specific meanings clearly point to iconic mappings between form and meaning in the lexicon of SSL, which pinpoints modalityspecific properties of the visual modality.
In this paper, we propose a general methodology for designing semantic role/relation system. Based on this methodology, we establish a succinct semantic relation system for consecutive predicative constituents for Chinese, which includes serial verb construction, discourse construction, and other constructions describing serial events. This semantic relation system has 13 middle-level classes and 24 fine-grained sub-classes in contrast to conventional complex classification schemes and meets the uniqueness and completeness criteria of semantic relation identification. We conduct experiments on our system by training four annotators in 1 h to label 200 sentences extracted from Sinica Treebank and HIT-CDTB. With the help of our predesigned feature-based decision tree and a connective markers checklist, the annotators attain a 73% consistency with the reference standard annotation and substantial agreement by Cohen’s kappa coefficient for middle-level labeling. By analyzing the labeling error types, we slightly revise our classification scheme and propose six methods to improve the classification and labeling system, hoping to achieve even better agreement in the future.
This paper describes the use of GrETEL for linguistic research. GrETEL is a linguistic search tool that enables users to look up constructions in syntactically annotated corpora or <i>treebanks</i>. It provides online access to the data, allowing users to query a treebank using either an example sentence or an XPath expression in order to look for similar constructions. A major asset of GrETEL is that it enables non-technical users to consult treebanks in a user-friendly way, which is also in line with the main CLARIN goal of applying the results of speech and language technology to research in the humanities and the social sciences. Besides a description of the querying procedure in GrETEL, this paper presents a selection of research in Dutch syntax and semantics that has been carried out using GrETEL. Furthermore, an overview is given of further developments.
OBJECTIVE: Intrusive negative affect and concurrent deficits in positive affect are hallmarks of posttraumatic stress disorder (PTSD). We sought to further extend the extant literature by exploring the experience of negative affect intrusion upon potentially positive situations (here termed, "negative affect interference," NAI). METHOD: Two studies with adults endorsing at least 1 traumatic event (Study 1, N = 294; Study 2, N = 286) examined how NAI and more general hedonic deficits (HD) relate to psychopathology, trauma exposure characteristics, and ratings of normed visual stimuli. RESULTS: Study 1 found that NAI and HD were positively correlated with PTSD symptoms and childhood trauma, and NAI incremented over depressive symptoms in predicting PTSD severity. Study 2 results indicated additional strong positive correlations between NAI and HD and anhedonia, affect regulation problems, negative affect, and neuroticism. NAI and HD were found to increment over trait NA in predicting PTSD symptoms. Individuals endorsing elevated NAI and HD rated positively valenced pictures (including food and erotic images) as less arousing, although not more negative. CONCLUSIONS: These findings expand conceptualizations of anhedonia and emotional numbing by drawing attention to negative affect in otherwise positive contexts. (PsycINFO Database Record
Public textual cyberbullying has become one of the most prevalent issues associated with\nonline safety of young people, particularly on social networks. To address this issue, we\nargue that the boundaries of what constitutes public textual cyberbullying needs to be first\nidentified and a corresponding linguistically motivated definition needs to be advanced.\nThus, we propose a definition of public textual cyberbullying that contains three necessary\nand sufficient elements: the personal marker, the dysphemistic element and the\ncyberbullying link between the previous two elements. Subsequently, we argue that one of\nthe cornerstones in the overall process of mitigating the effects of cyberbullying is the\ndesign of a cyberbullying lexical database that specifies what linguistic and cyberbullying\nspecific information is relevant to the detection process. In this vein, we propose a novel\ncyberbullying lexical database based on the definition of public textual cyberbullying. The\noverall architecture of our cyberbullying lexical database is determined semantically, and, in\norder to facilitate cyberbullying detection, the lexical entry encapsulates two new semantic\ndimensions that are derived from our definition: cyberbullying function and cyberbullying\nreferential domain. In addition, the lexical entry encapsulates other semantic and syntactic \ninformation, such as sense and syntactic category, information that, not only aids the\nprocess of detection, but also allows us to expand the cyberbullying database using\nWordNet (Miller, 1993).
Similarity calculation between business process models has an important role in managing repository of business process model. One of its uses is to facilitate the searching process of models in the repository. Business process similarity is closely related to semantic string similarity. Semantic string similarity is usually performed by utilizing a lexical database such as WordNet to find the semantic meaning of the word. The activity name of the business process uses terms that specifically related to the business field. However, most of the terms in business domain are not available in WordNet. This case would decrease the semantic analysis quality of business process model. Therefore, this study would try to improve semantic analysis of business process model. We present a new lexical database called B-BabelNet. B-BabelNet is a lexical database built by using the same method in BabelNet. We attempt to map the Wikipedia page to WordNet database but only focus on the word related to the domain of business. Also, to enrich the vocabulary in the business domain, we also use terms in the business-specific online dictionary (businessdictionary.com). We utilize this database to do word sense disambiguation process on business process model activity’s terms. The result from this study shows that the database can increase the accuracy of the word sense disambiguation process especially in particular terms related to the business and industrial domains.
This paper introduces the Universal Dependencies Treebank for Slovenian. We overview the existing dependency treebanks for Slovenian and then detail the conversion of the ssj200k treebank to the framework of Universal Dependencies version 2. We explain the mapping of part-of-speech categories, morphosyntactic features, and the dependency relations, focusing on the more problematic language-specific issues. We conclude with a quantitative overview of the treebank and directions for further work.
Discourse parsing has long been treated as a stand-alone problem independent from constituency or dependency parsing. Most attempts at this problem are pipelined rather than end-to-end, sophisticated, and not self-contained: they assume goldstandard text segmentations (Elementary Discourse Units), and use external parsers for syntactic features.
We present a method for automatically converting the Dutch Lassy Small treebank, a phrasal dependency treebank, to UD. All of the information required to produce accurate UD annotation appears to be available in the underlying annotation. However, we also note that the close connection between POS-tags and dependency labels that is present in UD is missing in the Lassy treebanks. As a consequence, annotation decisions in the Dutch<br/>data for such phenomena as nominalization<br/>and clausal complements of prepositions<br/>seem to differ to some extent from comparable data in English and German. Because the conversion is automatic, we can now also compare three state-of-theart dependency parsers trained on UD Lassy Small with Alpino, a hybrid Dutch parser which produces output that is compatible with the original Lassy annotations.
The study shows the results of an inventory of place names connected to Arbutus unedo L., a Mediterranean species, widespread throughout Sardinia. The main aim was to compare the past distribution of place names, referring to the strawberry tree, to the current distribution of the species on the island. In addition, we investigated the meaning and the diversity of these local place names in the various communities. The result was a collection of 432 phyto-toponyms. 248 of them were used for an analysis of their distribution in the habitats, indicated on the Map of the Nature System in Sardinia, defined on the basis of the current vegetation typology. The persistence of the species in the various habitats was either confirmed or negated with in site investigations and interviews. 47.5% of municipalities have place names related to the strawberry tree. Of the 248 phyto-toponyms, 127 fall in the habitats where the species currently persists proving a correspondence between their regional)
Background: HIV-associated vulnerabilities—especially those linked to psychological issues—and limited mental health–treatment resources have the potential to adversely affect the health statuses of individuals. The concept of resilience has been introduced in the literature to shift the emphasis from vulnerability to protective factors. Resilience, however, is an evolving construct and is measured in various ways, though rarely among underserved, minority populations. Herein, we present the preliminary psychometric properties of a sample of HIV-seropositive Puerto Rican women, measured using a newly developed health-related resilience scale. Methods and design: The Resilience Scales for Children and Adolescents, an instrument with solid test construction properties, acted as a model in the development (in both English and Spanish) of the HRRS, providing the same dimensions and most of the same subscales. The present sample was nested within the Hispanic-Latino longitudinal cohort o)
Sentence reading involves multiple linguistic operations including processing of lexical and compositional semantics, and determining structural and grammatical relationships among words. Previous studies on Indo-European languages have associated left anterior temporal lobe (aTL) and left interior frontal gyrus (IFG) with reading sentences compared to reading unstructured word lists. To examine whether these brain regions are also involved in reading a typologically distinct language with limited morphosyntax and lack of agreement between sentential arguments, an FMRI study was conducted to compare passive reading of Chinese sentences, unstructured word lists and disconnected character lists that are created by only changing the order of an identical set of characters. Similar to previous findings from other languages, stronger activation was found in mainly left-lateralized anterior temporal regions (including aTL) for reading sentences compared to unstructured word and character li)
The Khon Mueang represent the major group of people present in today’s northern Thailand. While linguistic and genetic data seem to support a shared ancestry between Khon Mueang and other Tai-Kadai speaking people, the possibility of an admixed origin with contribution from local Mon-Khmer population could not be ruled out. Previous studies conducted on northern Thai people did not provide a definitive answer and, in addition, have largely overlooked the distribution of paternal lineages in the area. In this work we aim to provide a comprehensive analysis of Y paternal lineages in northern Thailand and to explicitly model the origin of the Khon Mueang population. We obtained and analysed new Y chromosomal haplogroup data from more than 500 northern Thai individuals including Khon Mueang, Mon-Khmer and Tai-Kadai. We also explicitly simulated different demographic scenarios, developed to explain the Khon Mueang origin, employing an ABC simulation framework on both mitochondrial and Y mi)
Background: Disclosing medical errors is considered necessary by patients, ethicists, and health care professionals. Literature insists on the framing of this disclosure and describes the apology as appropriate and necessary. However, this policy seems difficult to put into practice. Few works have explored the function and meaning of the apology. Objective: The aim of this study was to explore the role ascribed to apology in communication between healthcare professionals and patients when disclosing a medical error, and to discuss these findings using a linguistic and philosophical perspective. Methods: Qualitative exploratory study, based on face-to-face semi-structured interviews, with seven physicians in a neonatal unit in France. Discourse analysis. Results: Four themes emerged. Difference between apology in everyday life and in the medical encounter; place of the apology in the process of disclosure together with explanations, regrets, empathy and ways to avoid repeating the)
Social skills training, performed by human trainers, is a well-established method for obtaining appropriate skills in social interaction. Previous work automated the process of social skills training by developing a dialogue system that teaches social communication skills through interaction with a computer avatar. Even though previous work that simulated social skills training only considered acoustic and linguistic information, human social skills trainers take into account visual and other non-verbal features. In this paper, we create and evaluate a social skills training system that closes this gap by considering the audiovisual features of the smiling ratio and the head pose (yaw and pitch). In addition, the previous system was only tested with graduate students; in this paper, we applied our system to children or young adults with autism spectrum disorders. For our experimental evaluation, we recruited 18 members from the general population and 10 people with autism spectrum dis)
Abstract: In this study, we compare statistical properties of ancient and modern Chinese within the framework of weighted complex networks. We examine two language networks based on different Chinese versions of the Records of the Grand Historian. The comparative results show that Zipf’s law holds and that both networks are scale-free and disassortative. The interactivity and connectivity of the two networks lead us to expect that the modern Chinese text would have more phrases than the ancient Chinese one. Furthermore, by considering some of the topological and weighted quantities, we find that expressions in ancient Chinese are briefer than in modern Chinese. These observations indicate that the two languages might have different linguistic mechanisms and combinatorial natures, which we attribute to the stylistic differences and evolution of written Chinese. [ABSTRACT FROM AUTHOR], Copyright of PLoS ONE is the property of Public Library of Science and its content may not be copied )
Despite popular media portraying hoarding to be a problem of extremely poor housekeeping, most hoarded homes are relatively clean – large amounts of stuff just prevent the home from being functional. Some hoarded homes, however, develop poor living conditions like filth or disrepair. To date, little is known about how homes end up this way. The current study identified unique predictors and generated ideas about complex processes involved in the development of poor living conditions in hoarding. Three community agencies shared in-home assessment data for mainly involuntary clients with problematic living conditions, such as hoarded or filthy homes. These community agencies were the Metropolitan Boston Housing Partnership (n=115) in Boston, MA, the Hoarding Action Response Team (n=137) in Vancouver, BC, and the Hamilton Gatekeepers Program (n=209) in Hamilton, ON. Each site completed in-home assessments from 2010-2014 to evaluate client characteristics (lack of insight, social isolation, state of mind) and conditions of the home (number of pets, clutter accumulation, unusable bathrooms or kitchens) using the HOMES: Multidisciplinary Hoarding Risk Assessment, the Clutter Image Rating Scale, or a similar measure. Site-specific regression analyses identified unique predictors of poor living conditions. Clients with high clutter accumulation were at increased risk for squalor at all three sites, while kitchen or bathroom problems uniquely predicted squalor at two sites. Within two agencies, number of pets was also a consistent predictor of one indicator of squalor, the presence of urine or feces. Few clients had household disrepair (9-12% within sites), but findings hint that disrepair is associated with high clutter accumulation. Findings related to poor insight being a predictor of squalor were mixed. This is the first study to directly examine poor living conditions in hoarding. Replicated study findings across sites suggest that common features of hoarding, such as clutter accumulation and unusable rooms, are unique predictors for squalor. Results from this study can help community agencies that deal with problematic living situations prioritize intervention goals, especially if staff believe clients are at risk for poor living conditions.
Vowel reduction is a prominent feature of American English, as well as other stress-timed languages. As a phonological process, vowel reduction neutralizes multiple vowel quality contrasts in unstressed syllables. For bilinguals whose native language is not characterized by large spectral and durational differences between tonic and atonic vowels, systematically reducing unstressed vowels to the central vowel space can be problematic. Failure to maintain this pattern of stressed-unstressed syllables in American English is one key element that contributes to a “foreign accent” in second language speakers. Reduced vowels, or “schwas,” have also been identified as particularly vulnerable to the co-articulatory effects of adjacent consonants. The current study examined the effects of adjacent sounds on the spectral and temporal qualities of schwa in word-final position. Three groups of English-speaking adults were tested: Miami-based monolingual English speakers, early Spanish-English bil)
A numeral classifier is required between a numeral and a noun in Chinese, which comes in two varieties, sortal classifer (C) and measural classifier (M), also known as ‘classifier’ and ‘measure word’, respectively. Cs categorize objects based on semantic attributes and Cs and Ms both denote quantity in terms of mathematical values. The aim of this study was to conduct a psycholinguistic experiment to examine whether participants process C/Ms based on their mathematical values with a semantic distance comparison task, where participants judged which of the two C/M phrases was semantically closer to the target C/M. Results showed that participants performed more accurately and faster for C/Ms with fixed values than the ones with variable values. These results demonstrated that mathematical values do play an important role in the processing of C/Ms. This study may thus shed light on the influence of the linguistic system of C/Ms on magnitude cognition. [ABSTRACT FROM AUTHOR], Copyright o)
Patients with Parkinson’s disease (PD) display a variety of impairments in motor and non-motor language processes; speech is decreased on motor aspects such as amplitude, prosody and speed and on linguistic aspects including grammar and fluency. Here we investigated whether verbal monitoring is impaired and what the relative contributions of the internal and external monitoring route are on verbal monitoring in patients with PD relative to controls. Furthermore, the data were used to investigate whether internal monitoring performance could be predicted by internal speech perception tasks, as perception based monitoring theories assume. Performance of 18 patients with Parkinson’s disease was measured on two cognitive performance tasks and a battery of 11 linguistic tasks, including tasks that measured performance on internal and external monitoring. Results were compared with those of 16 age-matched healthy controls. PD patients and controls generally performed similarly on the lingui)
Word recognition includes the activation of a range of syntactic and semantic knowledge that is relevant to language interpretation and reference. Here we explored whether or not the number of arguments a verb takes impinges negatively on verb processing time. In this study, three experiments compared the dynamics of spoken word recognition for verbs with different preferred argument structure. Listeners’ eye movements were recorded as they searched an array of pictures in response to hearing a verb. Results were similar in all the experiments. The time to identify the referent increased as a function of the number of arguments, above and beyond any effects of label appropriateness (and other controlled variables, such as letter, phoneme and syllable length, phonological neighborhood, oral and written lexical frequencies, imageability and rated age of acquisition). The findings indicate that the number of arguments a verb takes, influences referent identification during spoken word re)
This is the first study to examine the effect of phonetic contexts on children’s lexical tone production. Mandarin tones in disyllabic words produced by forty-four 2- to 6-year-old children and twelve mothers were low-pass filtered to eliminate lexical information. Native Mandarin-speaking adults categorized the tones based on the pitch information in the filtered stimuli. All mothers’ tones were categorized with ceiling accuracy. Counter to the findings in most previous studies on children’s tone acquisition and the prevailing assumption in models of speech development that children acquire suprasegmental features much earlier than segmental features, this study found that children as old as six years of age have not mastered the production of Mandarin tones. Children’s tones were judged with significantly lower accuracy than mothers’ productions. Tone accuracy improved, while cross subject variability in tone accuracy decreased, with age. Children’s tone accuracy was affected by the)
The Macedonian Recension of the Church Slavonic language from its gradual beginning during the 11th century, especially in the second half, reaches its full development in the 12th century, when there is a consolidation of the basic norms of the Macedonian Church Slavonic literacy. The consolidation of these norms is connected and in continuity with the Old Slavonic Glagolitic period in the work of the Ohrid literary center. The paper presents representative examples that characterize the language of the Macedonian Old Church Slavonic literacy on orthographic, phonological, morpho-syntactic and lexical level.
Teesid: Käesolev artikkel käsitleb uusklassikalist luulet ehk luulet, mis tärkab humanistliku hariduse pinnalt ja on loodud nn klassikalistes keeltes ehk vanakreeka ja ladina keeles. Artikli esimene pool toob välja paar üldist probleemi varauusaja poeetika käsitlemises nii Eestis kui mujal. Teises osas esitatakse alternatiivina mõned näited (autoriteks G. Krüger, H. Vogelmann, L. Luden, O. Hermelin ja H. Bartholin) Tartu ja Tallinna uusklassikalisest luulest värsstõlkes koos poeetika analüüsidega, avalikkusele tundmata luuletuste puhul esitatakse ka originaaltekstid. SUMMARYThis article discusses poetry in classical languages (Humanist Greek and Neo-Latin) belonging to the classical literary tradition while focusing on poetry from Tallinn and Tartu from the sixteenth and seventeenth centuries. It does not aim to present an overview of this tradition in Estonia (already an object of numerous studies), but rather to discuss some general problems connected to such studies—both in Europe and Estonia—and to show some alternative (or complementary) analyses of neo-classical poetics, together with verse translations and texts that are not easily available or are unknown to the scholars.The discussion of neo-classical poetry in Estonia finds problems in a detachment from poetics and the consequent discrepancies. Firstly, although scholarly treatises stress the value of casual poetry (forming the most eminent part of Estonian Neo-Latin and Humanist Greek poetry), the same treatises present this poetry from the viewpoint of its social background, focusing more on the authors and events than the poetic form. For example, in the Anthology of Tartu casual poetry and the corpus of Neo-Latin poetry from Tartu, texts are presented according to genre, which is defined only according to the classification of social events (epithalamia, epicedia, congratulations for rectorate, disputations, etc). Secondly, in most cases (the anthology, re-editions), this poetry is presented to readers as prose translations. As in the case of ancient Greek and Roman poetry, the established norm in Estonia is verse translation. Translating poetry into prose, therefore, signals that these works are not to be considered poetry. Thirdly, commentaries on this poetry tend to list lexical parallels with authors from classical antiquity without distinguishing actual quotations from the usage of poetic formulae while simultaneously (mostly) ignoring the impact of pagan and Christian texts from late antiquity and renaissance and humanist literature.One alternative is to present Neo-Latin and Humanist Greek poetry as verse translations and focus more on discussing poetic devices and the impact of its contemporary poetry. Therefore, the second part of this article presents five poems as translations of verse and a subsequent analysis of their poetics.The first example is from a manuscript in the Tallinn City Archives and represents the earliest collection of neo-classical poetry, containing one Latin and five Greek poems belonging to the epistolary poem genre. Its author, Gregor Krüger Mesylanus (a latinized Greek translation of the name of his birth-town Mittenwalde, near Berlin), worked as a priest in Reval after his studies in Wittenberg during the time of Ph. Melanchthon (which explains Krüger‘s chosen poetic form). The Greek cycle is regarded thematically as variations on the same subject of the author‘s longing for home and his unhappiness with the jealousy and hostility of his fellow citizens in Reval. His choice of meter is influenced by Latin poetry, the initial long elegy balanced by four shorter poems of different meters (iambic and choriambic patterns). The final poem of the Greek cycle (Enviless Moon) is presented together with a metrical translation and analysis to demonstrate how sonorous patterns orchestrate the thematic development of the poem: the author‘s wish to be like the moon, who receives its light from the brighter sun, but remains still happy and grateful to God for his own gift and ability to bring a smaller light to others.The second example analyzes the structure and poetic motives of a metrical translation of a Greek Pindaric Ode by Heinrich Vogelmann from 1633. The paper’s author also examines the European tradition of The second example analyzes the structure and poetic motives of a metrical translation of a Greek Pindaric Ode by Heinrich Vogelmann from 1633. The paper’s author also examines the European tradition of such odes (including more than sixty examples from 1548 until 2004). The third example discusses two alternative translations and additional translation possibilities of a recently discovered anagrammatic poem by Lorenz Luden. The fourth and fifth examples are congratulatory poems addressed to Andreas Borg for the publication of his disputation on civil liberty (in 1697). A Latin congratulatory poem by Olaus Hermelin is an example of politically engaged poetry, which addresses not the student but the subject of his disputation and contemporary political situation (the revolt of Estonian nobility against the Swedish king, who had recaptured donated lands, and the exile of its leader, Johann Reinhold Patkul). The Greek poem by H. Bartholin refers to the arts of Muses to demonstrate the changes in poetical representations of university studies: by the end of the seventeenth century the motives of the dancing and singing, flowery Muses is replaced with the stress of the toil in the stadium and the labyrinth of Muses.This article discusses poetry in classical languages (Humanist Greek and Neo-Latin) belonging to the classical literary tradition while focusing on poetry from Tallinn and Tartu from the sixteenth and seventeenth centuries. It does not aim to present an overview of this tradition in Estonia (already an object of numerous studies), but rather to discuss some general problems connected to such studies—both in Europe and Estonia—and to show some alternative (or complementary) analyses of neo-classical poetics, together with verse translations and texts that are not easily available or are unknown to the scholars.The discussion of neo-classical poetry in Estonia finds problems in a detachment from poetics and the consequent discrepancies. Firstly, although scholarly treatises stress the value of casual poetry (forming the most eminent part of Estonian Neo-Latin and Humanist Greek poetry), the same treatises present this poetry from the viewpoint of its social background, focusing more on the authors and events than the poetic form. For example, in the Anthology of Tartu casual poetry and the corpus of Neo-Latin poetry from Tartu, texts are presented according to genre, which is defined only according to the classification of social events (epithalamia, epicedia, congratulations for rectorate, disputations, etc). Secondly, in most cases (the anthology, re-editions), this poetry is presented to readers as prose translations. As in the case of ancient Greek and Roman poetry, the established norm in Estonia is verse translation. Translating poetry into prose, therefore, signals that these works are not to be considered poetry. Thirdly, commentaries on this poetry tend to list lexical parallels with authors from classical antiquity without distinguishing actual quotations from the usage of poetic formulae while simultaneously (mostly) ignoring the impact of pagan and Christian texts from late antiquity and renaissance and humanist literature. One alternative is to present Neo-Latin and Humanist Greek poetry as verse translations and focus more on discussing poetic devices and the impact of its contemporary poetry. Therefore, the second part of this article presents five poems as translations of verse and a subsequent analysis of their poetics. The first example is from a manuscript in the Tallinn City Archives and represents the earliest collection of neo-classical poetry, containing one Latin and five Greek poems belonging to the epistolary poem genre. Its author, Gregor Krüger Mesylanus (a latinized Greek translation of the name of his birth-town Mittenwalde, near Berlin), worked as a priest in Reval after his studies in Wittenberg during the time of Ph. Melanchthon (which explains Krüger‘s chosen poetic form). The Greek cycle is regarded thematically as variations on the same subject of the author‘s longing for home and his unhappiness with the jealousy and hostility of his fellow citizens in Reval. His choice of meter is influenced by Latin poetry, the initial long elegy balanced by four shorter poems of different meters (iambic and choriambic patterns). The final poem of the Greek cycle (Enviless Moon) is presented together with a metrical translation and analysis to demonstrate how sonorous patterns orchestrate the thematic development of the poem: the author‘s wish to be like the moon, who receives its light from the brighter sun, but remains still happy and grateful to God for his own gift and ability to bring a smaller light to others. The second example analyzes the structure and poetic motives of a metrical translation of a Greek Pindaric Ode by Heinrich Vogelmann from 1633. The paper’s author also examines the European tradition of This article discusses poetry in classical languages (Humanist Greek and Neo-Latin) belonging to the classical literary tradition while focusing on poetry from Tallinn and Tartu from the sixteenth and seventeenth centuries. It does not aim to present an overview of this tradition in Estonia (already an object of numerous studies), but rather to discuss some general problems connected to such studies—both in Europe and Estonia—and to show some alternative (or complementary) analyses of neo-classical poetics, together with verse translations and texts that are not easily available or are unknown to the scholars.The discussion of neo-classical poetry in Estonia finds problems in a detachment from poetics and the consequent discrepancies. Firstly, although scholarly
Language comprehension involves the simultaneous processing of information at the phonological, syntactic, and lexical level. We track these three distinct streams of information in the brain by using stochastic measures derived from computational language models to detect neural correlates of phoneme, part-of-speech, and word processing in an fMRI experiment. Probabilistic language models have proven to be useful tools for studying how language is processed as a sequence of symbols unfolding in time. Conditional probabilities between sequences of words are at the basis of probabilistic measures such as surprisal and perplexity which have been successfully used as predictors of several behavioural and neural correlates of sentence processing. Here we computed perplexity from sequences of words and their parts of speech, and their phonemic transcriptions. Brain activity time-locked to each word is regressed on the three model-derived measures. We observe that the brain keeps track of t)
There has been a recent boom in research relating semantic space computational models to fMRI data, in an effort to better understand how the brain represents semantic information. In the first study reported here, we expanded on a previous study to examine how different semantic space models and modeling parameters affect the abilities of these computational models to predict brain activation in a data-driven set of 500 selected voxels. The findings suggest that these computational models may contain distinct types of semantic information that relate to different brain areas in different ways. On the basis of these findings, in a second study we conducted an additional exploratory analysis of theoretically motivated brain regions in the language network. We demonstrated that data-driven computational models can be successfully integrated into theoretical frameworks to inform and test theories of semantic representation and processing. The findings from our work are discussed in light of future directions for neuroimaging and computational research.
Health organizations are increasingly using social media, such as Twitter, to disseminate health messages to target audiences. Determining the extent to which the target audience (e.g., age groups) was reached is critical to evaluating the impact of social media education campaigns. The main objective of this study was to examine the separate and joint predictive validity of linguistic and metadata features in predicting the age of Twitter users. We created a labeled dataset of Twitter users across different age groups (youth, young adults, adults) by collecting publicly available birthday announcement tweets using the Twitter Search application programming interface. We manually reviewed results and, for each age-labeled handle, collected the 200 most recent publicly available tweets and user handles’ metadata. The labeled data were split into training and test datasets. We created separate models to examine the predictive validity of language features only, metadata features only, l)
Human language is composed of sequences of reusable elements. The origins of the sequential structure of language is a hotly debated topic in evolutionary linguistics. In this paper, we show that sets of sequences with language-like statistical properties can emerge from a process of cultural evolution under pressure from chunk-based memory constraints. We employ a novel experimental task that is non-linguistic and non-communicative in nature, in which participants are trained on and later asked to recall a set of sequences one-by-one. Recalled sequences from one participant become training data for the next participant. In this way, we simulate cultural evolution in the laboratory. Our results show a cumulative increase in structure, and by comparing this structure to data from existing linguistic corpora, we demonstrate a close parallel between the sets of sequences that emerge in our experiment and those seen in natural language. [ABSTRACT FROM AUTHOR], Copyright of PLoS ONE is the)
The amount of data from languages spoken all over the world is rapidly increasing. Traditional manual methods in historical linguistics need to face the challenges brought by this influx of data. Automatic approaches to word comparison could provide invaluable help to pre-analyze data which can be later enhanced by experts. In this way, computational approaches can take care of the repetitive and schematic tasks leaving experts to concentrate on answering interesting questions. Here we test the potential of automatic methods to detect etymologically related words (cognates) in cross-linguistic data. Using a newly compiled database of expert cognate judgments across five different language families, we compare how well different automatic approaches distinguish related from unrelated words. Our results show that automatic methods can identify cognates with a very high degree of accuracy, reaching 89% for the best-performing method Infomap. We identify the specific strengths and weaknes)
Abstract English datives show two syntactic patterns, the double object dative (DOD) and the prepositional dative (PD). The alternation between DOD and PD is influenced by three contextual factors: lexical verbs, syntactic weights, and information structures. However, it has been observed that English dative alternation by second language (L2) learners significantly deviates from the native norm. Accordingly, this study examines whether the three factors are influential when L2 learners produce dative sentences, by analyzing a learner corpus and a native speaker corpus. Results show that the learners produced PD significantly more frequently than the native speakers did. Even when DOD should be contextually preferred, the learners produced many PD sentences. These results suggest that L2 learners have trouble noticing the contextual factors when structuring English datives. The finding is further discussed as it relates to the major tenets of L2 acquisition such as cross-linguistic transfer, constructional knowledge, and language processing.
We investigated categorical perception of rising and falling pitch contours by tonal and non-tonal listeners. Specifically, we determined minimum durations needed to perceive both contours and compared to those of production, how stimuli duration affects their perception, whether there is an intrinsic F0 effect, and how first language background, duration, directions of pitch and vowel quality interact with each other. Continua of fundamental frequency on different vowels with 9 duration values were created for identification and discrimination tasks. Less time is generally needed to effectively perceive a pitch direction than to produce it. Overall, tonal listeners’ perception is more categorical than non-tonal listeners. Stimuli duration plays a critical role for both groups, but tonal listeners showed a stronger duration effect, and may benefit more from the extra time in longer stimuli for context-coding, consistent with the multistore model of categorical perception. Within a cer)