Papers reviewed and determined not to be word norm studies. Use the flag icon to report errors or suggest re-inclusion.
18265 papers
The generalized valency patterns cover both obligatory arguments and optional adjuncts of valency carriers in authentic texts. Such a valency is defined as the number of all the dependents of the valency carrier. With a motif defined as the longest sequence with non-decreasing values, this paper chooses two news-genre dependency treebanks, one in Chinese and one in English and examines the motifs of generalized valencies in them and their lengths. They both are found to abide by the right truncated modified Zipf-Alekseev distribution. In addition, the Hyperpoisson model captures the interrelation between motif lengths and length frequencies. These research findings validate valency motifs as basic language entities and as results of a diversification process.
Related to Herman (2000) and Herman (1987=2006), the present study deals with the problem of the deletion of word-final -s as evidenced in Latin inscriptions of the Empire. By reconsidering all items of the omission of -s recorded to date in the Computerized Historical Linguistic Database of Latin Inscriptions of the Imperial Age, the morphosyntactic explanation proposed by Herman (1987=2006) as for relevant omissions will be replaced by a phonetic and phonosyntactic approach which evidences the all-time prevalence of the consonantal environment in the omission of word final -s. Accordingly, the phonosyntactically determined deletion of word final -s before a subsequent consonant existed continuously but to various degrees from the Old Latin age onward all along the history of Latin. This situation might have been inherited by the Romance languages, where different and complex morphological innovations led either to the discontinuation of the phenomenon of phonosyntactically determined deletion and the stabilization of word final -s (as in Western Romance), or to the completion of the deletion process and the complete loss of word final -s (as in Eastern Romance).
Speakers constantly learn language from the environment by sampling their linguistic input and adjusting their representations accordingly. Logically, people should attend more to the environment and adjust their behavior in accordance with it more the lower their success in the environment is. We test whether the learning of linguistic input follows this general principle in two studies: a corpus analysis of a TV game show, Jeopardy, and a laboratory task modeled after Go Fish. We show that lower (non-linguistic) success in the task modulates learning of and reliance on linguistic patterns in the environment. In Study 1, we find that poorer performance increases conformity with linguistic norms, as reflected by increased preference for frequent grammatical structures. In Study 2, which consists of a more interactive setting, poorer performance increases learning from the immediate social environment, as reflected by greater repetition of others’ grammatical structures. We propose that these results have implications for models of language production and language learning and for the propagation of language change. In particular, they suggest that linguistic changes might spread more quickly in times of crisis, or when the gap between more and less successful people is larger. The results might also suggest that innovations stem from successful individuals while their propagation would depend on relatively less successful individuals. We provide a few historical examples that are in line with the first suggested implication, namely, that the spread of linguistic changes is accelerated during difficult times, such as war time and an economic downturn.
Nous presentons de nouvelles instanciations de trois corpus arbores en constituants du francais, ou certains phenomenes syntaxiques a l’origine de dependances a longue distance sont representes directement a l’aide de constituants discontinus. Les arbres obtenus relevent de formalismes grammaticaux legerement sensibles au contexte (LCFRS). Nous montrons ensuite qu’il est possible d’analyser automatiquement de telles structures de maniere efficace a condition de s’appuyer sur une methode d’inference approximative. Pour cela, nous presentons un analyseur syntaxique par transitions, qui realise egalement l’analyse morphologique et l’etiquetage fonctionnel des mots de la phrase. Enfin, nos experiences montrent que la rarete des phenomenes concernes dans les donnees francaises pose des difficultes pour l’apprentissage et l’evaluation des structures discontinues.
Universal Dependencies (UD) is becoming a standard annotation scheme crosslinguistically, but it is argued that this scheme centering on content words is harder to parse than the conventional one centering on function words. To improve the parsability of UD, we propose a backand-forth conversion algorithm, in which we preprocess the training treebank to increase parsability, and reconvert the parser outputs to follow the UD scheme as a postprocess. We show that this technique consistently improves LAS across languages even with a state-of-the-art parser, in particular on core dependency arcs such as nominal modifier. We also provide an in-depth analysis to understand why our method increases parsability. 1
Cultural differences may influence interactions between humans with different social norms and cultural traits, incurring different emotional and behavioral responses. The same applies to human-robot interaction (HRI). We believe that controlling robot emotions based on the cultural context can help robots adapt to humans from culturally diverse backgrounds. Such culturally aligned robots are expected to be easily accepted by humans as part of daily life. In this paper, we aim at investigating the role of culture in representing robot emotions which are injected by humans during its early stage of development and subject to change through their own experience thereafter. Several public data sets of pictures labeled with affective ratings by Indian, American, and European subjects are presented to social humanoid Pepper robots. The result shows that robots can learn to behave socially in alignment with an individual's cultural background. Moreover, we have demonstrated that robots under the effect of different cultures can generate different behavioral responses to the same stimuli, which is considered one of the most important issues in socially assitive robotics.
In this paper, an effective machine translation system from Thai to Khmer language on a website is proposed. To create a web application for a high performance Thai-Khmer machine translation (ThKh-MT), the principles and methods of translation involve with lexical base. Word reordering is applied by considering the previous word, the next word and subject-verb agreement. The word adjustment is also required to attain acceptable outputs. Additional steps related to structure patterns are added in a combination with the classical methods to deal with translation issues. PHP is implemented to build the application with MySQL as a tool to create lexical databases. For testing, 5,100 phrases and sentences are selected to evaluate the system. The result shows 89.25 percent of accuracy and 0.84 for F-Measure which infers to a higher efficiency than that of Google and other systems.
As few studies on Laos, there has not built relatively large dependency Treebank. Compared with the rich and mature Chinese corpus, the syntactic analysis of Laos is more difficult and still a urgently controversial issue. The existing machine learning methods need a lot of training corpus, these methods are not fully applicable to Laos and the accuracy of Laos processing results is also low. This paper presents an approach of Chinese-Laos bilingual corpus of word alignment to built Laos Dependency Treebank method. Firstly, the aligned word processing was made by Chinese-Laos sentence pairs. Secondly, the dependency parsing was done with Chinese sentences. Finally, Laos Dependency Parsing Treebank was generated by Chinese-Laos Languages align relationship and Chinese Dependency Tree. The results show that the accuracy of this method has been improved significantly compared with machine learning methods. This approach not only simplifies the process of manual collection and annotation of Laos Treebank, but saves the manpower and time in building the Treebank as well. In the case of the scarcity of Laos corpus, this approach can automatically build the Laos Dependency Treebank with high quality.
The urgency of the study is determined by the need for a comprehensive analysis of the invective in the work of I. Franko and ascertaining its linguistic status. The object of the article is the Ukrainian invective as a linguistic and speech phenomenon. The subject of the article is a characterization of the invective in I. Franko’s novel “Cross Roads”. The processing of linguistic material was conditioned by the application of such general scientifi c methods: observation — for fi xing linguistic and extralinguistic expressions of invective, descriptive — for identifi cation and identifi cation of characteristic features of Ukrainian invective, analysis and synthesis of factual material, which allowed to systematize and objectify the linguistic qualifi cation of factual material. Conclusions. The invective in the novel gives a certain color to the narrative, is one of the means of realistic depiction of everyday situations. It is a mark, emotionally and expressively colored, evaluative and expresses a negative phenomenon, sometimes allows the most complete transfer of all nuances of everyday speech. Since the aim of the invective is to make the opponent feel the whole abyss of his insignifi cance, the invective value is produced as a result of a kind of negative low creative process that occurs because of the desire of the addressee to reproduce the compatibility of words, phrases, and sentences that contradict stylistic norms. This process is not standard. Thus, the addressee fences off the realities of reality, because they are non-standard, contradictory. At the heart of the invective is a gross negative nomination, which is off ensive to the addressee. The selection of such nominations for the purpose of comparison creates an expressive imagery that contains the potential for impact on the listener due to the cynical characteristics of the object and is marked by an exquisite negative evaluation. Invective means are selected depending on the purpose and nature of the statement, from the person to whom it is addressed, since each situation requires appropriate lexical fi lling and stylistic means. With the help of appropriate language tools and techniques, it is possible to lay special information in the subconscious of a person that can become a part of his psychic essence. But it is undeniable that the word from all language means has the greatest infl uence on the addressee.
This article traces conceptual metaphors of anger in George Eliot’s ‘Middlemarch’ and argues for the importance of their role in narrative realism. It does so by showing that the figurative language of the novel is both embodied (i.e., arises in direct analogy with the bodily experience of anger) and culturally embedded. Notwithstanding the physiological and cultural conventionality of these expressions, Eliot employs a high degree of conceptual and linguistic management in her novels. This article suggests the need to broaden theories of narrative realism to take account of the network of metaphoric expressions that contribute to the form of narrative representation designated here distinctly by the term ‘embodied realism’. This discursive technique is understood to have evolved mimetically to enhance a reality effect that captures emotional recognition and authenticity, and to rely on entrenched and lexicalized mental models of anger shared by both author and reader. It is argued here that the strategy of metaphorically representing anger is most likely sub-consciously adopted yet consistently applied by Eliot. The result is a familiar sensation encoded in experientially-motivated language which prompts the reader to construe meanings that are both experientially logical and also narratively relevant. The involvement of readers in pre-existing anger schemas that are encapsulated within fairly stable metaphorical patterns and activated in the course of reading thus provides a highly reliable roadmap for interpretation, whilst also lending linguistic validity to the capability of the realist novel to elicit emotional engagement and to enforce behavioural norms.
Speakers not always maintain all grammatical norms in oral speech. However, before the Internet came into being the written public speech was predominantly based on preservation of Armenian language norms, which was largely a subject to censorship and editing. Today the situation has dramatically changed. Alongside with numerous mistakes (spelling, punctuation, lexical), a great variety of cases regarding the violation of grammatical norms are evident. The given article is dedicated to the study of these cases. Particularly, the mistakes concerning the plural, the incorrect usage of dates, pronouns and link words, cases of violation of verbal norms, as well as penetration of some common expressions of everyday language into written speech have been considered.
The Macedonian Recension of the Church Slavonic language from its gradual beginning during the 11th century, especially in the second half, reaches its full development in the 12th century, when there is a consolidation of the basic norms of the Macedonian Church Slavonic literacy. The consolidation of these norms is connected and in continuity with the Old Slavonic Glagolitic period in the work of the Ohrid literary center. The paper presents representative examples that characterize the language of the Macedonian Old Church Slavonic literacy on orthographic, phonological, morpho-syntactic and lexical level.
This study investigates the cumulative effect of (non-)native intonation, rhythm, and speech rate in utterances produced by Spanish learners of Dutch on Dutch native listeners’ perceptions. In order to assess the relative contribution of these language-specific properties to perceived accentedness and comprehensibility, speech produced by Spanish learners of Dutch was manipulated using transplantation and resynthesis techniques. Thus, eight manipulation conditions reflecting all possible combinations of L1 and L2 intonation, rhythm, and speech rate were created, resulting in 320 utterances that were rated by 50 Dutch natives on their degree of foreign accent and ease of comprehensibility.<br/>Our analyses show that all manipulations result in lower accentedness and higher comprehensibility ratings. Moreover, both measures are not affected in the same way by different combinations of prosodic features: For accentedness, Dutch listeners appear most influenced by intonation, and intonation combined with speech rate. This holds for comprehensibility ratings as well, but here the combination of all three properties, including rhythm, also significantly affects ratings by native speakers. Thus, our study reaffirms the importance of differentiating between different aspects of perception and provides insight into those features that are most likely to affect how native speakers perceive second language learners.
There has been a recent boom in research relating semantic space computational models to fMRI data, in an effort to better understand how the brain represents semantic information. In the first study reported here, we expanded on a previous study to examine how different semantic space models and modeling parameters affect the abilities of these computational models to predict brain activation in a data-driven set of 500 selected voxels. The findings suggest that these computational models may contain distinct types of semantic information that relate to different brain areas in different ways. On the basis of these findings, in a second study we conducted an additional exploratory analysis of theoretically motivated brain regions in the language network. We demonstrated that data-driven computational models can be successfully integrated into theoretical frameworks to inform and test theories of semantic representation and processing. The findings from our work are discussed in light of future directions for neuroimaging and computational research.
Emotional response to music is often represented on a two-dimensional arousal-valence space without reference to score information that may provide critical cues to explain the observed data. To bridge this gap, we present IMMA-Emo, an integrated software system for visualising emotion data aligned with music audio and score, so as to provide an intuitive way to interactively visualise and analyse music emotion data. The visual interface also allows for the comparison of multiple emotion time series. The IMMA-Emo system builds on the online interactive Multi-modal Music Analysis (IMMA) system. Two examples demonstrating the capabilities of the IMMA-Emo system are drawn from an experiment set up to collect arousal-valence ratings based on participants' perceived emotions during a live performance. Direct observation of corresponding score parts and aural input from the recording allow explanatory factors to be identified for the ratings and changes in the ratings.
We describe the first effort to annotate a signed language with syntactic dependency structure: the Swedish Sign Language portion of the Universal Dependencies treebanks.The visual modality presents some unique challenges in analysis and annotation, such as the possibility of both hands articulating separate signs simultaneously, which has implications for the concept of projectivity in dependency grammars.Our data is sourced from the Swedish Sign Language Corpus, and if used in conjunction these resources contain very richly annotated data: dependency structure and parts of speech, video recordings, signer metadata, and since the whole material is also translated into Swedish the corpus is also a parallel text.
The role of social media in giving voice to public opinion is impossible to ignore. Increasingly, platforms such as Facebook and Twitter are used for the mobilisation of action against public figures, corporations and organisations which have attracted negative public attention. For the social actors suffering this treatment, such actions may have devastating consequences.While previous studies have demonstrated that this mobilisation is in part the result of the real-time nature of social media, which allows for ‘rapid mass self-communication’ (Van der Meer & Verhoeven 2013) and the instant spreading of coherent frames across diverse groups of publics, the underlying conceptual dimension of these frames has only been studied to a limited degree, and primarily within crisis communication research (Ngai et al. 2015; Van der Meer et al. 2014). However, studying the conceptual grounding may offer additional and valuable explanations for the salience of particular frames and their ability to inspire collective action across different groups of publics. A previous, small-scale study indicates, for instance, that when commonly held notions of right and wrong are challenged, this leads to the establishment of strong and coherent frames that evoke socially and culturally embedded norms, and which not only have the purpose of condemning the culprit and his actions but also will unite publics in their call for corrective action (Author 2015). This paper reports on an explorative study that investigates the conceptual grounding of frames in instances of organisational and personal action that is deemed reproachful on social media. By examining a corpus of entries posted on Facebook in connection with two major organisational crises, the study confirms previous findings and demonstrates that the strength of frames may result from the evocation and foregrounding of basic social norms and values shared across public groups, which are otherwise considered to have different outlooks and perceptions.The theoretical foundation of the analysis is framing (Fillmore 1982; Hallahan 1999) combined with social media research (Liu 2010; Liu et al. 2011; Van der Meer & Verhoeven 2013), which provides the analyst with tools for investigating the conceptual and linguistic levels of communication on social media. Being concerned with the cognitive information processing of the receivers of text, framing can be instantiated through a number of lexical items, including metaphor (e.g. Lakoff and Johnson [1980]2003; Kövecses 2015). Due to its grounding in a bodily, situational, and discourse context as well as its richness in expression, metaphor is particularly relevant to this study and will receive special attention in the investigation of frames.
In this study we aim to analyze how a story is built in today media space. Besides the linguistics norms concerning the meaning issues, the media story is a source of catharsis which can consume psychosocial energies. The public event becomes a media event and so it becomes an aesthetic event. The contemporary soul burns the pain and the anger watching TV, doing symbolic gestures, looking for uniformity. For our case study we chose to analyze the stories around the disaster from a Romanian nightclub; during a concert, a fire started and over 60 people died. We aim to describe the lexical ritual that we identified in the media discourse. We describe the patterns that generate meaning. We show how the subjectivity and the ideology bring closer the media discourse and the fictional one, and so we see that the report on such a tragedy means more empathy and less information, more emotional release and less storage of meanings. We can speak about the haste of a collective self to impose the ritual anger as a unique direction in front of a disaster.
We present an Access database that will help to analyze and classify Portuguese MWEs according to their linguistic properties. We used a subset of an existing lexicon of Portuguese MWEs that was extracted from a corpus using Mutual Information as a statistical measure, followed by manual validation. The lexicon is organized in a three-level structure: the main lemmas, the group lemmas and the variants of those groups. This subset was imported to the database together with linguistic information about the MWEs and their variants (morphosyntactic structure, syntactic category, frequency, MI, grammatical function, discursive function, etc.). A semantic and syntactic fine-grained typology was established, and, by selecting a particular MWE, we can classify it with respect to its degree of semantic decomposability and syntactic transformation. The database is highly customizable and enables the addition/deletion of semantic or syntactic categories considered important throughout the analysis. In the end, all the information of the database will be exported to a XML format, resulting in a lexicon enriched with linguistic information.
This article gives an overview of the state of art of tools and resources for syntactic analysis of Estonian. A morphosyntactic disambiguator, surface-syntactic analyzer and dependency parser are all based on the Constraint Grammar formalism. As for language resources, a 400,000-word manually annotated dependency treebank has been created, its annotation scheme is compatible with the output of the Constraint Grammar dependency parser. Part of the treebank has been converted to the Universal Dependencies annotation scheme. Our tools have also been tested by large-scale corpus annotation.
In the spirit of socialist self-management and the aspirations to satisfy the working class’s right to inform and be informed, numerous gazettes were started by companies and local organisations in Istria during the second half of the 20th century. An important place amongst them, as a valuable source for the study of language in Istria of the socialist period, was held by the Uljanik, the magazine of the eponymous Pula shipyard, which was published in Pula from 1954 to 1990. This paper analyses the orthographical, morphological, syntactical, word-formational, lexical and stylistic features of the magazine, which are then compared to the normative rules of the handbooks and other publications with the aim to determine the extent to which the language of the <i>Uljanik</i> was in accord with the norm of the period. The analysis of all the features provides insight into the real situation of the language and orthography and evaluates the influence of the language policy and its tendencies on the language of the media in Istria in this period. A special emphasis is put on stylistic characteristics which are susceptible to the influence of the political expression in this type of texts. In addition, a content analysis of the articles has been carried out in the attempt to see how much language problems were discussed among workers.
The increased globalization of science and technology and the growing number of bilinguals and multilinguals in the world have made research with multiple languages a mainstay for scholars who study human function and especially those who focus on language, cognition, and the brain. Such research can benefit from large-scale databases and online resources that describe and measure lexical, phonological, orthographic, and semantic information. The present paper discusses currently-available resources and underscores the need for tools that enable measurements both within and across multiple languages. A general review of language databases is followed by a targeted introduction to databases of orthographic and phonological neighborhoods. A specific focus on CLEARPOND illustrates how databases can be used to assess and compare neighborhood information across languages, to develop research materials, and to provide insight into broad questions about language. As an example of how using large-scale databases can answer questions about language, a closer look at neighborhood effects on lexical access reveals that not only orthographic, but also phonological neighborhoods can influence visual lexical access both within and across languages. We conclude that capitalizing upon large-scale linguistic databases can advance, refine, and accelerate scientific discoveries about the human linguistic capacity.
In this paper, we present the results of searching for long-distance dependencies in an automatically annotated treebank for Dutch. We concentrate on phenomena that have recently been subject to debate, and where conflicting claims have been made regarding the question whether these constructions actually occur with some frequency in spontaneous language use. Long-distance dependencies involving a tensed or infinitival subordinate clause are quite rare and show collocational effects. Resumptive prolepsis and R-pronominal parasitic gaps are outside the scope of the computational grammar. We show that access to syntactic annotation even in such cases helps to find positive examples relatively quickly.
Listeners can adjust and recalibrate their phonetic boundaries based on exposure to new speech input (Norris et al., 2003). In this study, we investigate whether social factors external to the speech signal during exposure can affect this phonetic recalibration. Specifically, we test whether phonetic recalibration is modulated by the facial expression of the speaker. Existing studies show that speech production and perception are dynamically sensitive to social characteristics of the speaker (Niedzielski, 1997; Johnson et al., 1999; Babel 2012, i.a.), but it has not been studied whether perceptual learning (i.e., phonetic recalibration) is similarly sensitive to social factors. During a training phase, participants were presented auditorily with (i) 60 words with a word-medial /d/ (e.g., academia), (ii) 60 with a word-medial /t/ (e.g., politician), and (iii) 60 filler words containing neither /d/ nor /t/. An additional set of 180 non-word fillers contained neither /d/ nor /t/. The auditory material was produced by a female native speaker of American English. The task of the participants was to make a lexical decision for the 360 spoken words and non-words. Crucially, the /t/ sounds in the t-words were carefully manipulated – in particular, by shortening VOT and closure length – to be ambiguous between /t/ and /d/, and this manipulation was verified in a separate norming study. The /d/ sounds were not manipulated. During this training phase, a picture of a woman was presented on the screen. In one between-subjects condition (Smile), the woman was smiling; in the other condition (No Smile), the same woman was not smiling. After the training phase, the participants performed a categorization task for tokens on an 11-step /ata/-/ada/ continuum to assess whether their category boundary between /t/ and /d/ had shifted. Since the /t/ sounds in the training are closer to /d/ than usual, if perceptual learning occurs, the category boundary should shift towards the /d/-end of the continuum. Results from 18 female participants are shown in Figure 1. (Data collection is ongoing and the study will include a total of 32 female and 32 male participants.) Listeners in the No Smile condition showed a positive effect of perceptual learning, in that they tended to choose /t/ more often for higher continuum steps than a control group did (z = 1.9, p = 0.06), shifting the category boundary to the /d/-end. (The baseline was obtained from a separate group of female participants who did not undergo training.) Listeners in the Smile condition, on the other hand, showed no evidence for perceptual learning (z = -0.9, p = 0.4). This finding is somewhat counter to studies on learning that report better learning outcomes with more attractive or likable instructors (Westfall et al., 2016), though Babel (2012) shows that greater likeability and attractiveness can sometimes result in reduced phonetic imitation. The current study provides a novel finding that phonetic recalibration is affected by speech-external social factors, though more research is needed to understand the role of specific facial expressions.
Teesid: Käesolev artikkel käsitleb uusklassikalist luulet ehk luulet, mis tärkab humanistliku hariduse pinnalt ja on loodud nn klassikalistes keeltes ehk vanakreeka ja ladina keeles. Artikli esimene pool toob välja paar üldist probleemi varauusaja poeetika käsitlemises nii Eestis kui mujal. Teises osas esitatakse alternatiivina mõned näited (autoriteks G. Krüger, H. Vogelmann, L. Luden, O. Hermelin ja H. Bartholin) Tartu ja Tallinna uusklassikalisest luulest värsstõlkes koos poeetika analüüsidega, avalikkusele tundmata luuletuste puhul esitatakse ka originaaltekstid. SUMMARYThis article discusses poetry in classical languages (Humanist Greek and Neo-Latin) belonging to the classical literary tradition while focusing on poetry from Tallinn and Tartu from the sixteenth and seventeenth centuries. It does not aim to present an overview of this tradition in Estonia (already an object of numerous studies), but rather to discuss some general problems connected to such studies—both in Europe and Estonia—and to show some alternative (or complementary) analyses of neo-classical poetics, together with verse translations and texts that are not easily available or are unknown to the scholars.The discussion of neo-classical poetry in Estonia finds problems in a detachment from poetics and the consequent discrepancies. Firstly, although scholarly treatises stress the value of casual poetry (forming the most eminent part of Estonian Neo-Latin and Humanist Greek poetry), the same treatises present this poetry from the viewpoint of its social background, focusing more on the authors and events than the poetic form. For example, in the Anthology of Tartu casual poetry and the corpus of Neo-Latin poetry from Tartu, texts are presented according to genre, which is defined only according to the classification of social events (epithalamia, epicedia, congratulations for rectorate, disputations, etc). Secondly, in most cases (the anthology, re-editions), this poetry is presented to readers as prose translations. As in the case of ancient Greek and Roman poetry, the established norm in Estonia is verse translation. Translating poetry into prose, therefore, signals that these works are not to be considered poetry. Thirdly, commentaries on this poetry tend to list lexical parallels with authors from classical antiquity without distinguishing actual quotations from the usage of poetic formulae while simultaneously (mostly) ignoring the impact of pagan and Christian texts from late antiquity and renaissance and humanist literature.One alternative is to present Neo-Latin and Humanist Greek poetry as verse translations and focus more on discussing poetic devices and the impact of its contemporary poetry. Therefore, the second part of this article presents five poems as translations of verse and a subsequent analysis of their poetics.The first example is from a manuscript in the Tallinn City Archives and represents the earliest collection of neo-classical poetry, containing one Latin and five Greek poems belonging to the epistolary poem genre. Its author, Gregor Krüger Mesylanus (a latinized Greek translation of the name of his birth-town Mittenwalde, near Berlin), worked as a priest in Reval after his studies in Wittenberg during the time of Ph. Melanchthon (which explains Krüger‘s chosen poetic form). The Greek cycle is regarded thematically as variations on the same subject of the author‘s longing for home and his unhappiness with the jealousy and hostility of his fellow citizens in Reval. His choice of meter is influenced by Latin poetry, the initial long elegy balanced by four shorter poems of different meters (iambic and choriambic patterns). The final poem of the Greek cycle (Enviless Moon) is presented together with a metrical translation and analysis to demonstrate how sonorous patterns orchestrate the thematic development of the poem: the author‘s wish to be like the moon, who receives its light from the brighter sun, but remains still happy and grateful to God for his own gift and ability to bring a smaller light to others.The second example analyzes the structure and poetic motives of a metrical translation of a Greek Pindaric Ode by Heinrich Vogelmann from 1633. The paper’s author also examines the European tradition of The second example analyzes the structure and poetic motives of a metrical translation of a Greek Pindaric Ode by Heinrich Vogelmann from 1633. The paper’s author also examines the European tradition of such odes (including more than sixty examples from 1548 until 2004). The third example discusses two alternative translations and additional translation possibilities of a recently discovered anagrammatic poem by Lorenz Luden. The fourth and fifth examples are congratulatory poems addressed to Andreas Borg for the publication of his disputation on civil liberty (in 1697). A Latin congratulatory poem by Olaus Hermelin is an example of politically engaged poetry, which addresses not the student but the subject of his disputation and contemporary political situation (the revolt of Estonian nobility against the Swedish king, who had recaptured donated lands, and the exile of its leader, Johann Reinhold Patkul). The Greek poem by H. Bartholin refers to the arts of Muses to demonstrate the changes in poetical representations of university studies: by the end of the seventeenth century the motives of the dancing and singing, flowery Muses is replaced with the stress of the toil in the stadium and the labyrinth of Muses.This article discusses poetry in classical languages (Humanist Greek and Neo-Latin) belonging to the classical literary tradition while focusing on poetry from Tallinn and Tartu from the sixteenth and seventeenth centuries. It does not aim to present an overview of this tradition in Estonia (already an object of numerous studies), but rather to discuss some general problems connected to such studies—both in Europe and Estonia—and to show some alternative (or complementary) analyses of neo-classical poetics, together with verse translations and texts that are not easily available or are unknown to the scholars.The discussion of neo-classical poetry in Estonia finds problems in a detachment from poetics and the consequent discrepancies. Firstly, although scholarly treatises stress the value of casual poetry (forming the most eminent part of Estonian Neo-Latin and Humanist Greek poetry), the same treatises present this poetry from the viewpoint of its social background, focusing more on the authors and events than the poetic form. For example, in the Anthology of Tartu casual poetry and the corpus of Neo-Latin poetry from Tartu, texts are presented according to genre, which is defined only according to the classification of social events (epithalamia, epicedia, congratulations for rectorate, disputations, etc). Secondly, in most cases (the anthology, re-editions), this poetry is presented to readers as prose translations. As in the case of ancient Greek and Roman poetry, the established norm in Estonia is verse translation. Translating poetry into prose, therefore, signals that these works are not to be considered poetry. Thirdly, commentaries on this poetry tend to list lexical parallels with authors from classical antiquity without distinguishing actual quotations from the usage of poetic formulae while simultaneously (mostly) ignoring the impact of pagan and Christian texts from late antiquity and renaissance and humanist literature. One alternative is to present Neo-Latin and Humanist Greek poetry as verse translations and focus more on discussing poetic devices and the impact of its contemporary poetry. Therefore, the second part of this article presents five poems as translations of verse and a subsequent analysis of their poetics. The first example is from a manuscript in the Tallinn City Archives and represents the earliest collection of neo-classical poetry, containing one Latin and five Greek poems belonging to the epistolary poem genre. Its author, Gregor Krüger Mesylanus (a latinized Greek translation of the name of his birth-town Mittenwalde, near Berlin), worked as a priest in Reval after his studies in Wittenberg during the time of Ph. Melanchthon (which explains Krüger‘s chosen poetic form). The Greek cycle is regarded thematically as variations on the same subject of the author‘s longing for home and his unhappiness with the jealousy and hostility of his fellow citizens in Reval. His choice of meter is influenced by Latin poetry, the initial long elegy balanced by four shorter poems of different meters (iambic and choriambic patterns). The final poem of the Greek cycle (Enviless Moon) is presented together with a metrical translation and analysis to demonstrate how sonorous patterns orchestrate the thematic development of the poem: the author‘s wish to be like the moon, who receives its light from the brighter sun, but remains still happy and grateful to God for his own gift and ability to bring a smaller light to others. The second example analyzes the structure and poetic motives of a metrical translation of a Greek Pindaric Ode by Heinrich Vogelmann from 1633. The paper’s author also examines the European tradition of This article discusses poetry in classical languages (Humanist Greek and Neo-Latin) belonging to the classical literary tradition while focusing on poetry from Tallinn and Tartu from the sixteenth and seventeenth centuries. It does not aim to present an overview of this tradition in Estonia (already an object of numerous studies), but rather to discuss some general problems connected to such studies—both in Europe and Estonia—and to show some alternative (or complementary) analyses of neo-classical poetics, together with verse translations and texts that are not easily available or are unknown to the scholars.The discussion of neo-classical poetry in Estonia finds problems in a detachment from poetics and the consequent discrepancies. Firstly, although scholarly
This paper describes a Romanian Dependency Treebank, built at the Al. I. Cuza University (UAIC), and a special OCR techniques used to build it. The corpus has rich morphological and syntactic annotation. There are few annotated representative corpora in Romanian, and the existent ones are mainly focused on the contemporary Romanian standard. The corpus described below is focused on the nonstandard aspects of the language, the Regional and the Old Romanian. Having the intention to participate at the PROIEL project, which aligns oldest New Testaments, we annotate the first printed Romanian New Testament (Alba Iulia, 1648). We began by applying the UAIC tools for the morphological and syntactic processing of Contemporary Romanian over the books first quarter (second edition). By carefully manually correcting the result of the automated annotation (having a modest accuracy) we obtained a sub-corpus for the training of tools for the Old Romanian processing. But the first edition of the New Testament is written in Cyrillic letters. The existence of books printed in the Old Cyrillic alphabet is a common problem for Romania and The Republic of Moldova, countries where the Romanian is spoken; a problem to solve by the joint efforts of the NLP researchers in the two countries.
Background Previous studies have shown that cultural context has an influence on emotion and cognition. In this study the emotional response to international affective picture system (IAPS) was compared between Iranians and normative ratings of Americans young adults. Method One hundred and thirty eight Iranian university students (85 women, 48 men) age 18 to 52 (average= 31, SD = 7.76) enrolled in the study. Participants’ emotional response to IAPS images were rated in three dimensions (valence, arousal, dominance) using self-assessment Manikin (SAM) system. Then, valence, arousal, dominance scores were compared to those of 100 American undergraduates (50 females, 50 males) of the same age group, enrolled at Florida university and surveyed by Prof. PJ Lang in 2008. Result Our results indicate that there is complete correlation between the mean ratings of valence, arousal and dominance between Iranian and American participants. Also the results showed similarities in valence ratings, but arousal ratings especially in female participants were different. The relationship between arousal and valence showed a similar boomerang shaped distribution seen with the North American sample. Iranian sample showed positively offset and negative bias comparable to the American counterparts. Conclusion The results are promising in the sense that IAPS images can be used in studies within Iranian cultural context. However, arousal values require a modification for their proper application in Iranian cultural context. Disclosure of interest The author has not supplied his/her declaration of competing interest.
Previous works proposed annotation projection in parallel corpora to inexpensively generate treebanks or propbanks for new languages. In this approach, linguistic annotation is automatically transferred from a resource-rich source language (SL) to translations in a target language (TL). However, annotation projection may be adversely affected by translational divergences between specific language pairs. For this reason, previous work often required careful qualitative analysis of projectability of specific annotation in order to define strategies to address quality and coverage issues. In this demonstration, we present THE PROJECTOR, an interactive GUI designed to assist researchers in such analysis: it allows users to execute and visually inspect annotation projection in a range of different settings. We give an overview of the GUI, discuss use cases and illustrate how the tool can facilitate discussions with the research community.
We propose new computational models for analyzing self-reported emotional diary texts of pregnant women to support maternal care. We gathered affective ratings outside clinical setting and developed new models to facilitate interpretation and communication of affective expressions between persons representing different affective ratings. Relying on constructed emotion theory, models of dimensional emotion categories and affective ratings of Self Assessment Manikin, we demonstrate our new proposal to analyze linguistic data with computational models exploiting vector space and clustering methods. 35 persons having Finnish as a native language provided affective ratings for 195 emotional adjectives and 16 pregnancy-related nouns in Finnish in dimensions of pleasure, arousal and dominance. We developed new models to represent dependencies and differences of affective ratings between various population subgroup categorizations, including "women without children", "women with children" and "men without children" that we consider important population segments to be addressed in maternal care. Our affective ratings showed significant correlations between pleasure and dominance (like Warriner et al., 2013) and with previous data collections (Söderholm et al., 2013; Eilola & Havelka, 2010; Warriner et al., 2013). Our affective ratings had significant effects on categorizations based on gender, gender-parental role and the time of the day and duration of giving ratings. Our results indicate accordance with significant affectivity differences of gender and age (Warriner et al., 2013) and motherhood (Rosebrock et al., 2015). Our proposed models aim to support health-related communication. Our results suggest gathering next the affective ratings of patients of maternal care in a real clinical setting.
This paper presents a frame annotation scheme for Danish nouns, with VerbNet-derived frames and semantic roles covering both frame arguments and satellites. The scheme was implemented as a new module for a Danish frame tagger and applied to a 90, 000-token Danish treebank with ongoing manual revision. In addition to explicit frames, Constraint Grammar rules are used to map free semantic roles on noun dependents without pre-defined frames, using general syntactic-semantic context clues. We discuss the annotation scheme and present a statistical breakdown and linguistic evaluation of the assigned noun frames and adnominal roles in the corpus.
Omorfi is free and open source project containing various tools and data for handling Finnish texts in a linguistically motivated manner. The main components of this repository are: 1) a lexical database containing hundreds of thousands of words (c.f. lexical statistics), 2) a collection of scripts to convert lexical database into formats used by upstream NLP tools (c.f. lexical processing), 3) an autotools setup to build and install (or package, or deploy): the scripts, the database, and simple APIs / convenience processing tools, and 4) a collection of relatively simple APIs for a selection of languages and scripts to apply the NLP tools and access the database
The article begins with a presentation of a selection of electronic monolingual and bi/multilingual lexicographic resources and corpora available today to contemporary users of Slovene. The focus is on works combined with English and designed for translation purposes which provide information on the meaning of words and wider lexical units, i.e., e-dictionaries, lexical databases, web translation tools and various corpora. In a separate sub-section the most common translation technologies are presented, together with an evaluation of their role in the modern translation process. Sections 2 and 3 provide a brief outline of the changes that have affected classical dictionary planning, compilation and use in the new digital environment, as well as of the relationship between dictionaries and related resources, such as lexical databases. Some stereotypes regarding dictionary use are identified and, in conclusion, the existing corpus-based databases for the Slovenian-English pair are presented, with a view to determining priorities for the future interlingual infrastructure action plans in Slovenia.
The Prague Dependency Treebank of Spoken Czech 2.0 (PDTSC 2.0) is a corpus of spoken language, consisting of 742,316 tokens and 73,835 sentences, representing 7,324 minutes (over 120 hours) of spontaneous dialogs. The dialogs have been recorded, transcribed and edited in several interlinked layers: audio recordings, automatic and manual transcripts and manually reconstructed text. These layers were part of the first version of the corpus (PDTSC 1.0). Version 2.0 is extended by an automatic dependency parser at the analytical and by the manual annotation of “deep” syntax at the tectogrammatical layer, which contains semantic roles and relations as well as annotation of coreference.
Modifying the style of movements will be an important component of robotic interaction as more and more robots move into human-facing scenarios where humans are (consciously or unconsciously) constantly monitoring the motion profile of counterparts in order to make judgments about the state of these counterparts. This thesis includes two main contributions: (1) the development of two MATLAB tools that are designed to aid in the creation and simulation of stylized movement trajectories in varied contexts and (2) three user studies that explore the effects of environmental context on a human’s perception of stylized movement. \n \nFirst and foremost, the results from all of the user studies indicate that environmental contexts and stylized walking sequences both impact affect recognition. In the first two studies, participants were asked to categorize stimuli as one of seven affective labels. The results show that the labels were not applied consistently and so it was concluded that the affect of a multi-dimensional stimuli cannot be adequately categorized using a single affective label. In the third study the stimuli were evaluated on multiple scales and classified using ratings of valence and arousal rather than affective labels. The results were used to create a least squares model for the dataset that decomposed the affect ratings of animations to display the compound effects of stylized walking sequences and environmental contexts on affective ratings.
The article analyzes the use of innovative teaching technologies in the methodology of teaching the Ukrainian language for professional orientation in agrarian universities. The authors consider innovative forms of conducting lectures, practical classes and organization of independent work of students. Considerable attention is paid to the peculiarities of the preparation and conduct of interviews, discussions, public speaking, telephone conversations, etc., which contributes to the development of the individual style of the language of professional communication. The role of systematic work of the student from different types of lexicographic works is defined for the enrichment of lexical and phraseological resources, improvement of language literacy, increase ofculture and increase of the possibility of assimilation of new concepts. The system of communicative tasks is given in order to correct students' speech according to literary norms; enrichment of the lexical stock with speech blocks, clichés, characteristic of business communication; development of skills to select appropriate verbal and non-verbal means for registration of acts of communication taking into account the social status of a partner; Development of control and stimulating competence of organization of speech activity. It is noted that the implementation of the above communicative tasks contributes to the implementation of a set of pedagogical goals: the grammar material is worked out, students' speeches are adjusted and enriched, communicative skills are intensified and polished: to master the communication situation; plan your speech, that is, outline the meaning of the utterance; to select adequate language means for reproduction of the content; provide feedback.
It has recently been demonstrated that the reported tastes/flavours of food/beverages can be modulated by means of external visual and auditory stimuli such as typeface, shapes, and music. The present study was designed to assess the role of the emotional valence of the product-extrinsic stimuli in such crossmodal modulations of taste. Participants evaluated samples of mixed fruit juice whilst simultaneously being presented with auditory or visual stimuli having either positive or negative valence. The soundtracks had either been harmonised with consonant (positive valence) or dissonant (negative valence) musical intervals. The visual stimuli consisted of images of emotional faces from the International Affective Picture System (IAPS) with valence ratings matched to the soundtracks. Each juice sample was rated on two computer-based scales: One anchored with the words sour and sweet, while the other scale required hedonic ratings. Those participants who tasted the juice sample while presented with the positively-valenced stimuli rated the juice as tasting sweeter compared to negatively-valenced stimuli, regardless of whether the stimuli were visual or auditory. These results suggest that the emotional valence of food-extrinsic stimuli can play a role in shaping food flavour evaluation and liking.
The objective of this paper is to investigate the possibility of using key word analysis of corpus linguistics methodology in the translation quality assessment in legal genre. The conventional approach in translation quality assessment focuses on the achievement of equivalence between the source text and the target text. However, this kind of strong emphasis on preserving the letter of law often disrupts the understanding of the target reader, thus causing default in securing the same legal effect intended in the translated legal text. Against this backdrop, this study adopts more reader-oriented approach in legal translation quality assessment based on the concept of textual fit proposed mainly by Biel (2014). Focusing more on the expectancy norm of the target audience, this mode of assessment evaluates the extent to which the translation fits into the non-translated convention of the relevant sub-genre of legal language, thus lessening the cognitive processing efforts of the target reader. This paper suggests a model for assessing textual fit by examining the overused lexical, grammatical and semantic patterns of English-translated Korean statutes compared to non-translated English statutes, based on hierarchical key word analysis results provided by Wmatrix.
In the last decades, dialectometry has emerged as a new field of dialectology. As this kind of research requires large amounts of data, many dialectometric studies used data from “traditional” dialect atlases (e. g. ALF, AIS, RND) which were collected by investigating representatives of the oldest dialects available in the survey locations (i.e. the so-called NORMs, cf. Chambers & Trudgill 2004: 29). Moreover, these data contained mostly lexical and phonological (and sometimes morphological) variables, while syntactic phenomena are largely absent in traditional atlases. In this paper we would like to present results of a dialectometric study that focuses on three aspects which have not been given much attention in previous research. The first aspect concerns the research area, German-speaking Switzerland. Although it is one of the liveliest and at the same time best researched dialect areas in Central Europe, until recently (cf. Goebl et al. 2013, Scherrer & Stoeckle accepted) there have been very few dialectometric studies in this area (cf. Kelle 2001). The second aspect regards the investigated linguistic level: our analyses are based on syntax data from the Syntactic Atlas of German-speaking Switzerland (‘Syntaktischer Atlas der deutschen Schweiz', SADS; cf. Glaser & Bart 2015) which were collected between 2000 and 2002 in 383 locations German-speaking Switzerland. A special characteristic of this atlas – which leads to the third aspect we will focus on – lies in the large number of informants and their varying socio-demographic backgrounds. Whereas in traditional atlas projects, generally one or two representatives were interviewed at each survey location, in the SADS a total of almost 3200 informants participated in the survey (i. e. on average about 8 speakers per location). This gives us not only the possibility to work with frequency instead of binary data for each location, but more importantly, this setting allows us to include socio-demographic variables into our analyses. In other geographic and sociolinguistic contexts, extralinguistic variables other than geography turned out to be important explanatory factors for dialect variation (cf. Hansen-Morath 2016, Hansen-Morath & Stoeckle 2014). As for German-speaking Switzerland, various studies focusing on single phenomena from the SADS revealed high correlations between syntactic and socio-demographic variation (cf. Stoeckle accepted, Friedli 2012, Richner-Steiner 2011). However, it is still unclear whether this correlation can be observed for aggregated data and what role socio-demographic variables play in explaining syntactic variation. In order to answer these questions, we will pursue a twofold approach. On the one hand, we will create different subsets with respect to socio-demographic variables and perform dialectometric analyses for each of these subsets. A comparison of the results will help to answer the question whether a change in the geographic dialect structuring can be observed. On the other hand, we will perform regression analyses in order to determine the importance of different extralinguistic factors in explaining linguistic variation. Finally, the results will have to be interpreted in the light of the specific Swiss-German diaglossic situation, where (contrary to many other contexts) change toward both dialectal and standard structures can be observed.
The present study examined the significance of viewing images of neutral faces versus images of neutral objects on zygomatic muscle activity using facial EMG. Participants (60% women) from a pool of introductory psychology courses had their facial EMG recordings measured in response to images of neutral faces and neutral objects. Participantsâ valence rating of each image was also recorded using the Self-Assessment Manikin (SAM) in order to rate their emotional response to each image. The primary hypothesis was that participants would have greater activity in the zygomatic muscle region when presented with images of neutral faces as opposed to lessor activity when presented with images of neutral objects. It was also hypothesized that if participants preferred seeing images of faces as compared to objects, their positive feelings would produce higher SAM ratings. Results from the present study indicated images of neutral faces showed no significant difference in EMG activity compared to images of neutral objects. Self-report data also showed no significant difference in pleasantness or emotional valence between ratings of neutral faces and ratings of neutral objects.
In next years, it is necessary to draft and adopt thousands of standards and other normative documents identical to the European ones. This makes it important to formulate and adopt clear and unambiguous rules for drafting these documents. These rules should fully meet the norms of the modern Ukrainian language. One of the problems is related to the usage rules for verbs with affixes -sia (hereinafter referred to as the sia-verbs), which represent about a third (33%) of the total number of Ukrainian verbs. The essence of the problem with these verbs is that under the influence of the Russian language sia-verbs are widely used in passive constructions, which, according to leading Ukrainian linguists, don’t meet the norms of the modern Ukrainian language. The problem with these verbs is that under the influence of the Russian language sia-verbs are still widely used in passive constructions, which, according to leading Ukrainian linguists, don’t meet the norms of the modern Ukrainian language. The purpose of this article is to suggest consistent terms and definitions of basic concepts, which are needed to draft these rules, and clear criteria that would allow clearly distinguish inherent Ukrainian constructions from intruded ones. In the article, the terms for denoting verbs with affixes –sia are analysed and the advantages of the term “sia-verb” over other terms are shown. The confusion behind the usage of the terms “process” and “action”, which are very important for the formulation of rules, is investigated. It is suggested to use the term “process” as a generic term denoting the categorical meaning of the verb as parts of speech, regardless of the specific lexical meanings of an individual verb, and to use the term “action” as specific term denoting the kind of process, which is generated and directly stimulated by a logical subject. It is noted that using these terms for denotation of other concepts is inappropriate, because it can lead to confusion. The difference is shown between the transitivity/intransitivity of a process as a semantic concept and the transitivity/intransitivity of verbs that name these processes. In semantics, the criterion of process transitivity is the direction of the process and its extension to a logical object other that the logical subject. Classification of verbs by transitivity is solely based on a formally morphological criterion associated with a grammatical object, which may or may not be required by the verb used in a certain meaning. Examples are given, which demonstrate that the semantic and grammatical approaches to transitivity do not always match. It is shown that for sia-verbs, the main and primary meaning is the reflexive one (broadly speaking, this is the meaning of an intransitive process, which is focused, looped within the realm of the logical subject that, at the same time, can be the logical object). There have been selected nine sub-meanings of the reflexive meaning, that convey different shades of reflexivety – from processes focused on the logical subject to the processes having a very wide general relation to it, including ones that convey permanent and defining intransitive possessive abilities (properties). The names for these sub-meanings present in the literature have been analyzed and a consistent system of Ukrainian terms is suggested for them. These terms are built based on a pattern, which, on the one hand, makes these specific concepts’ relation with the generic concept “reverse meaning” obvious thanks to the generic characteristic, and, on the other hand, explicitly shows the difference of every specific concept from other subordinate concepts via their delimiting characteristics. Five of these terms are generally accepted, one is chosen from the options available in the literature, but three more terms are suggested from the scratch to meet the requirement of being systematic. Thereby, the Ukrainian language naturally uses the sia-verb in the situations, where the speaker treats the process as intransitive one, i.e. there is no logical subject separated from the logical object. Therefore intransitivity / transitivity of processes is the criterion that makes it possible to distinguish inherent Ukrainian reflexive and impersonal constructions from intruded ones.
OBJECTIVE: There is an evolving debate about pathological affective responses in patients with anorexia nervosa (AN). We examined startle responses in different stages of AN. METHODS: We applied a startle reflex paradigm with standardized visual stimuli (International Affective Pictures System; food and body pictures) in 64 female participants (17 acute AN, 16 chronically ill AN, 15 long-term recovered AN, 16 healthy controls). We measured subjective ratings of valence and anxiety, and electromyographic startle responses. RESULTS: Participants with acute and chronic AN displayed the same subjective valence ratings to affective stimuli but showed less startle reactivity to affective pictures (F(6, 116) = 2.75, p =.02) compared with healthy control. Food pictures were rated as more unpleasant and higher anxiety provoking by currently ill AN (F(3, 59) = 3.32, p=.03). DISCUSSION: We observed diverging subjective and psychophysiological reactions in different stages of AN. Psychophysiological methods can help to attain a more comprehensive understanding of biological alterations in the long-term course of AN. Copyright © 2017 John Wiley & Sons, Ltd and Eating Disorders Association.
galegoEste traballo describe o procedemento de deseno e construcion dun corpus ingles- galego lematizado e desambiguado semanticamente con respecto aos sentidos das palabras definidas nunha base de datos lexica. Partese dun conxunto de textos en ingles xa anotados coas etiquetas correspondentes aos nomes, verbos, adxectivos e adverbios; estes textos traducense ao galego e as palabras galegas anotanse co lema e o sentido lexico. O resultado, o corpus SensoGal, representa un recurso util que calquera usuario pode consultar e reutilizar, ao tempo que facilita a presenza do idioma galego no ambito das tecnoloxias. Nas seguintes seccions presentase o proceso de elaboracion en que se identifican as dificultades atopadas nas fases de traducion e anotacion e se rexistran as decisions tomadas por se poden servir como referencia na esperable continuacion do proxecto. Tamen se detalla o sistema de consultas e se fai unha reflexion sobre os resultados obtidos e o posible traballo futuro. EnglishThis paper presents the design and elaboration of an English-Galician corpus lemmatized and semantically disambiguated with respect to the meanings of the words defined in a lexical database. To perform the task, we used a group of texts in English where nouns, verbs, adjectives and adverbs were already tagged; these texts were translated into Galician and the words tagged with their lemma and lexical meaning. The result is the corpus SensoGal: a useful resource for users and linguists that facilitates the presence of Galician language in the field of technology. In the following sections the elaboration process will be described in the phases of translation and labelling by registering the difficulties met and the decisions taken to serve as a reference in the foreseeable continuation of the project. The search system will be also explained. Finally, a reflection about the results and the future work will be done.
Abstract English datives show two syntactic patterns, the double object dative (DOD) and the prepositional dative (PD). The alternation between DOD and PD is influenced by three contextual factors: lexical verbs, syntactic weights, and information structures. However, it has been observed that English dative alternation by second language (L2) learners significantly deviates from the native norm. Accordingly, this study examines whether the three factors are influential when L2 learners produce dative sentences, by analyzing a learner corpus and a native speaker corpus. Results show that the learners produced PD significantly more frequently than the native speakers did. Even when DOD should be contextually preferred, the learners produced many PD sentences. These results suggest that L2 learners have trouble noticing the contextual factors when structuring English datives. The finding is further discussed as it relates to the major tenets of L2 acquisition such as cross-linguistic transfer, constructional knowledge, and language processing.
Multiword expressions (MWEs) make up a significant portion of the lexicon and have distinctive characteristics of non-compositionality, non-substitutability, non-modifiability. They have been widely recognized as a very problematic part of natural language processing (NLP) as the current linguistic databases often do not have enough coverage on MWEs. This paper attempts to fill a gap in research by looking at how difficult it is to retrieve and process Japanese MWEs. The research presents an overview of 360 entries obtained through automatic-retrieval (AR) and manual retrieval (MR) from the corpus. These entries are then compared across seven databases; goo dictionary, imiwa? dictionary, the JDMWE, the WWWJDIC, NINJAL, wordnet, and the N-gram count corpus to test for whether thery are MWEs. The results obtained from this study suggest that the coverage of the database used, the differences in how phrases are represented in the dictionary, complications caused by the different writing systems present in Japanese, as well as the need for human judgement, are some of the main problems in determining whether a phrase is an MWE.
Research shows that the complexity approach to phonological treatment has a stronger evidence-base than other treatment options yet implementation in clinical practice has been missing, most likely due to a lack of familiarity with this approach. This session provided a tutorial on the main sound (accuracy, implicational universals, developmental norms, stimulability, sonority sequencing principle for clusters) and word characteristics (frequency, density, age-of-acquisition, lexicality) that guide treatment planning in the complexity approach. Case studies were used to provide practice selecting sounds and words within a complexity approach for a variety of different cases. Practical issues in using this approach (i.e., how to actually teach complex sounds and words) along with clinical materials were shared to support greater implementation of this evidence-based approach in attendee’s clinical practice.
Musical metalanguage shares the multilayered structure of general language which depends on the contextually conditioned discourse levels (registers) which enable communication in functionally different situations. Register variety is most obvious on the lexical level (which often points towards the professional identity of musicians) and is thus considered to be worth further terminological survey. Multiplicity of linguistic levels can be regarded as cultural richness, but at the same time it often hinders the communication in professional and scientific contexts, which should offer the utmost compliance with the linguistic norm. The purpose of the present research is identifying some characteristic features of terminological usage among music professionals in the Republic of Croatia depending on their social and professional identity, with special respect to the terminological norm. It has been shown that the probability of use of recommended terms correlates with the examinees' professional role, teaching and/or scientific activity, regional distribution and age.
This paper reports the rationale of a lexicographic project, which aims to implement a theoretically and empirically motivated methodology for compiling a bilingual lexical database by studying 10 polysemous English verbs of motion and their prima facie equivalents in Greek. Attention is focused on the 2 template entries created to ensure comprehensive coverage of the lexical units of the verbs and maintain consistency throughout the database. Parts of the database entries are selected to exemplify (a) the systematic treatment of polysemy and phraseology, and (b) the unified way of comparing distributions of semantic concepts among different lexical units.
The UAIC-RoDia-DepTb is a balanced treebank, containing texts in non-standard language: 2,575 chats sentences, old Romanian texts (a Gospel printed in 1648, a codex of laws printed in 1818, a novel written in 1910), regional popular poetry, legal texts, Romanian and foreign fiction, quotations. The proportions are comparable; each of these types of texts is represented by subsets of at least 1,000 phrases, so that the parser can be trained on their peculiarities. The annotation of the treebank started in 2007, and it has classical tags, such as those in school grammar, with the intention of using the resource for didactic purposes. The classification of circumstantial modifiers is rich in semantic information. We present in this paper the development in progress of this resource which has been automatically annotated and entirely manually corrected. We try to add new texts, and to make it available in more formats, by keeping all the morphological and syntactic information annotated, and adding logicalsemantic information. We will describe here two conversions, from the classic syntactic format into Universal Dependencies format and into a logical-semantic layer, which will be shortly presented.
Although it is possible to observe when another person is having an emotional moment, we also derive information about the affective states of others from what they tell us they are feeling. In an effort to distill the complexity of affective experience, psychologists routinely focus on a simplified subset of subjective rating scales (i.e., dimensions) that capture considerable variability in reported affect: reported valence (i.e., how good or bad?) and reported arousal (e.g., how strong is the emotion you are feeling?). Still, existing theoretical approaches address the basic organization and measurement of these affective dimensions differently. Some approaches organize affect around the dimensions of bipolar valence and arousal (e.g., the circumplex model), whereas alternative approaches organize affect around the dimensions of unipolar positivity and unipolar negativity (e.g., the bivariate evaluative model). In this report, we (a) replicate the data structure observed when collected according to the two approaches described above, and reinterpret these data to suggest that the relationship between each pair of affective dimensions is conditional on valence ambiguity, and (b) formalize this structure with a mathematical model depicting a valence ambiguity dimension that decreases in range as arousal decreases (a triangle). This model captures variability in affective ratings better than alternative approaches, increasing variance explained from ~60% to over 90% without adding parameters.