Papers reviewed and determined not to be word norm studies. Use the flag icon to report errors or suggest re-inclusion.
18265 papers
A simple multiple imputation-based method is proposed to deal with missing data in exploratory factor analysis. Confidence intervals are obtained for the proportion of explained variance. Simulations and real data analysis are used to investigate and illustrate the use and performance of our proposal.
The syntax and semantics of human language can illuminate many individual psychological differences and important dimensions of social interaction. Accordingly, psychological and psycholinguistic research has begun incorporating sophisticated representations of semantic content to better understand the connection between word choice and psychological processes. In this work we introduce ConversAtion level Syntax SImilarity Metric (CASSIM), a novel method for calculating conversation-level syntax similarity. CASSIM estimates the syntax similarity between conversations by automatically generating syntactical representations of the sentences in conversation, estimating the structural differences between them, and calculating an optimized estimate of the conversation-level syntax similarity. After introducing and explaining this method, we report results from two method validation experiments (Study 1) and conduct a series of analyses with CASSIM to investigate syntax accommodation in social media discourse (Study 2). We run the same experiments using two well-known existing syntactic metrics, LSM and Coh-Metrix, and compare their results to CASSIM. Overall, our results indicate that CASSIM is able to reliably measure syntax similarity and to provide robust evidence of syntax accommodation within social media discourse.
This paper presents a treebank for the healthcare domain developed at ezDI. The treebank is created from a wide array of clinical health record documents across hospitals. The data has been de-identified and annotated for constituent syntactic structure. The treebank contains a total of 52053 sentences that have been sampled for subdomains as well as linguistic variations. The paper outlines the sampling process followed to ensure a better domain representation in the corpus, the annotation process and challenges, and corpus statistics. The Penn Treebank tagset and guidelines were largely followed, but there were many syntactic contexts that warranted adaptation of the guidelines. The treebank created was used to re-train the Berkeley parser and the Stanford parser. These parsers were also trained with the GENIA treebank for comparative quality assessment. Our treebank yielded great-er accuracy on both parsers. Berkeley parser performed better on our treebank with an average F1 measure of 91 across 5-folds. This was a significant jump from the out-of-the-box F1 score of 70 on Berkeley parser’s default grammar.
How to make the most of multiple heterogeneous treebanks when training a monolingual dependency parser is an open question. We start by investigating previously suggested, but little evaluated, strategies for exploiting multiple treebanks based on concatenating training sets, with or without fine-tuning. We go on to propose a new method based on treebank embeddings. We perform experiments for several languages and show that in many cases fine-tuning and treebank embeddings lead to substantial improvements over single treebanks or concatenation, with average gains of 2.0-3.5 LAS points. We argue that treebank embeddings should be preferred due to their conceptual simplicity, flexibility and extensibility.
International audience
This paper presents a sequence to sequence (seq2seq) dependency parser by directly predicting the relative position of head for each given word, which therefore results in a truly end-to-end seq2seq dependency parser for the first time. Enjoying the advantage of seq2seq modeling, we enrich a series of embedding enhancement, including firstly introduced subword and node2vec augmentation. Meanwhile, we propose a beam search decoder with tree constraint and subroot decomposition over the sequence to furthermore enhance our seq2seq parser. Our parser is evaluated on benchmark treebanks, being on par with the state-of-the-art parsers by achieving 94.11% UAS on PTB and 88.78% UAS on CTB, respectively.
Enlightenment added the issues of the language system and the reader’s perception to the debate over translation problems. The Word was no longer a Divine mystery, but it was materialized in specifi c features, which were critically penetrated by translators. The contribution of Ukrainian translators (Teofan Prokopovych, Havrylo Buzhynskyi, Symon (Petro) Kokhanovskyi, Hryhoriy Polytyka, Petro Pidhoretskyi) to the framing of the Russian Empire instead of their homeland stimulated the discussion of translation as a way to defi ne tasks and specifi c features of searching for and fi xing up the own identity of a nation. Petro Lodiy’s main translation principle was to use all the registers of his native language so as to express the content of the original. On the basis of Hryhoriy Skovoroda’s texts, it is not possible to precisely determine the features of his translation term system due to lack of contexts, although he used fi ve Latin terms designating translation. It is not entirely clear if one should understand them as the hypernym “verto” / “converto” and the hyponyms “transfero (translator)” / “exprimo” and “interpreto (interpres)”, or as a coherent paradigm of “transfero (translator)” / “exprimo” – “interpreto (interpres)” – “verto”, which can be subject to overlap the paradigm of John Dryden (1680): “metaphrase” – “paraphrase” – “imitation”. Romanticism enriched translation discussions with the subject of linguistic identity: the mentality of a nation is refl ected in its language, and the reader lives – feels, perceives, understands – according to the linguistic norms, and by them only (Hryhoriy Kvitka-Osnovyanenko, Petro Hulak-Artemovskyi, Yakiv Holovatskyi, and later Oleksandr Potebnia and Panteleimon Kulish). Thus, untranslatability was advanced to the forefront of translation theory. From the mid-19th century, translation criticism incorporated the practice of comparing texts and commenting on the results of this operation, which boosted the search for the means of interpretative justifi cation. Back at this time Ukrainian scholars (Orest Novytskyi, Mykhailo Maksymovych, Pavlo Hrabovskyi) began applying the contextual and historical / etymological methods of semantic analysis. The translators (Mykhailo Starytskyi, Borys Hrinchenko) were managing to develop the lexical meanings of the Ukrainian language for its conceptual enrichment, and their views served as criteria for defi ning a successful correspondence in Ukrainian-language translations. Keywords: translation theory, translation criticism, translation quality assessment, translatability.
This article proposes a surface-syntactic annotation scheme called SUD that is near-isomorphic to the Universal Dependencies (UD) annotation scheme while following distributional criteria for defining the dependency tree structure and the naming of the syntactic functions. Rule-based graph transformation grammars allow for a bi-directional transformation of UD into SUD. The back-and-forth transformation can serve as an error-mining tool to assure the intralanguage and inter-language coherence of the UD treebanks.
This paper proposes a state-of-the-art recurrent neural network (RNN) language model that combines probability distributions computed not only from a final RNN layer but also from middle layers. Our proposed method raises the expressive power of a language model based on the matrix factorization interpretation of language modeling introduced by Yang et al. ( The proposed method improves the current state-of-the-art language model and achieves the best score on the Penn Treebank and WikiText-2, which are the standard benchmark datasets. Moreover, we indicate our proposed method contributes to two application tasks: machine translation and headline generation.
We compare and analyze sequential, random access, and stack memory architectures for recurrent neural network language models. Our experiments on the Penn Treebank and Wikitext-2 datasets show that stack-based memory architectures consistently achieve the best performance in terms of held out perplexity. We also propose a generalization to existing continuous stack models (Joulin & Mikolov,2015; Grefenstette et al., 2015) to allow a variable number of pop operations more naturally that further improves performance. We further evaluate these language models in terms of their ability to capture non-local syntactic dependencies on a subject-verb agreement dataset (Linzen et al., 2016) and establish new state of the art results using memory augmented language models. Our results demonstrate the value of stack-structured memory for explaining the distribution of words in natural language, in line with linguistic theories claiming a context-free backbone for natural language.
Film clips are proven to be one of the most efficient techniques in emotional induction. However, there is scant literature on the effect of this procedure in older adults and, specifically, the effect of using different positive stimuli. Thus, the aim of the present study was to examine emotional differences between young and older adults and to know how a set of film clips works as mood induction procedure in older adults, especially, when trying to elicit attachment-related emotions. To this end, we use this procedure to analyze differences in subjective emotional response between young and older adults. A sample of 57 older adults and 83 young adults watched a film set previously validated in young population. Their responses were studied in an individual laboratory session to elicit 6 target emotions (disgust, fear, sadness, anger, amusement and tenderness) and neutral state. Self-reported emotional experience was measured using the Self-Assessment Manikin (SAM). Our results show that film clips are capable of evoking positive and negative emotions in older adults. Furthermore, older adults experienced more intensely negative emotions than young adults, especially in response to disgust and fear clips. They also reported higher arousal than young adults, especially in the case of sadness, anger and tenderness clips. Nevertheless, the older adults recovered more easily from the effects of the emotion induction. The young adults reported higher arousal ratings than older adults in response to amusement film clips. On the other hand, this study reflects the importance of controlling the baseline state to study the real strength of mood induction. Overall, current data suggests significant differences occur in emotional response in adult age and that film clips are an effective tool for studying positive and negative emotions in aging research.
Cognitive variation due to language and culture has been shown in a range of domains, including visual perception,emotions, theory of mind, economic strategies, decision making, and categorization. While such patterns are robust,individuals within a given culture are affected by these cultural patterns differentially. One possible cause for theseindividual differences is personality (e.g. extroversion or agreeableness). The personality traits of individuals will affecthow they interact with and adopt cultural patterns. To explore this possibility, we perform analyses on online data fromindividuals with self-identified Myers-Briggs personality types (a popularized personality measure that is widely self-reported in social media). In particular, we examine how personality type predicts the rate at which individuals adopt novellexical items and conform to the linguistic norms of their surrounding community. The results make explicit predictionsabout which individuals will be more affect by cultural and linguistic patterns.
The main aim of the PhD study A computational syntactic analysis of Setswana(AS Berg, May 2018) is the computational syntactic analysis of the Setswana simple sentence, using Lexical Functional Grammar (LFG) as framework and XLE as the associated grammar development platform. The computational grammar is tested with a hand-crafted test suite constructed with 828 test items and consists of Setswana phrases and simple sentences. The analyses of these test items are stored in the the following available formats:.SExp for trees,.lfg for functional structures and.pl for trees and functional structures in prolog. The treebank consists of 2903 trees and functional structures for the 828 phrases and sentences.
The present paper describes and illustrates the main naming strategies attested in a lexical database of 1233 Kakataibo names of plant and animals. Seven naming strategies are proposed for Kakataibo ethnobiological nomenclature: coining, morphological derivation, borrowing, ethnobiological polysemy, compounding and grammatical nominalization (the latter two being exclusively associated with lexically complex forms). Kakataibo ethnobiological terminology overally follows the general word-formation patterns available in the language, but it will be argued that some types of compounds and grammatical nominalizations found in the database are constraint to names of plants and animal. Indeed, one particular type of lexicalized grammatical nominalization seems to be cross-linguistically unusual.
We introduce a class of convolutional neural networks (CNNs) that utilize recurrent neural networks (RNNs) as convolution filters. A convolution filter is typically implemented as a linear affine transformation followed by a nonlinear function, which fails to account for language compositionality. As a result, it limits the use of high-order filters that are often warranted for natural language processing tasks. In this work, we model convolution filters with RNNs that naturally capture compositionality and long-term dependencies in language. We show that simple CNN architectures equipped with recurrent neural filters (RNFs) achieve results that are on par with the best published ones on the Stanford Sentiment Treebank and two answer sentence selection datasets. 1
Rhapsodie is a 33000-word treebank of spoken French that is annotated for syntax and prosody. It breaks down into 57 five-minute long samples produced by 89 male and female speakers. The discourse profile of each sample is captured by six variables: event structure (dialogue vs. monologue), social context (public vs. private), genre (argumentation, description, narrative, oratory, and procedural), interactivity (interactive, non-interactive, and semi-interactive), channel (broadcasting and face-to-face), and planning type (planned, semi-spontaneous, and spontaneous). The prosodic profile of each sample is captured by two sets of three variables. The first set consists of primary (i.e. structurally objective) variables, namely the mean number per second of pauses (fPauses), conversational overlaps (fOverlap), and gap fillers (fEuh). The second set is based on a model consisting of secondary variables determined a priori by the authors because they are likely to occur in certain discourse genres. They are the mean numbers per second of prosodic prominences (fProm), intonational periods (fIPE), intonation packages (fIPA). Our main research question is whether discourse types in French can be characterized and ultimately predicted by prosodic features. We also address two side questions. First, does the fact that the corpus is relatively small, heterogeneous, and not necessarily balanced affect the representativeness of our results? Second, are the secondary prosodic features representative of discourse genres? We compiled a data table that consists of 57 observations (the corpus samples) and the twelve above listed variables. We visualized the table with RhapVis, a tool we designed on purpose (http://ressources.modyco.fr/sm/RhapVis/), explored it with principal component analysis (http://ressources.modyco.fr/sm/RhapVis/PCA.html), and looked for confirmed tendencies with non-parametric one-way ANOVAs (Kruskal-Wallis H tests). Our exploration shows that argumentative and narrative sequences are prosodically marked, whereas descriptive and procedural sequences are not. A discourse genre is prosodically marked when it is characterized by a high frequency of prosodic features, namely the simultaneous occurrence of overlaps, prominences, and intonation packages. We also claim that a discourse genre is prosodically marked when it is atypical with respect to the other speech genres. This is the case with oratory speech, which is characterized by a high frequency of intonational periods and pauses and is consequently isolated from the other types. These results were partially confirmed by the ANOVAs. Focusing on primary variables, running an ANOVA on fPause showed a significant main effect of Genre (p < 0.05). Further inspection indicates that while the lowest fPause score was found in Narration (M = 0.32; SD = 0.04), the highest score was observed in Oratory (M = 0.42; SD = 0.01). For fOverlap, the main effect of Genre reached the level of significance (p < 0.001), indicating that fOverlap also varies according to Genre. The descriptive data showed that the fOverlap score was the highest for both Argumentation (M = 0.05, SD = 0.04) and Narration (M = 0.02, SD = 0.01). Conversely, no overlap was found in both Oratory and Procedural samples. References Lindqvist, Christina. Corpus transcrits de quelques journaux televises francais, Stockholm, Elanders Gotab, 2001, 289 pages Portele T, Heuft B, Widera C, Wagner P, Wolters M (2000) Perceptual Prominence In: Speech and Signals. Aspects of Speech Synthesis and Automatic Speech Recognition. Festschrift dedicated to Wolfgang Hess on his 60th birthday. Forum Phoneticum, 69. Hektor, Frankfurt a.M.: 97-116. Wagner, P. et al. (2015b), « Disentangling and connecting different perspectives on prosodic prominence », Communication a ICPL, International Conference Prominence in Language, 2015, Cologne, ICPH, 2015
The article is an initial complex study of the lexical field norm in Ancient Chinese with focus on the classical (Warring States) period. It attempts to bring together as many terms with the meaning ‘norm, standard, rule’ as possible, classify them according to their origin and conceptual background and describe them from various perspectives, including the etymological and metaphorical one. A brief comparative glimpse on the state of affairs in Ancient Greek and Latin is offered at the end of the text, and further directions of research are suggested.
The paper presents the largest Polish Dependency Bank in Universal Dependencies format -PDBUD -with 22K trees and 352K tokens. PDBUD builds on its previous version, i.e. the Polish UD treebank (PL-SZ), and contains all 8K PL-SZ trees. The PL-SZ trees are checked and possibly corrected in the current edition of PDBUD. Further 14K trees are automatically converted from a new version of Polish Dependency Bank. The PDBUD trees are expanded with the enhanced edges encoding the shared dependents and the shared governors of the coordinated conjuncts and with the semantic roles of some dependents. The conducted evaluation experiments show that PDBUD is large enough for training a high-quality graph-based dependency parser for Polish.
OBJECTIVES: Sufficient prefrontal top-down control of limbic affective areas, especially the amygdala, is essential for successful effortful emotion regulation (ER). Difficulties in effortful ER have been seen in patients with bipolar disorder (BD), which could be suggestive of a disturbed prefrontal-amygdala regulation circuit. The aim of this study was to investigate whether BD patients show abnormal effective connectivity from the prefrontal areas to the amygdala during effortful ER (reappraisal). METHODS: Forty participants (23 BD patients and 17 healthy controls [HC]) performed an ER task during functional magnetic resonance imaging. Using dynamic causal modeling, we investigated effective connectivity from the dorsolateral prefrontal cortex (DLPFC) and ventrolateral prefrontal cortex (VLPFC) to the amygdala, as well as connectivity between the DLPFC and VLPFC during reappraisal. RESULTS: Both BD patients and HC showed decreased negative affect ratings following reappraisal compared to attending negative pictures (P <.001). There were no group differences (P =.10). There was a differential modulatory effect of reappraisal on the connectivity from the DLPFC to amygdala between BD patients and HC (P =.04), with BD patients showing a weaker modulatory effect on this connectivity compared to HC. There were no other group differences. CONCLUSION: The disturbance in BD patients in effective connectivity from the DLPFC to the amygdala while reappraising is indicative of insufficient prefrontal control. This impairment should be studied further in relation to cycling frequency and polarity of switches in BD patients.
Temporal lobe epilepsy with amygdala enlargement (TLE-AE) is increasingly recognized as a distinct adult electroclinical syndrome. However, functional consequences of morphological alterations of the amygdala in TLE-AE are poorly understood. Here, two emotional stimulation designs were employed to investigate subjective emotional rating and skin conductance responses in a sample of treatment-naïve patients with suspected or confirmed autoimmune TLE-AE (n = 12) in comparison to a healthy control group (n = 16). A subgroup of patients completed follow-up measurements after treatment. As compared to healthy controls, patients with suspected or confirmed autoimmune TLE-AE showed markedly attenuated skin conductance responses and arousal ratings, especially pronounced for anxiety-inducing stimuli. The degree of right amygdala enlargement was significantly correlated with the degree of autonomic arousal attenuation. Furthermore, a decline of amygdala enlargement following prompt aggressive immunotherapy in one patient suffering from severe confirmed autoimmune TLE-AE with a very recent clinical onset was accompanied by a significant improvement of autonomic responses. Findings suggest dual impairments of autonomic and cognitive discrimination of stimulus arousal as hallmarks of emotional processing in TLE-AE. Emotional responses might, at least partially, recover after successful treatment, as implied by first single case data.
The speech of the language pathologist may serve as a linguistic model in terms of application of language norms, unambiguous use of lexis and register for the achievement of effective professional communication. The aim of the present survey was to integrate the lexical and the corpus-based approach and to compile a collection of terminological units for the purposes of teaching English as a specialized language in the domain of Logopedics.The ESP syllabus brings together a broad range of subjects such as linguistics, phonetics, and medical sciences anatomy and physiology, psychology, neurology, and speech and language pathology. To introduce the core vocabulary and the main issues in several fields and to compensate for the lack of bilingual reference materials available to students in Logopedics, the creation of a bilingual glossary is naturally justified.On the one hand are the key components of speech production such as phonation; resonance; fluency; intonation, voice, and the components of language (phonology, morphology, syntax, semantics, and pragmatics). On the other hand, the focus on doctor-patient communication, history taking, voice, mechanics of breathing, syndromes of communicative disorders, should be introduced in accordance with the academic style conventions. For the achievement of such complex the teaching and learning goals a series of ESP language practice materials were developed, incorporating the listening, reading and speaking skills. Terminological units were extracted from authentic publications, grouped thematically in bilingual glossaries and published as an educational resource on the university platform. Key lexical items were incorporated into learning tasks and specially designed exercises for practicing pronunciation, vocabulary, and extensive oral practice through audio visuals, discussions, simulation and role-play activities, students’ Power point presentations by topics and other communication-based activities that transform the classroom into an interactive place. Further to the development of foreign language fluency, the researcher/lecturer believes that writing summaries and translating short professional texts should be part of the language seminars for logopedics.The bilingual glossaries provide clear and concise definitions, sample sentences that illustrate usage, and translation equivalents in Bulgarian, while the language practice resources reinforce the relevant terms in the field.
We present LEAR (Lexical Entailment Attract-Repel), a novel post-processing method that transforms any input word vector space to emphasise the asymmetric relation of lexical entailment (LE), also known as the IS-A or hyponymy-hypernymy relation. By injecting external linguistic constraints (e.g., WordNet links) into the initial vector space, the LE specialisation procedure brings true hyponymyhypernymy pairs closer together in the transformed Euclidean space. The proposed asymmetric distance measure adjusts the norms of word vectors to reflect the actual WordNetstyle hierarchy of concepts. Simultaneously, a joint objective enforces semantic similarity using the symmetric cosine distance, yielding a vector space specialised for both lexical relations at once. LEAR specialisation achieves state-of-the-art performance in the tasks of hypernymy directionality, hypernymy detection, and graded lexical entailment, demonstrating the effectiveness and robustness of the proposed asymmetric specialisation model.
The article presents translation analysis of the texts within tourism discourse. According to the authors, the Internet is the most popular source of information and thus tourist websites are aimed at forming tourism attractiveness of a certain region as well as promoting regional branding. As illustrated by examples of multilingual hotel websites, the language component of website content is an essential factor for translation. As a result, the analysis of data shows that in many translations various errors are made, which are characterized by a violation of stylistic, lexical, grammatical, spelling and punctuation norms or rules, consequently, translated texts do not correspond to their original communicative and pragmatic function. Having studied the original examples, the authors prove that the translated text in the tourism discourse performs its main function, i.e. attracts a large number of potential customers only when a professional translator while translating generates a new text, taking into account grammatical and linguistic norms of the language of translation, as well as maintaining stylistic imagery and colour in accordance with a specific lingua-culture of a foreign recipient.
One of the relatively recent trends in learner corpora research is building and exploiting learner translator corpora. Within corpus-based translation studies (CTS) translations are approached as a special variety of the target language. They are usually represented by texts produced by professional translators and are studied as manifestations of the current translational norm. Learner translations can be seen as a more specific variant of the said variety, which is likely to deviate from the accepted translational norm. As of now, typical linguistic features of learner translations as opposed to professional ones are only tentatively described. We hypothesize that these texts should demonstrate heavier translationese features due to the lack of professional translational skills, comparatively poor source language processing competence and target language production skills. The aim of this research is to compare learner and professional Russian translations of English mass-media texts with the reference Russian corpus of non-translations to reveal lexical differences between the three. We found that learner translations consistently showed more distance from non-translations than their professional counterparts, while both learner and professional translations undoubtedly had discursive features which made them linguistically different from naturally occurring language. These findings might help define (non)professionalism in translation and shed light on correlation between the linguistic features of a given text and translation quality, as well as contribute to pedagogical approaches to translator education.
Abstract The present paper presents the findings from the analysis of the Greek corpus of European Union directives spanning the years 1999–2008 (corpus A) and the corpus of the legal instruments used to transpose them into Greek law (corpus B). The aim of the analysis is to verify the existence of a Greek Eurolect, born through translation, and to highlight the differences between this new legal variety and the corresponding Greek legal variety. The findings of the study are particularly interesting as they point to the existence of a Greek Eurolect characterised by Europeisms on a lexical level; morphosyntactic preferences which do not conform to the Greek legal language conventions and norms; an extensive use of the future tense as a result of translating English shall into Greek; and an oscillation between the use of Κatharevousa and Demotiki, that is an H-variety and an L-variety of the Greek language.
The research is based on the documents of 1735-1755 from the State Archive of the Volgograd Region and is aimed at revealing pragmatic features of alterations and corrections that are preserved in the drafts of administrative correspondence and texts. Linguistic interpretation of reasons for corrections and selection of a certain variant of the utterance for the written text drafts resulted in distinguishing three types of corrections: factual, stylistic, and communicatively pragmatic. Factual corrections touch upon the content of the document and are explained by necessity to depict the past, present or future events in accordance with the relevance of situation. That is determined by exactness as a major feature of the document. These corrections appear to be text cut-ins of various length which help to clarify or confirm information presented in the document. Stylistic corrections are aimed at improving the style of information delivery and language performance at the lexical, grammatical, textual levels of the document. They are alterations that help to bring the text into compliance with the norms of officialese, exactness, logical and textual coherence; the corrections are provided by lexical insertions or alterations, removing dialectal and common words, restoring direct word order, changing verbal forms, adding discourse units that are to explicate logical relations between syntactic parts when the transformation of oral speech into written is required. Stylistic improvements are presented as the ones that demonstrate intentions of the writer to develop varieties in the genre. Pragmatic alterations communicate the intention to orient the text towards its addressee, to follow the norms of speech etiquette, to simplify the content and its perception, besides, in case of word order correction due to etiquette norms, pruning complicated speech units, adding etiquette clichés and emotionally colored phrases, they contribute to the impact the text is supposed to produce. In conclusion it is stated that reconstruction of the human language history may be viewed through reconstruction of mental-and-speech activity of officialese writers, in particular by discovering such features as language and professional competences, choice of language style, knowledge on correctness and speech norms, that all together reflect the direction in the development of the Russian literary language in the 18 th century.
The primary aim of this paper is to present the main lexical, stylistic, morphological and syntactic characteristics of the language used in football match reports of the media in Spanish. Due to the fact that football is the sport with the most followers around the world, in recent years we have witnessed the increase of consumption of relevant texts, especially with the emergence of the specialized media on the Internet.The paper first addresses the issue of defining the genre of la crónica futbolística and its formal aspects, and then we proceed with the presentation of the most prominent linguistic features. Being a specialized genre, it is based on a specific terminology, so we further present the main characteristics of some of the terms from the morphological, lexical and phraseological points of view, based on recent corpus research. Finally, some relevant syntactic phenomena of the genre are mentioned, with emphasis on the most frequent deviations from the standard norm, as it’s demonstrated in the literature.Additionally, the secondary aim of this article is to present the findings of some exhaustive new studies that address different aspects of football slang, as we are convinced that they can be very useful as a basis for future empirical research.In conclusion, it is evident that the texts on football in general and especially the field of football slang are extremely suitable for any type of linguistic research. The prolific production of media texts helps establishing a broader authentic corpus, and thus facilitates any study of the real language use in this specific domain.Key words: specific language, style, terminology, match report, football.
Intuitively, deriving meaning from an abstract image is a uniquely human, idiosyncratic experience. Here we show that, despite having no universally recognised lexical association, abstract images spontaneously elicit specific concepts conveyed by words, with a consistency akin to that of concrete images. We presented a group of naïve participants with abstract picture-word pairs construed as 'related' or 'unrelated' according to a preliminary norming procedure conducted with different participants. Surprisingly, the naïve participants with no prior exposure to the abstract images or any hints regarding their possible meaning, displayed a reaction time priming effect for 'related' versus 'unrelated' picture-word pairs. Critically, this behavioural priming effect, and an associated decrease in N400 mean amplitude indexing semantic priming, both correlated significantly with the degree of relatedness established in the preliminary norming procedure. Given that ratings and electrophysiological measures were obtained in different groups of individuals, our results show that abstract images evoke consistent meaning across observers, as has been shown in the case of music.
Abstract The article deals with basic requirements to the translation for specific purposes, namely legal translation. The problem posed here is defining object and theoretical basis of legal translation. The question of the necessity of information search as an integral part of translation strategy has been raised. Detailed analysis revealed that the requirements of professional translators include knowledge of lexical and grammatical peculiarities of both languages in legal sphere; deep understanding of the concepts employed by specialists in particular field and the specialist terms used to express these concepts and their relationships in the source and target languages. It is recommended that evaluation of the translation may be done on the following principles: communicative pragmatic norms of translation; equivalent norms of translation; absence of contextual, cultural, functional, lexico-grammatical mistakes.
This corpus contains parallel English-Montenegrin subtitles collected in the scope of conducting a linguistic and translatological research by Petar Božović for his PhD thesis "Audiovisual Translation and Elements of Culture: A Comparative Analysis of Transfer with Reception Study in Montenegro". The data and permission to redistribute were obtained from the Radio and Television of Montenegro (http://www.rtcg.me), the public service broadcaster of Montenegro. \nThe corpus consists of English and Montenegrin subtitles of three TV series: House of Cards (686 minutes), Damages (2878 minutes), and Tudors (1999 minutes). The corpus covers 10 seasons, 110 episodes, and 5,563 minutes in terms of duration. \nSentence alignment and basic encoding were performed inside the OPUS project (http://opus.nlpl.eu/MontenegrinSubs.php), while MSD tagging, lemmatisation, and TEI conversion were performed by the CLARIN.SI infrastructure. The English texts were tagged by TreeTagger (http://www.cis.uni-muenchen.de/~schmid/tools/TreeTagger/) and the Montenegrin texts by ReLDI Tagger (https://github.com/clarinsi/reldi-tagger) using the Serbian language model. The TreeTagger (Penn Treebank) tagset was mapped to the SPOOK MSD tagset for English (https://nl.ijs.si/spook/msd/html-en/msd-en.html). \nThe corpus is available in TEI format and derived vertical format used by CQP and Manatee (Sketch Engine). The alignments in the vertical file are given separately as tables linking the alignment elements of the two languages.
L ’objectif de ce travail est d’évaluer le dédeveloppement lexical, précoce chez les enfants bilingues et d’explorer le lien possible entre la, taille du vocabulaire et les fonctions exécutives. Nous avons testé 15,bilingues français-portugais (7 de 16 mois et 8 de 24 mois). Leur, développement langagier a été évalué avec l'Inventaire du développement, communicatif français et portugais (adaptations du CDI MacArthur-Bates, Fenson et al., 2007). Des questionnaires parentaux ont été utilisés pour,évaluer la dominance linguistique (PaBiQ, Tuller, 2015), les stades de, développement (ASQ-3™, Squires et al., 2009) et les fonctions exécutives,(BRIEF-P, Gioia, Aspy, … Isquith, 2003). Nous avons calculé la taille du, vocabulaire dans chacune des langues, le vocabulaire total et le vocabulaire, conceptuel total et comparé avec les normes des monolingues. Presque, tous les participants ont un vocabulaire total dans chacune des langues,(français ou portugais) et un vocabulaire conceptuel total similaire à celui, des monolingues portugais et français. Leur vocabulaire total,(français+portugais) est par contre supérieur à celui des monolingues. Il, existe une corrélation entre la taille du vocabulaire et la mémoire de travail,(Stokes & Klee, 2009), mais aucune avec l'inhibition. Ces résultats donnent, un meilleur aperçu du processus de développement du langage bilingue.
Considering the importance of rational, effective professional language (for thinking and communication), the subject of the study, the results of which we presented in this article, is the analysis of the consistency of logic terms with the norms of Ukrainian terminological standards. This article is about the observance of linguistic norms, in particular, giving preference to Ukrainian-speaking terms against foreign-language ones (a very large percentage of foreign-language terms in the field of logic is evident) as well as the delimitation of the names of action, events and consequences by the form of the word. Having achieved these objectives would, at the same time, lead to the adoption of terms of logic and unification of formally-linguistic means during the creation of terms. The research was to do the following: 1) the discovery of those terms of logic that do not meet the requirements of the DSTU on terminology; 2) the analysis of the possible ways to achieve the correspondence between terms and terminological standards of Ukraine. As a result of the research, we have constructed the series of interrelated process logic terms: name of the action by the verb – name of the action by verbal noun – name of the completed action (event) by the verb – name of the event by the verbal noun – name of the consequence of action by the verbal noun. We constructed such rows for terms that are the names of operations for obtaining new knowledge: generalization, restriction, derivation, proof, refutation, making a conclusion, deduction, making an assumption, making of a hypothesis, implication. We used the following requirements during the construction of these series of terms: 1) all the terms of each series must be created on the same lexical basis; 2) As the terms we should use such words, in which the meaning of the word is consistent with its form, that is, if the main purpose in accordance with a particular form of the word is an action, an event, or a consequence of an event, then the meaning of a word must be the action, the event or the consequence of the event, respectively. If, for example, the main purpose of the verbal noun with suffixes ‑annia, -ennia is to denote an unfinished or completed action, then it is incorrect to use it to denote the consequence of an event. When creating a system of terms, we have established the following relationships between them: proof – is a deduction in the case, when as a conclusion we try to confirm the truthfulness of a given thesis; refutation – is a deduction in the case, when as a conclusion we try to confirm the falseness of a predefined thesis; derivation – is a deduction in the case, when the desired conclusion is not predefined; making an assumption – is the creation of allegedly true affirmations; making of a hypothesis – is the creation of probably true affirmations in the field of science. As a result of the study of the coherence of widely used logic terms with the requirements of terminological standards, we have found out that in the Ukrainian terminology system of logic there is a large number of terms that are not consistent with the terminological standards and, therefore, we proposed new linguistically correct terms (in a number of cases we proposed linguistically correct termsduplicates in order to be able to choose the most appropriate one).
Abstract In the context of the current heated debate surrounding the pervasive influence of the English language and Anglo-American culture on other languages, as well as the widespread purist attitude towards some contact-induced language change phenomena, both abroad and in Romania, our article discusses the situation of English lexical borrowings in present-day Romanian, focusing on the perception and processing of the so-called luxury Anglicisms ( Sections 2 and 3 ) by young Romanian native speakers, in an attempt to see whether such an analysis can help clarify their acceptability and diffusion across our target population. We propose an alternative cognitive, psycholinguistic approach to the study of contact-induced lexical borrowings, aiming to show that there is no difference in the young Romanian native speakers’ processing of sentences containing luxury Anglicisms and their established Romanian counterparts. Such findings may support our claim that the acceptability and diffusion of such Anglicisms are pervasive across our target population, even if the official position generally condemns such uses, considering them gratuitous and a burden in communication, even making it unintelligible sometimes. Our analysis starts from the observation that most (but not all) Romanian academics, whether linguists or not, tend to embrace a purist attitude, while on the other hand young Romanians accept such Anglicisms and tend to use them extensively. In fact, such uses are not limited to young people, who have been the subjects of our research, but are the ‘norm’ in daily conversations and elsewhere across the general population ( Stoichițoiu Ichim 2006 ). Thus, there seems to be a gap between the actual acceptability and diffusion of luxury Anglicisms among Romanians and the ‘official’ recommendations. Based on the results of a sensicality task, meant to show how 188 Romanians, aged 18–22, process and perceive sentences with or without luxury Anglicisms (see Section 6 ), we will try to show that luxury Anglicisms are accepted and, by recurrent use, diffused among the Romanian community. For a more accurate picture of their diffusion, the findings will be further correlated with data from CoRoLa, the only official corpus of present-day Romanian (beginning 1989) made available under the auspices of the Romanian Academy, as well as a corpus currently in the making, and the Internet (see Section 7 ). Besides showing that luxury Anglicisms cannot really be blamed for burdening or impairing processing, and thus communication, and explaining why such uses should not be censured or disapproved, we hope that our study of acceptability and diffusion will demonstrate that we are dealing with a complex, multi-layered phenomenon that can be better understood by going beyond a diachronic and synchronic analysis of particular words and a frequency count, and should incorporate more experimental data. Last but not least, we suggest that, on the practical side, such experimental studies as the one described here could be used as an additional criterion for the lexicographic inclusion of lexical borrowings.
In NLP data drives research, as evidenced by the frequency with which seminal works of database engineering such as the Penn Treebank have been employed as a basis for experimentation. Traditionally large-scale expertly annotated corpora are expensive and time consuming to produce. This paradigm drove researchers to adopt automated methods for generating labeled data with available tools such as Freebase, DBpedia, and the "infoboxes" found on Wikipedia pages. These knowledge bases have been, or are in the process of being, subsumed by Wikidata, an initiative to concentrate such disparate data repositories in an organized machine readable format. This resource is an important research tool. In this paper, we review our experience using Wikidata in constructing a large annotated corpus under distant supervision, moreover we make the materials, the code used to generate our annotations, freely available to all interested parties.
We present a new Arabic annotation web application that aims to associate tags to the words of the vocalized Hadiths corpus. The application assures a collaborative space between many annotators. Each annotator treats one or more Hadith from his/her personal session. We analyze each Hadith using AlKhalil morphological analyzer. The annotator rectifies the analysis and assigns the accurate annotation values. Annotation includes the segmentation of a word, the POS tag, the root and the pattern. The final annotation results of the corpus are structured in XML files. We can display, also, the annotation results in a Treebank. Our tool is considered as a crowdsourcing solution that uses contributions from various annotators to obtain an annotated Arabic coprus. This resource may be used as a standard collection to experiment NLP researches.
Abstract The first part of this paper outlines the relevant aspects of functional structuralism serving lexicographers as a departure point for building a model of lexical meaning useable in the Dictionary of Contemporary Slovak Language. This section also points to some aspects of Klára Buzássyová’s research on lexis and wordformation that have enriched the functionalstructuralist paradigm. The second section shows other theoretical and methodological frameworks, such as linguistic pragmatics, cognitive linguistics and corpus linguistics (all of them departing in some respect from the structuralism and, in other aspects, being complementary with it) that can enhance the structuralist basis of the model. The third section outlines an extended model of lexical meaning that represents a synthesis of all those theoretical frameworks and, at the same time, represents a reflection of three language constituents: 1. The social constituent is present in consideration of communicative functions of utterances, naming functions of lexical units, functional styles and registers, language norms, and situational contexts; 2. The psychological component takes the form of consideration of the prototype effect, the abolition of boundaries between linguistic meaning and other parts of cognition; 3. Thanks to the structural/systematic component, a description of paradigmatic and syntagmatic behaviour of words can be performed, and an inventory of formalcontent units and categories (lexemes, lexies, wordforming and grammatical structures) can be provided. In our dictionary practice, the abovementioned model is reflected in the methodological procedures as follows: 1. Systemization of repetitive (regular, standardized) phenomena; 2. Prototypicalization of meaning description; 3. Contextualization/encyclopedization of meaning description; 4. Pragmatization of meaning description; 5. Continualized presentation of language phenomena, i.e., introduction of numerous phenomena of transient and indeterminate nature and indicating the existence of a semanticpragmatic and lexicalgrammatical continuum; 6. “Discretization” of combinatorial continuum, i.e., identification and description of entrenched word combinations with naming functions.
Discourse analysis is necessary for different tasks of Natural Language Processing (NLP). As two of the most spoken languages in the world, discourse analysis between Spanish and Chinese is important for NLP research. This paper aims to present the first open Spanish-Chinese parallel corpus annotated with discourse information, whose theoretical framework is based on the Rhetorical Structure Theory (RST). We have evaluated and harmonized each annotation part to obtain a high annotated-quality corpus. The corpus is already available to the public.
A digital Persian text suffers from two simple but important problems. The first problem concerns multi-token units to which the individual words are attached. The other problem concerns multi-unit tokens that result from the detachment of elements of a word. This paper introduces an algorithm to reduce these problems automatically and to achieve a standard text. The proposed algorithm has three steps. In the first step, the multi-token units are split into individual words and the multi-unit tokens are then attached together. For this step, a core algorithm based on language modeling is introduced to split multi-token units into independent words. The algorithm is modified with respect to the possible challenges of improving the performance[m2]. Furthermore, this step utilizes a morphological analyzer to study derivational and inflectional affixes and exact matching in a word list to resolve the problem of the multi-token units. In the second step, an exact word matching strategy is used to resolve the multi-token unit problem of verbs. The third step repeats the algorithm in the first step to fix new problems raised by running the second step. The introduced algorithm was tested in tokenizing the data in the Persian Linguistic DataBase (PLDB). The algorithm achieved 72.04% correction of the errors in the test set with 97.8% accuracy and 0.02% error production in the spelling.
Ambivalence is a common experience that permeates a broad range of research. Unfortunately, quantifying ambivalence has proven a daunting task, with researchers limited to studying vacillating ambivalence, VA (i.e., temporal oscillations between favor/disfavor evaluations of an attitude object). Here, we demonstrate the use of the density matrix to measure both VA and what we term “simultaneous ambivalence” (SA): ambivalence that manifests itself as “in the moment” concurrent favor/disfavor evaluations. In a methodological study we gave participants the option of either single-responding or double-responding to questionnaire items regarding a controversial topic (i.e., affirmative action). Since standard statistical procedures provide no means for analyzing double responses, such data are routinely treated as “bad.” As demonstrated here, the density matrix provides an unambiguous and relatively easy means of accounting for double responses, which is our indicator of SA. Our data are well explained by a mixture model, with participants divided into two nearly equal groups of SA and non-SA participants, and provide evidence that the general phenomenon of SA transcends differences of gender and ethnicity. Further, the density matrix data are consistent with viewing SA and VA as distinct ambivalence constructs.
Coherent texts, whether written or spoken, are built upon discourse relations linking utterances together through causal, temporal or contrastive connections, among many other types (Mann & Thompson 1988). Different types of relations are signalled by different types of markers, although there is no one-to-one mapping. These markers often belong to the functional category of discourse-relational devices, or “connectives”, such as however, because or in fact. Writers and speakers also have the option to use other signalling devices (e.g. lexical or syntactic patterns) or even to leave a discourse relation implicit (e.g. Taboada 2009). This study focuses on another strategy for discourse marking, namely the use of underspecified connectives (Spooren 1997). More particularly, we investigate the role of the additive conjunction and to signal relations of addition but also of consequence, contrast and concession. In these cases, the discourse relation is more specific than the information strictly provided by the connective: a consequence or contrast is more informative than a mere additive relation. Despite the low informative value of and, it is quite often found in authentic contexts where such enriched interpretations were assigned to the discourse relation (6% of and express a result in the Penn Discourse TreeBank 2.0, Prasad et al. 2008), which calls for more research on the conditions under which and can be used as an underspecified connective. Our research objective is thus to compare the linguistic and contextual features of utterances linked by and which either express addition, consequence, contrast or concession. To do so, we first need corpus-based data where such discourse relations are reliably identified. Discourse relation annotation is extremely costly in time and human resources, it requires heavy training and, even so, agreement scores are often rather low (Spooren & Degand 2010). As a result, researchers have recently started to turn to crowdsourcing as an alternative method to gather discourse relation disambiguations through a low-cost, non-expert workforce (Kawahara et al. 2014; Rohde et al. 2016). Scholman & Demberg (2017) report on the results of a connective insertion task which they used as an indirect method to annotate discourse relations: the sense of a relation can be retrieved through the selection of unambiguous connectives from a list to fill in a blank between utterances, provided this task is repeated by a large number of participants (around 20). The authors discuss the validity of this method and conclude that crowdsourcing connective insertions is reliable enough as an alternative to expert annotations. For connective insertion tasks, it's important to distinguish between originally implicit vs. explicit relations. For originally implicit relations, inserting a connective resembles the approach taken in PDTB annotation (except that the choice of connectives is more restricted in the crowdsourcing step in order to allow for disambiguation of relation type). In originally explicit relations, however, the meaning and interpretation may substantially change by removing the connective, see examples (1)-(3) below. In this study, where we investigate originally explicitly marked relations with the connective “and”, we will therefore compare a crowdsourced connective insertion task with a crowdsourced connective replacement task, in which the original connective “and” is not removed from the stimulus. (1) I am not going back to Germany. Therefore, I will not eat spätzle ever again. (consequence) (2) I am not going back to Germany. In fact, I will not eat spätzle ever again. (addition) (3) I am not going back to Germany. I will not eat spätzle ever again. (?cause) Our working hypothesis is that such differences in interpretations, triggered by the connective (or absence thereof), apply to utterances containing and as well. In this respect, we challenge previous experimental research on and which showed that and has a very small informative value and little or no facilitating effect on reading times (Murray 1994) or comprehension (Cain & Nash 2011). By contrast, we expect that connective insertion tasks will be affected by the presence of and, thus supporting the claim that and does trigger enriched pragmatic inferences (Blakemore & Carston 1999). We therefore propose to use crowdsourcing for the study of underspecified and by comparing stimuli with and without the original connective. In other words, we want to test whether connective elicitations will differ across stimuli which are identical except for the presence or absence of the conjunction and. To this end, we ran two crowdsourcing experiments on the Prolific Academic online platform. In the first one, we used 83 authentic pairs of utterances originally containing and. We collected them from the Loyola Corpus of Computer-Mediated Communication (Goldstein-Stewart et al. 2008), in order to avoid the high formality of existing corpora such as the Penn Discourse Treebank (economy newspaper articles). This corpus contains blogs and chat conversations between college students about topics such as gay marriage, gender discrimination or privacy rights. This data was pre-annotated by the first author as either expressing a relation of addition, contrast, concession or consequence. The stimuli are grouped in four lists of about 20 items each, which are balanced with respect to the pre-annotated relation type (about 10 addition, 6 consequence, 1 contrast, 3 concession in each list). The participants (paid 1€ per list) can choose from a list of eight connectives to fill in a blank between the two utterances (the original and has been removed). The connectives are in addition, plus, therefore, as a result, by contrast, whereas, nevertheless and yet. The second experiment uses exactly the same lists of items, except that the stimuli now show the original and connecting the utterances, and the participants are therefore instructed to substitute this and with one of the connectives from the same list of options. In the analysis, we first compare the connectives chosen by the participants with the relation type pre-identified by the expert annotator, in order to see whether they converge (e.g. therefore or as a result selected in case of a relation of consequence). This first step provides us with a dataset of utterance pairs with their disambiguated discourse relation, without resorting to costly (and partly subjective) expert annotations. The items thus classified into one of the four categories (addition, consequence, contrast, concession) will allow us to test the effect of additional variables (e.g. register) in further studies (Crible & Demberg 2017). We will replicate this analysis with the results of the second experiment. We will then compare whether the connectives selected by the participants are the same when and is present in the stimuli and when it is not. Preliminary results show that, when the relation was pre-annotated as additive, the participants tend to equally choose a consequence or an additive connective, with no significant difference, which suggests that consequence is often interpreted in the absence of a connective. For all other relation types (i.e. consequence, concessive and contrast), the great majority of participants’ choices match the pre-annotation, thus confirming that and can be used in contexts which express more than mere addition. The results from the second experiment are still pending, and should lead to interesting comparisons on the effect of and in connective elicitation (or connective substitution). This study has a number of implications on the informative value of and, which may be higher than what previous studies have suggested, and on the use of connective elicitation with or without the original connective included in the stimuli. We argue that, when dealing with authentic corpus-based stimuli, including the original connective in the experiment is a more accurate representation of the data and better reproduces the interpretation mechanisms as they would be processed in natural conditions. The experiments reported in this paper constitute the first step of a larger project on the contextual and cognitive constraints to the production and interpretation of underspecified connectives. They also relate to ongoing crosslinguistic projects on the meaning variation of and and its use across spoken registers (Crible, in press) and in translation (Abuzcki et al. 2017).
Abstract An important branch of linguistics, namely, sociolinguistics, considers “languages” as normative social constructs and not as fixed communication tools characterised by an identifiable set of core features. The latter position was defended in the early sociolinguistic studies of Joshua Fishman on language decline and maintenance in the second half of the twentieth century and in the influential work on generative grammar of Noam Chomsky. In contradiction or contrast to this position, today’s sociolinguistics, which I will refer to as “mainstream sociolinguistics” in this chapter, claims that, in the default case, languages have no fixed boundaries and that they are, in fact, “fluid”. The main argument which mainstream sociolinguistics puts forward to support this claim is that the output of speech production, i.e., referred to as “languaging”, taps from linguistic resources and not from identifiable languages. This chapter argues against this theory. Firstly, multilingual phenomena, including hybrid varieties of global English and hybrid expressions in linguistic landscapes, including that of the Dutch city of Utrecht which claimed to support the languaging-approach, are, in fact, traditional cases of code-switching and code-mixing involving identifiable languages. Secondly, the languaging-approach, in contrast to the languages-approach, makes the wrong predictions. The absence of identifiable languages predicts the “flattening” of linguistic power relations. However, it will be argued that, even in a linguistically highly diverse context, a re-arrangement of power relations between languages takes place and that language hierarchies pop up. Hence, the theory which recognises individual languages makes the correct predictions. Thus, there is no reason to abandon the language concept of early sociolinguistics or Chomskyan linguistics that languages, or at least some modules of language, especially in the domain of semantics and pragmatics, are socially constructed, but that, at the same time, they are characterised by a prototypical grammatical and lexical basic core.