Papers reviewed and determined not to be word norm studies. Use the flag icon to report errors or suggest re-inclusion.
18265 papers
The semantic analysis field has a crucial role to play in the research related to text analytics. Calculating the semantic similarity between sentences is a long-standing problem in the area of natural language processing, and it differs significantly as the domain of operation differs. In this paper, we present a methodology that can be applied across multiple domains by incorporating corpora-based statistics into a standardized semantic similarity algorithm. To calculate the semantic similarity between words and sentences, the proposed method follows an edge-based approach using a lexical database. When tested on both benchmark standards and mean human similarity dataset, the methodology achieves a high correlation value for both word (r = 0.8753) and sentence similarity (r = 0.8793) concerning Rubenstein and Goodenough standard and the SICK dataset (r = 0.83241) outperforming other unsupervised models.
It has been shown that implicit connectives can be exploited to improve the performance of the models for implicit discourse relation recognition (IDRR). An important property of the implicit connectives is that they can be accurately mapped into the discourse relations conveying their functions. In this work, we explore this property in a multi-task learning framework for IDRR in which the relations and the connectives are simultaneously predicted, and the mapping is leveraged to transfer knowledge between the two prediction tasks via the embeddings of relations and connectives. We propose several techniques to enable such knowledge transfer that yield the state-of-the-art performance for IDRR on several settings of the benchmark dataset (i.e., the Penn Discourse Treebank dataset).
We present our CHARLES-SAARLAND system for the SIGMORPHON 2019 Shared Task on Crosslinguality and Context in Morphology, in task 2, Morphological Analysis and Lemmatization in Context. We leverage the multilingual BERT model and apply several fine-tuning strategies introduced by UDify demonstrating exceptional evaluation performance on morpho-syntactic tasks. Our results show that fine-tuning multilingual BERT on the concatenation of all available treebanks allows the model to learn cross-lingual information that is able to boost lemmatization and morphology tagging accuracy over fine-tuning it purely monolingually. Unlike UDify, however, we show that when paired with additional character-level and word-level LSTM layers, a second stage of fine-tuning on each treebank individually can improve evaluation even further. Out of all submissions for this shared task, our system achieves the highest average accuracy and f1 score in morphology tagging and places second in average lemmatization accuracy.
Tree-LSTMs have been used for tree-based sentiment analysis over Stanford Sentiment Treebank, which allows the sentiment signals over hierarchical phrase structures to be calculated simultaneously. However, traditional tree-LSTMs capture only the bottom-up dependencies between constituents. In this paper, we propose a tree communication model using graph convolutional neural network and graph recurrent neural network, which allows rich information exchange between phrases constituent tree. Experiments show that our model outperforms existing work on bidirectional tree-LSTMs in both accuracy and efficiency, providing more consistent predictions on phrase-level sentiments.
Animal phobias are one of the most prevalent mental disorders. We analysed how fear and disgust, two emotions involved in their onset and maintenance, are elicited by common phobic animals. In an online survey, the subjects rated 25 animal images according to elicited fear and disgust. Additionally, they completed four psychometrics, the Fear Survey Schedule II (FSS), Disgust Scale - Revised (DS-R), Snake Questionnaire (SNAQ), and Spider Questionnaire (SPQ). Based on a redundancy analysis, fear and disgust image ratings could be described by two axes, one reflecting a general negative perception of animals associated with higher FSS and DS-R scores and the second one describing a specific aversion to snakes and spiders associated with higher SNAQ and SPQ scores. The animals can be separated into five distinct clusters: (1) non-slimy invertebrates; (2) snakes; (3) mice, rats, and bats; (4) human endo- and exoparasites (intestinal helminths and louse); and (5) farm/pet animals. However, only snakes, spiders, and parasites evoke intense fear and disgust in the non-clinical population. In conclusion, rating animal images according to fear and disgust can be an alternative and reliable method to standard scales. Moreover, tendencies to overgeneralize irrational fears onto other harmless species from the same category can be used for quick animal phobia detection.
This paper suggests one way to enhance the ability of Chinese learners to analyze sentences. It is building a treebank and visualizing it as a syntactic tree(or parsed tree) and providing it to learners. The process of building a treebank, which is the most important key in this method, is divided into three parts and described in detail.
BACKGROUND: The tendency to inhibit anger (anger-in) is associated with increased pain. This relationship may be explained by the negative affectivity hypothesis (anger-in increases negative affect that increases pain). Alternatively, it may be explained by the cognitive resource hypothesis (inhibiting anger limits attentional resources for pain modulation). METHODS: A well-validated picture-viewing paradigm was used in 98 healthy, pain-free individuals who were low or high on anger-in to study the effects of anger-in on emotional modulation of pain and attentional modulation of pain. Painful electrocutaneous stimulations were delivered during and in between pictures to evoke pain and the nociceptive flexion reflex (NFR; a physiological correlate of spinal nociception). Subjective and physiological measures of valence (ratings, facial/corrugator electromyogram) and arousal (ratings, skin conductance) were used to assess reactivity to pictures and emotional inhibition in the high anger-in group. RESULTS: The high anger-in group reported less unpleasantness, showed less facial displays of negative affect in response to unpleasant pictures, and reported greater arousal to the pleasant pictures. Despite this, both groups experienced similar emotional modulation of pain/NFR. By contrast, the high anger-in group did not show attentional modulation of pain. CONCLUSIONS: These findings support the cognitive resource hypothesis and suggest that overuse of emotional inhibition in high anger-in individuals could contribute to cognitive resource deficits that in turn contribute to pain risk. Moreover, anger-in likely influenced pain processing predominantly via supraspinal (e.g., cortico-cortical) mechanisms because only pain, but not NFR, was associated with anger-in.
In the present article, a novel emotional complexity marker is proposed for classification of discrete emotions induced by affective video film clips. Principal Component Analysis (PCA) is applied to full-band specific phase space trajectory matrix (PSTM) extracted from short emotional EEG segment of 6 s, then the first principal component is used to measure the level of local neuronal complexity. As well, Phase Locking Value (PLV) between right and left hemispheres is estimated for in order to observe the superiority of local neuronal complexity estimation to regional neuro-cortical connectivity measurements in clustering nine discrete emotions (fear, anger, happiness, sadness, amusement, surprise, excitement, calmness, disgust) by using Long-Short-Term-Memory Networks as deep learning applications. In tests, two groups (healthy females and males aged between 22 and 33 years old) are classified with the accuracy levels of [Formula: see text] and [Formula: see text] through the proposed emotional complexity markers and and connectivity levels in terms of PLV in amusement. The groups are found to be statistically different ( p << 0.5) in amusement with respect to both metrics, even if gender difference does not lead to different neuro-cortical functions in any of the other discrete emotional states. The high deep learning classification accuracy of [Formula: see text] is commonly obtained for discrimination of positive emotions from negative emotions through the proposed new complexity markers. Besides, considerable useful classification performance is obtained in discriminating mixed emotions from each other through full-band connectivity features. The results reveal that emotion formation is mostly influenced by individual experiences rather than gender. In detail, local neuronal complexity is mostly sensitive to the affective valance rating, while regional neuro-cortical connectivity levels are mostly sensitive to the affective arousal ratings.
This chapter aims to provide a large-scale collection of digital tape recordings of Cantonese speech and establishing an archive of Cantonese texts based on transcriptions of these recordings. It explains a corpus of Cantonese syllables and words together with other polysyllabic Chinese expressions. The chapter considers generation of relevant lexical information of Cantonese Chinese speech, determination of the processing and production unit of Cantonese speech, and estimation of the code-switched situation in Hong Kong. Sources of the natural Cantonese speech include dialogues of Radio call-in programs, conversations of TV programs, casual chatting among the students in canteen. The chapter also provide useful information of the pervasive code-switching situation in Hong Kong that clearly confounded the traditional language teaching methods in Hong Kong education sector.
In order to automatically extend a treebank of Old French (9 th -13 th c.) with new texts in Old and Middle French (14 th -15 th c.), we need to adapt tools for syntactic annotation. However, these stages of French are subjected to great variation, and parsing historical texts remains an issue. We chose to adapt a symbolic system, the French Metagrammar (FRMG), and develop a lexicon comparable to the Lefff lexicon for Old and Middle French. The final goal of our project is to model the evolution of language through the whole period of Medieval French (9 th -15 th c.).
The article examines the Universal Dependencies (UD) annotation scheme. The UD project is an international initiative to produce treebanks of the world’s languages, whereby the treebanks have been annotated in a cross-linguistically consistent manner. A central aspect of the UD annotation scheme is its analysis of function words. The scheme advocates subordinating function words to content words. This article discusses linguistic and practical motivations behind the UD decision to subordinate function words to content words. It demonstrates that UD choices in this area are not supported linguistically. At the same time, the near convertibility of the UD treebanks to a more linguistically motivated annotation format means that the UD initiative remains of great value to linguistics in general.
The vast majority of successful deep neural networks are trained using variants of stochastic gradient descent (SGD) algorithms. Recent attempts to improve SGD can be broadly categorized into two approaches: (1) adaptive learning rate schemes, such as AdaGrad and Adam and (2) accelerated schemes, such as heavy-ball and Nesterov momentum. In this paper, we propose a new optimization algorithm, Lookahead, that is orthogonal to these previous approaches and iteratively updates two sets of weights. Intuitively, the algorithm chooses a search direction by looking ahead at the sequence of ``fast weights generated by another optimizer. We show that Lookahead improves the learning stability and lowers the variance of its inner optimizer with negligible computation and memory cost. We empirically demonstrate Lookahead can significantly improve the performance of SGD and Adam, even with their default hyperparameter settings on ImageNet, CIFAR-10/100, neural machine translation, and Penn Treebank.
A recently proposed balanced-bracket encoding (Yli-Jyrä and GómezRodríguez 2017) has given us a way to embed all noncrossing dependency graphs into the string space and to formulate their exact arcfactored inference problem (Kuhlmann and Johnsson 2015) as the best string problem in a dynamically constructed and weighted unambiguous context-free grammar. The current work improves the encoding and makes it shallower by omitting redundant brackets from it. The streamlined encoding gives rise to a bounded-depth subset approximation that is represented by a small finite-state automaton. When bounded to 7 levels of balanced brackets, the automaton has 762 states and represents a strict superset of more than 99.9999% of the noncrossing trees available in Universal Dependencies 2.4 (Nivre et al. 2019). In addition, it strictly contains all 15-vertex noncrossing digraphs. When bounded to 4 levels and 90 states, the automaton still captures 99.2% of all noncrossing trees in the reference dataset. The approach is flexible and extensible towards unrestricted graphs, and it suggests tight finite-state bounds for dependency parsing, and for the main existing parsing methods.
The article discusses some problems that arose while implementing two sets of annotation rules for treebanking Ancient Greek (the first set by Bamman and Crane, the second by Celano). Educational uses of treebanking are discussed, together with the most common problems that occurred while treebanking the Iudicium vocalium by 2nd-century writer Lucian of Samosata.
There are different types of discourses and each of them possesses its own particular tools and wars of their linguistic implementation. These phenomena are in the focus of close attention of linguists. The relevance of the chosen study is determined by the importance of high-rate spreading scientific dental knowledge. The effect of such exchange and the quality of the scientific text depends on language skills, basic grammatical norms and communicative qualities of scientific speech. Scientists, who deliberately and consciously have mastered the normative basis of the Ukrainian language, are able to formulate his opinion correctly. It is necessary to know and to use the basic grammatical norms for the correct presentation of scientific thoughts. The purpose of this article is to search for and analyze the frequency of grammatical mistakes recorded in fragments of different genres of scientific dentistry texts: original articles, abstracts, manuals, textbooks, monographs. General scientific methods (observation, comparison, generalization, synthesis, description) and linguistic methods (functional-stylistic, semantic-stylistic, discourse analysis, etc.) were used. In this work the emphasis is placed on the importance of common rules of effective using dictionaries, and the need for an analysis of scientific dental discourse in the normative aspect has been determined. We have pointed out that the grammatical skills of doctors, their linguistic senses and skills to produce high-quality scientific texts are essential components of professional communicative competence. The article provides the details for correct use of verbal nouns and forms of active adjectives as based on our study these are the weak grammar point for young dental academic writers. The frequency of their representation in the analyzed analytical sources has been analyzed. In addition, we offer practical language recommendations that can be helpful for medical and dental professionals.
The guidelines of the Ancient Greek Dependency Treebank 2.0 have been written to annotate Ancient Greek texts.The epigraphic texts, however, pose a challenge for those carrying out morphosyntactic annotation: should we remain as close as possible to the actual epigraphic text, or represent it in an interpreted and normalized version?How should all epigraphic peculiarities which do not have standard editorial representation, such as, for example, punctuation marks, be treated?A small corpus such as that of the inscriptions of the Euboean colonies of Sicily of the archaic and classical period has allowed us to test different options and evaluate the annotation challenges.This contribution is the result of a discussion about the advantages and disadvantages of often opposed annotation possibilities.We present here our first proposal for an adaptation of the guidelines to analyse the morphosyntax of inscriptions, which we hope will stimulate discussion between epigraphists and linguists.In particular, we propose to try to stick to the epigraphic evidence as far as possible and therefore render its complexities (e.g., local alphabets, dialect variants not attested in literary texts, ellipsis, punctuation marks, and word forms which can be linguistically interpreted differently), while trying to preserve consistency with the annotation of literary texts.
The development of code-mixing (CM) NLP systems has significantly gained importance in recent times due to an upsurge in the usage of CM data by multilingual speakers. However, this proves to be a challenging task due to the complexities created by the presence of multiple languages together. The complexities get further compounded by the inconsistencies present in the raw data on social media and other platforms. In this paper, we present a neural stack based dependency parser for CM data of Bengali and English by utilizing pre-existing resources for closely related Hindi and English CM treebank as well as monolingual treebanks for Bengali, Hindi and English. To address the issue of scarcity of annotated resources for Bengali-English CM pair, we present a rule based system to computationally generate a synthetic code-mixing treebank for Bengali and English (Syn-BE) which is used to further improve the accuracy of our dependency parser. For evaluation purpose, we present a dataset of 500 Bengali-English tweets annotated under Universal Dependencies scheme.
Lengths (in words) of projective and non-projective sentences from a Czech UD dependency treebank are compared. It is shown that non-projective sentences are significantly longer (in addition, the same result was obtained in this study also for Arabic, Polish, Russian, and Slovak). The hyperpascal distribution, which was suggested as the model for frequency distribution of sentence length measured in words, fits well the data from both projective and non-projective sentences; however, its parameters attain different values for the two groups. Proportions of non-projective sentences in the treebanks used are presented, together with a discussion on factors which can influence them.
International audience
Standard approaches to treebanking traditionally employ a waterfall model (Sommerville, 2010), where annotation guidelines guide the annotation process and insights from the annotation process in turn lead to subsequent changes in the annotation guidelines.This process remains a very expensive step in creating linguistic resources for a target language, necessitates both linguistic expertise and manual effort to develop the annotations and is subject to inconsistencies in the annotation due to human errors.In this paper, we propose an alternative approach to treebanking-one that requires writing grammars.This approach is motivated specifically in the context of Universal Dependencies, an effort to develop uniform and cross-lingually consistent treebanks across multiple languages.We show here that a bootstrapping approach to treebanking via interlingual grammars is plausible and useful in a process where grammar engineering and treebanking are jointly pursued when creating resources for the target language.We demonstrate the usefulness of synthetic treebanks in the task of delexicalized parsing, a task of interest when working with languages with no linguistic resources and corpora.Experiments with three languages reveal that simple models for treebank generation are cheaper than human annotated treebanks, especially in the lower ends of the learning curves for delexicalized parsing, which is relevant in particular in the context of lowresource languages.
The article aims to be an introduction to the dependency treebanks currently available for Ancient Greek and Latin, i.e., the Ancient Greek and Latin Dependency Treebank (AGLDT), the Index Thomisticus Treebank (IT-TB), the PROIEL Treebank, and the SEMATIA Treebank.Their pipelines for creation of morphosyntactic annotations are presented so as to highlight major commonalities and differences.All treebanks share the same basic underlying formalism, whereby syntactic words are connected to each other to form labeled directed acyclic graphs, and their annotation schemes, although different, are comparable to a very large extent. An introduction to the dependency treebank formalismA dependency treebank is a corpus containing a symbolic representation of the syntax of one or more texts.It can be defined as a set of sentences parsed according to the linguistic formalism of dependency grammar.Most treebanks for Ancient Greek and Latin, i.e., the Ancient Greek and Latin Dependency Treebank (AGLDT), the Index Thomisticus Treebank (IT-TB), the PROIEL Treebank, and the SEMATIA Treebank, are dependency treebanks.Even if, in the present article, I deal only with dependency treebanks, most of what follows in the present section could also be applied, mutatis mutandis, to describe constituency treebanks, such as the Nestle 1904 and SBNLGT Treebanks, 1 the major difference being that in dependency treebanks all nodes except the ROOT node are paired with tokens, 2 while in constituency treebanks nonterminal nodes, which represent phrases, such as VPs or PPs, are also licensed.The parsed sentences in a treebank are formally represented as labeled directed acyclic graphs, where each token, excluding the ROOT node, is annotated
We report on the conversion of the Hamburg Dependency Treebank The HDT consists of more than 200.000 sentences annotated with dependency structure, making every attempt at manual conversion or manual post-processing extremely costly. The conversion employs an unranked tree transducer. This formalism allows to express transformation rules in a concise way, guarantees the well-formedness of the output and is predictable to the rule writers. Together with the release of a converted subset of the HDT spanning 3 million tokens, we release an interactive workbench for writing and refining tree transducer rules. Our conversion achieves a very high labeled accuracy with respect to a manually converted gold standard of 97.3%. Up to now, the conversion effort took about 1000 hours of work.
This paper expands on recent studies of very large treebank collections aiming to find empirical evidence for language universals, specifically for the functionally motivated Dependency Length Minimization (DLM) hypothesis. According to DLM grammars are set up to support the expression of utterances in a way that minimizes the distance between heads and dependents. We construct several incremental baselines that lead from the random free order linearization to the real language by adding various word order constraints. We conduct detailed analyses on 55 treebanks and find that all of the constraints contribute to DLM. We show that DLM on the one hand shapes the regularity and on the other motivates the attested exceptions from canonical word order. The findings contribute to a more fine-grained, differentiated picture of the role of DLM in the interaction of competing constraints on grammar and language use.
This paper attempts to evaluate some of the systematic differences in Uralic Universal Dependencies treebanks from a perspective that would help to introduce reasonable improvements in treebank annotation consistency within this language family. The study finds that the coverage of Uralic languages in the project is already relatively high, and the majority of typically Uralic features are already present and can be discussed on the basis of existing treebanks. Some of the idiosyncrasies found in individual treebanks stem from language-internal grammar traditions, and could be a target for harmonization in later phases.
Background \n \nBack pain originating from multiple spinal disorders is the leading cause of disability worldwide. Adult spinal deformity (ASD) is a complex three-dimensional entity of multiple anatomical and functional disorders and predisposes to decreased health-related quality of life (HRQoL). With rising life expectancy and a growing elderly population it is predicted that an increasing number of symptomatic ASD patients will need surgical treatment. It is noteworthy that the prevalence of different grades of ASD, impact on HRQoL and patient-reported outcome (PRO) of ASD surgery has not earlier been reported in the Finnish population. \n \nThe aims of this study were to assess the reliability and repeatability of radiographic diagnostic imaging of ASD and to produce a culturally adapted and valid Finnish version of the Scoliosis Research Society (SRS) Questionnaire version 30 for spinal deformities. Thereafter the prevalence of ASD, applicability of a simplified version of the SRS-Schwab ASD classification and the Finnish SRS-30 were evaluated in a symptomatic adult patient cohort with prolonged degenerative spinal disorders. Finally, the long-term outcome, complications, patient satisfaction and predictive factors for poor outcome of ASD surgery, were investigated. \n \nMethods \n \nOver one year, a consecutive cohort of adult patients was recruited to Studies I-III after referral to the Central Hospital of Central Finland spine clinic due to prolonged degenerative spinal disease. 637 patients returned the completed HRQoL questionnaires and digital full spine radiographs were obtained. The radiographs of 49 patients were randomly selected for the reliability assessment, and a repeatability study of sagittal spinopelvic measurements with basic software tools was performed by three raters differing in their experience of image rating. The SRS-30 underwent translation and cross-cultural adaptation into Finnish and was subsequently validated and psychometrically tested among 274 patients. The SRS-Schwab ASD classification was graded and simplified dividing into mild, moderate and marked groups. The division was tested along with the Finnish SRS-30 questionnaire during evaluation of the prevalence and HRQoL of patients with sagittal malalignment among symptomatic adult patients with spinal degenerative disease but no pre-known deformity. The 79 patients in Study IV were operated during 2007-2016 in our clinic. The clinical and radiographic outcome, patient satisfaction, predictive factors for poor outcome and complications were analysed using the diagnostic tools renovated and tested in Studies I-III. \n \nResults \n \nThe intra-and interrater reliability of the sagittal spinopelvic measurements proved reliable and repeatable with intraclass correlation coefficients (ICC) between 0.78-0.99 and standard error of measurement (SEM) of 0.80-6.2° or 2.2-5.8mm. Greater rater experience in performing the radiographic measurements decreases, and greater complexity of the measurement landmarks increases, intra- and inter-rater bias. \n \nThe reproducibility and internal consistency (ICC 0.905, SEM 0.17, Cronbach α 0.885) of the Finnish version of the SRS-30 was good. The SRS-30 had discriminative validity in the pain, self-image and satisfaction with management domains compared with other questionnaires. A statistically significant difference between the moderate and marked deformity groups in the SRS-30 domains of function/activity (p=0.022) and self-image/appearance (p=0.016) was found. \n \nOf the 637 patients in the consecutive cohort, 25% had moderate and 11% marked spinal deformities. The patients with marked deformity were significantly older, more overweight and more physically inactive than the others in the study population. The 3-class categorization of the SRS-Schwab ASD classification determined well the severity of sagittal deformity and concomitant loss of function, activity (p=0.004), and self-image/appearance (p=0.030) measured with the SRS-30, and disability with the ODI (p=0.033). \n \nASD operation decreased disability (ODI) and pain (VAS) significantly (p=0.001). Postoperative improvement in radiographic sagittal parameters was significant and maintained at 4-5 years of follow-up (p≤0.001). The mechanical failure of instrumentation of bone resulted in reoperation risk of 13.9% within the first and 29.8% during the 5-year follow-up. According to SRS-30, 49 (62.0%) patients were satisfied or very satisfied with the treatment and 57 (72.1%) would have the same operation again. Depression predicted poor outcome with an odds ratio of 6.97 (p=0.018). \n \nConclusions \n \nThe study comprised an unselected consecutive cohort of adult patients with prolonged degenerative spinal diseases, and thus the results can be generalized. Rater experience had a positive influence on the otherwise good reliability and repeatability of the spinopelvic measurements taken from full spine radiographs. The deformity-specific Finnish SRS-30 translation proved reliable and valid among the study cohort. The simplified categories of the SRS-Schwab ASD classification can detect different grades of deformity and related loss of HRQoL. Long-term radiographic and patient-reported clinical outcomes after the ASD surgery remained significantly better than preoperative scores. Risk for reoperation was highest during the first postoperative year. However good patient satisfaction and outcomes could be achieved irrespective of adverse effects. Depression was the only significant predictive factor for poor outcome after ASD surgery. \n \n \n \nKeywords: adult spinal deformity, ASD, scoliosis, kyphosis, full-spine radiograph, reliability, repeatability, validation, outcome, health-related quality of life, Scoliosis Research Society questionnaire 30, SRS-30, SRS-Schwab ASD classification, spine surgery, pelvic incidence, pelvic tilt, sagittal vertical axis, lumbar lordosis, thoracic kyphosis
The article considers the modernstate of computer (electronic) lexicography, which is traditionally divided into corpus and electronic ones. At first glance theincreasing global computerization has greatly facilitated the work of lexicographers and linguists, however, a number ofproblems have arisen at once. Among the first ones are the principles of corpora compilation, the requirements of which haveconsiderably expanded since the appearance of the first ones, and at present stage of corpus technologies development corporashould be able to answer a wide range of possible inquiries and meet the needs of different users. Most of the principlesdeveloped up till now are standardized and unified for the convenience of both scientific and non-scientific research. The nextkey issue that exists nowadays is not only the inclusion of the main lexicographic postulates, but also consideringtechnological and operational features of the new media (smartphones and tablets). Provided practical lexicography isprimarily intended to meet the needs of the end user, the procedures for determining the needs of users are becomingincreasingly relevant today by involving them in surveys and analysis of logs. It is noted that more and more non-specialists ofthe field are involved in the actual development of lexicographic projects. They perform simple tasks, mainly as volunteers.Such practices also require well-balanced and clear rules in order to avoid creating a low-quality final product. There exist afew online platforms where you can “hire” a number of volunteers to perform such tasks. Another aspect of the developmentof new lexicography, mainly in the English-speaking world, is a special type of lexicography - lexicography for fun. The mainobject of research here is the vocabulary associated with certain hobbies, books, films or computer games. The articleemphasizes the importance and relevance of this subject with anthropocentric cognitive paradigm of modern linguistics. Thestate of research of this phenomenon abroad and in Ukraine is clarified and disclosed. The novelty of the work is to analyzethe current state of research, development and methods of organization of lexicographic work, as well as to indicate theperspectives of its development in Ukraine. References Burkhanov, Igor. 1999. Linguistic Foundations of Ideography: Semantic Analysis and IdeographicDictionaries. Rzeszow: WSP. Cibej, Jaka, Darja Fise, Kosem Iztok. “The role of crowdsourcing in lexicography” (paper presented at eLEX 2015 conference in Herstmonceux Castle (UK) from 11 to 13 August 2015). Down, Ellie. 2005. The Unofficial Guide to Harry Potter. Chichester: Summersdale Publishers Ltd. Gao, Yongwei “The Appification of Dictionaries: From a Chinese Perspective”. (paper presented at eLex 2013 conference in Tallinn (Estonia) from 17 to 19 October 2013). Garside Roger, Geoffrey Leech, and Tony McEnery. 1997. Corpus Annotation: Linguistic Information from Computer Text Corpora. London: Longman. Gouws, Rufus Hjalmar, Ulrich Heid, Wolfgang Schweickard, et al. 2013. An International Encyclopedia ofLexicography. Supplementary. Volume: Recent Developments with Focus on Electronic and Computational Lexicography. Berlin; Boston: De Gruyter Mouton. Granger, Sylviane. 2012. “Introduction: Electronic lexicography-from challenge to opportunity”. Electronic Lexicography, edited by Sylviane Granger, and Magali Paquot. Oxford: Oxford University Press. Hidalgo, Pablo. 2017. Star Wars: The Last Jedi. The Visual Dictionary. N. Y.: DK Publishing. Holmer Louise, Martens von, Monica, and Skoldberg Emma. Making a dictionary app from a lexical database: the case of the Contemporary Dictionary of the Swedish Academy. (paper presented at eLEX 2015 conference in Herstmonceux Castle (UK) from 11 to 13 August 2015) Horot, Yevheniya, Lesya Malimon. 2015. “Suchasna leksykohrafiya: problemy y perspektyvy”. Aktualnipytannya inozemnoyi filolohiyi 3: 42–48. Danchevska, Yuliya. 2014. “Korpusy tekstiv u linhvodydaktytsi: zdobutky ta perspektyvy”. Novapedahohichna dumka 1: 58–60. Dubichynskyy, Volodymyr. 2013. “Ukraynskaya leksykohrafyya: istoriya i sovremennost”. Slavyanskayaleksykohrafyya. Moskva: Azbukovnik. Kulchytska, Tetyana. 1999. Ukrayinska leksykohrafiya XIX – XX st.: bibliohraf. pokazhchyk. NAN Ukrayiny: Lviv. nauk. b-ka im. V. Stefanyka. Kulchytskyy Ihor, Yulia Danchevska, and Ihor Likhnyakevych. 2013. “Deyaki aspekty stvorennya tavykorystannya paralelnykh korpusiv”. Naukovyy visnyk VNU im. Lesi Ukrayinky 20: 48–52. Kulchytskyy Ihor. 2017. “Informatsiyna tekhnolohiya poperedn’oho opratsyuvannya pryrodomovnykh tekstiv. Informatsiyni tekhnolohiyi ta vzayemodiyi”. Kulchytskyy, Ihor. 2002. Kompyuterno-tekhnolohichni aspekty stvorennya suchasnykh leksykohrafichnykh system. – K.: Nats. b-ka Ukrayiny im. V. I. Vernadskoho NAN Ukrayiny Kulchytskyy, Ihor. 2015. “Tekhnolohichni aspekty ukladannya korpusiv tekstiv”. Dani tekstovykh korpusiv u linhvistychnykh doslidzhennyakh. Lviv: Vydavnytstvo Lvivskoi politekhniky. Kupriyanov, Yevhen. 2008. “Kompyuterna leksykohrafiya yak problema suchasnoho movoznavstva(istorychnyy aspekt)”. Visnyk Kharkivskoho natsionalnoho universytetu im. V.N. Karazina 53: 12–16. Levchenko Olena, Ihor Kulchytskyy. 2013. “Tekhnolohiya peretvorennya p’atymovnoho slovnyka porivnyan u elektronnu formu”. Visnyk Natsionalnoho universytetu Lvivska politekhnika 770: 129–138. Lew, Robert. Space restrictions in paper and electronic dictionaries and their implications for the design of production dictionaries. Accessed February 25, 2019. https://repozytorium.amu.edu.pl/bitstream/10593/799/1/Lew_space_restrictions_in_paper_and_electronic_dictionaries.pdf McArthur, Tom. 1986. Worlds of Reference. Cambridge: Cambridge University Press. Meyer, Christian, Andrea Abel. 2017. “User participation in the Internet era”. The Routledge Handbook of Lexicography. Abingdon; New York: Routledge. Perebyynis, Valentyna, Viktor Sorokin. 2009. Tradytsiyna ta kompyuterna leksykohrafiya. Kyiv: Vyd. tsentr KNLU. Polyuha, Lev. 2006. “Ukrayinske slovnytstvo na perelomi tysyacholit”. Ukrayinoznavchi studiyi 6-7:17–25. Reynolds, David. 2000. Star Wars: the Visual Dictionary. The Ultimate Guide to Star Wars Characters and Creatures. N.Y.: Dorling Kindersley. Rundell, Michael. Redefining the dictionary: From print to digital. Accessed February 20, 2019.https://www.kdictionaries.com/kdn/kdn21_2013.pdf Rusanivskyy, Vitaliy, Volodymyr Shyrokov. 2002. “Informatsiyno-linhvistychni osnovy suchasnoyitlumachnoyi leksykohrafiyi”. Movoznavstvo 6: 7–48. Shyrokov, Volodymyr. 2011. Kompyuterna leksykohrafiya. Kyiv: Nauk. dumka. Syvokozova, Viktoriya. 2013. “Do pytannya periodyzatsiyi ukrayinskoyi tlumachnoyi leksykohrafiyi”. Movni i kontseptualni kartyny svitu 43: 61–62. Snizhko, Nataliya. 2017. “Nova dzherelna baza ukrayinskoyi leksykohrafiyi v systemi intehralnykhlinhvistychnykh doslidzhen”. Lyudyna. Kompyuter. Komunikatsiya. Lviv: Vydavnytstvo Lvivskoyipolitekhniky 47–51. Starko, Vasyl. 2017. “Kompyuterni linhvistychni proekty hurtu r2u: stan ta zastosuvannya”. Ukrayinka mova. 3: 86–97. Svensen, Bo. 1993. Practical Lexicography: Principles and Methods of Dictionary-Making. Translated by John Sykes and Kerstin Schofield. Oxford: Oxford University Press. Tarp, Sven. 2012. “Online dictionaries: today and tomorrow”. Lexicographica. De Gruyter. 28: 253–268. Trap-Jensen, Lars. Lexicography between NLP and Linguistics: Aspects of Theory and Practice (paperpresented at eLEX 2017 conference in Leiden (the Netherlands) from 19 to 21 September 2017.
This dissertation investigates the frequency, semantic, and functional characteristics of recurring discontinuous formulaic language in a learner corpus of argumentative and literary essays. Discontinuous sequences of words, or ‘frames’, are recurrent sequences of words that have one or more variable slots. For example, in the * of and it is * to, where the asterisks represent variable slots in the sequences of words. A corpus of English argumentative essays authored by native speakers of English, Japanese, and Spanish is analyzed using modern methods in corpus linguistics to determine which frames are used frequently. Frequent frames are compared between the L1 groups to investigate trends in structure, frequency, and variability.\nThe focus of analysis then narrows to a group of 30 recurring frames comprised of only function words, that is, function word frames such as in the * of, the * of the, and to the * that. The 30 frames are grouped based on structural characteristics, such as noun and preposition-based frames, for further analysis to better understand their semantic and functional characteristics. A lexical database is used to explore the semantic characteristics of fillers of the frames. To determine discourse functions of all instances of each of the 30 target frames, a well-known taxonomy previously applied to continuous sequences of words is adapted and applied to the present context. Discourse functions of the frames are then used as dependent variables in a multinomial logistic regression conducted with four distinct predictor variables: (1) L1, (2) proficiency level, (3) topic of essay, and (4) specific frame. The purpose of the regression is to see which predictor best accounts for discourse function fulfilled by frames from the structural groups.\nFindings from the various analyses first indicate that Japanese learners of English use function word frames at far lower rates than the L1 English and Spanish speakers. Secondly, fillers of the structural groups of function word frames tend to be abstract nouns and the frames largely serve the discourse function of intangible framing of a following noun phrase. In terms of predictor variables, the frames themselves, particularly prepositions, best predict discourse function. The results lend support to the idea that function words, despite carrying little meaning in isolation, are semantically motivated and systematically contribute meaning to larger sequences of words. Pedagogical implications of this study include the teaching of function words from a more phraseological perspective as well as highlighting the connection between frames, fillers, and discourse functions.
On the one hand, use and spread of slang as an important part of the mode of communication has been mostly associated with teenagers' language who by passing the accepted linguistic norms seek proving themselves. Moreover, origination of slang is said to be one of the properties of teenagers at the turn of the century. On the other hand, although language of genders has been studied by several scholars, it gets more interesting when it is paired with this entertaining word play, which due to its very nature-being coarse and direct-has mostly been attributed to males. Females, however, are getting inventive and trying to come up with some slang of their own and it seems that for teenage girls’ slangs are evolving at an even faster clip. Are there really any differences between male and female teenagers in their use of the slang words? If there are, what kind of difference may be found in societies which are perceived to be male dominant? Being concern of the present study, the researchers used both empirical and ethnographic elements of research to find relationships, if any, between use of slang words and the gender of Iranian teenagers, in a society which is perceived to be mostly male-dominant. Two high schools were therefore randomly chosen and six male and six female high schoolers who aged between 16 and 17 and came from upper middle-class families were randomly selected and interviewed using sociolinguistic interview protocols. Preliminary results, based on Chi-Square tests, revealed significant differences between the linguistic behavior of Iranian male and female teenagers. Females were interestingly found to be both more direct and more creative in their use of slang words. Our findings shed doubt on the generally accepted view of scholars on the linguistic (and social) dominance of males in modern Iranian society.
Many government schemes were unsuccessful because lack of proper feedback on the ongoing schemes, where billion dollars investment is going to be in vain. Sentiment analysis is one of best approach to analyse opinions of the peoples on various government schemes. Sentiment analysis and machine learning techniques emerged to analyse huge social media corpora to track people's views on government policies, products and services. Sentiment analysis process consists of various phases which include data discovery, data collection, data pre-processing, and data analysis. Stemming is a process to generate the morphemes in natural language sentences for various applications such as sentiment analysis, information retrieval, and domain analysis. The stemming process involved two major errors, which are over-stemming and under-stemming errors. Most of sentiment analysis natural languages processing applications used Lancaster and Porter stemming algorithms where more than one word inflected into same morpheme, which causes the etymology behaviour of the stemming word and prone to classify the tweets false positives and false negative. The proposed un-prejudice light stemming algorithm prevent etymology behaviour of morpheme and sustain its meaning during stemming process by selecting a word which has maximum number of synonyms in lexical database.
The similarity between words constitutes significant support to tasks in natural language processing. Several works use Lexical resources such as WordNet for semantic similarity and synonym identification. Nevertheless, words out-of-vocabulary or missing links between senses are perceived problems of this approach. Distributional-based proposals like word embeddings have successfully been used to meet such problems, but the lack of contextual information can prevent the achievement of even better results. The distributional models that include contextual information can bring advantages to this area, but these models are still scarcely explored. Therefore, this work studies the advantages of incorporating syntactic information in the distributional models, fostering for better results in semantic similarity approaches. For that purpose, the current work explore existing lexical and distributional techniques regarding the measurement of word similarity in Brazilian Portuguese. Experiments were carried out with the lexical database WordNet, using different techniques over a standard dataset. The results indicate that word embeddings can cover words out of vocabulary and have better results in comparison with lexical approaches. The main contribution of this article is a new approach to apply syntactic context in the training process of word embeddings to a Brazilian Portuguese corpus. The comparison of this model with the outcome of the previous experiments shows sound results and presents relevant complementary aspects.
When biblicalhumanities.org launched in 2013, we naively assumed that software developers would know how to take advantage of high quality, innovative data if it were published in open formats under free licenses. We understood the importance of building communities that know how to take advantage of this data, software and frameworks that make it easily accessible, and tailoring data to specific use cases, but we assumed that this would happen if we just made the data freely available. Freely licensed base texts, morphologies, contextual glosses, syntax treebanks, discourse analyses, lexicons, textual variants, images, conjectural emendations, and grammars are available now, and some are at least as good as commercial resources, but they are not yet widely used. In the last year or so, Jupyter Notebooks and academic conferences have started to bring this data into academic study, but we have barely begun to tap its potential for translation, language learning, and scripture engagement for those who know the languages. To realize this potential, we need to understand the use cases and mindset of potential users and create software and other resources tailored to their needs. Over time, we need to gather feedback from these same users, walk with them as they learn and grow, and welcome them into communities that will create the next generation of resources.
This contains data files needed for FinMeter. This data is complementary for FinMeter Python library described in: Mika Hämäläinen and Khalid Alnajjar (2019). Let's FACE it. Finnish Poetry Generation with Aesthetics and Framing. In <em>the Proceedings of The 12th International Conference on Natural Language Generation</em>. Sources: The pretrained vectors for Finnish (es - I know) and English (en) are from E. Grave, P. Bojanowski, P. Gupta, A. Joulin, T. Mikolov, <em>Learning Word Vectors for 157 Languages. Creative Commons Attribution-Share-Alike License 3.0</em>. See https://fasttext.cc/docs/en/crawl-vectors.html The word2vec model trained on the Finnish Internet ParseBank is from Kanerva, Jenna; Luotolahti, Juhani; Laippala, Veronika; Ginter, Filip: Syntactic N-gram Collection from a Large-Scale Corpus of Internet Finnish. Proceedings of the Sixth International Conference Baltic HLT. 2014. paper. Creative Commons Attribution-ShareAlike 4.0 International License. See http://bionlp.utu.fi/finnish-internet-parsebank.html The Finnish concreteness data has been automatically translated from Brysbaert, Marc, Amy Beth Warriner, and Victor Kuperman. "Concreteness ratings for 40 thousand generally known English word lemmas." <em>Behavior research methods</em> 46.3 (2014): 904-911. Creative Commons Attribution-NonCommercial-NoDerivs 3.0 Unported License. see http://crr.ugent.be/archives/1330
The present study examined the effects of labeling on pain tolerance, sensation, and affect for individuals who are high or low pain catastrophizers, as measured through the Pain Catastrophizing Scale (PCS). Pasticipants completed the PCS and were randomly assigned to 1 of 3 labeling conditions: A maximizing, a minimizing, and a neutral label condition. All participants then took part in a cold-pressor test. The cold-pressor measure of pain tolerance, as well as visual analog scales of sensory and affective ratings of pain, provided the dependent measures. Participants also completed the Anxiety Sensitivity Index (ASI) and the Somatic Amplification Questionnaire (SAQ). Results indicated that high pain catastrophizers have significantly reduced pain tolerance, increased pain sensations, and increased pain unpleasantness compared with low pain catastrophizers. In addition, significant correlations were found between the dependent measures. Main effects for labeling, and interaction effects between pain catastrophizing and labeling, were not supported.
This contains data files needed for FinMeter. This data is complementary for FinMeter Python library described in: Mika Hämäläinen and Khalid Alnajjar (2019). Let's FACE it. Finnish Poetry Generation with Aesthetics and Framing. In <em>the Proceedings of The 12th International Conference on Natural Language Generation</em>. Sources: The pretrained vectors for Finnish (es - I know) and English (en) are from E. Grave, P. Bojanowski, P. Gupta, A. Joulin, T. Mikolov, <em>Learning Word Vectors for 157 Languages. Creative Commons Attribution-Share-Alike License 3.0</em>. See https://fasttext.cc/docs/en/crawl-vectors.html The word2vec model trained on the Finnish Internet ParseBank is from Kanerva, Jenna; Luotolahti, Juhani; Laippala, Veronika; Ginter, Filip: Syntactic N-gram Collection from a Large-Scale Corpus of Internet Finnish. Proceedings of the Sixth International Conference Baltic HLT. 2014. paper. Creative Commons Attribution-ShareAlike 4.0 International License. See http://bionlp.utu.fi/finnish-internet-parsebank.html The Finnish concreteness data has been automatically translated from Brysbaert, Marc, Amy Beth Warriner, and Victor Kuperman. "Concreteness ratings for 40 thousand generally known English word lemmas." <em>Behavior research methods</em> 46.3 (2014): 904-911. Creative Commons Attribution-NonCommercial-NoDerivs 3.0 Unported License. see http://crr.ugent.be/archives/1330
English Abstract: The article analyses the relationship between etymology of political neologisms, ways of their formation and methods of rendering them from English into Russian. Translation strategy is directed at the recipient of the translated text and should therefore be pragmatic and based on the functional and stylistic norms of the translation language. The author concludes that the most efficient way of word-formation in the English language is compounding (morphological neologisms) and derivation of new meaning for the already existing words (semantic neologisms). The translation methods demonstrating high potential are transliteration, calquing (loan translation) and descriptive translation. Adequate translation is based on the following criteria: conciseness, single interpretation, conformity with the linguistic norms of the translation language. Russian Abstract: В работе затрагиваются вопросы взаимосвязи между этимологией, способами образования политических неологизмов и приемами их передачи с английского языка на русский. Выбор переводческой стратегии ориентирован на получателя текста перевода, поэтому должен обеспечивать прагматические задачи в соответствии с функционально-стилистическими нормами языка перевода. В работе делается вывод о том, что наиболее продуктивным способом образования новых слов является процесс словосложения (морфологические неологизмы) и наделения уже существующих в языке слов новым значением (семантические неологизмы); самыми эффективными приемами перевода неологизмов являются транслитерация, калькирование и описательный перевод. Критериями адекватного перевода политических неологизмов служат краткость, однозначность толкования, соответствие нормам языка перевода.
This chapter provides a history of 'the King's English' as a context for an analysis of language, history and power in The Merry Wives of Windsor and the second tetralogy. The trope is used as a rhetorical and ideological tool in performatives. Associated with temperance and honesty, 'the King's English' belongs to a set of defining values of true Englishness. The project to produce this linguistic norm coincides with a homologous project to produce a stable, monetary system of 'good' coin through exclusion of 'bad', 'counterfeit' or 'clipped' coin. These projects testify to a shift of the centre of economic and cultural gravity from the court to the merchant citizen class. Shakespeare's one English comedy centred on English citizens which features his one use of 'the King's English' is shown to engage critically with this ideology, and to set against it an idea of 'our English' as an inclusive mix, the 'gallimaufry' loved by the linguistically extravagant gentleman John Falstaff. The comedy draws out the implications of the banishment of Falstaff in the second tetralogy, which sets history against the project of cultural reformation ideology to produce (the) 'true' English.
This paper examines the relation between gendered language and the processing of non-stereotypical gender representation through a psycholinguistic priming experiment consisting of a self-paced reading test. The experiment tests two things: The processing ease of 3rd person singular pronouns that either match or mismatch the stereotypical gender of their referents; and whether sentences with gendered language affect this. Processing ease is measured by reading time. Sentences with gendered language are used as priming; and pronouns with matching or mismatching stereotypicality as targets. The results showed no difference in the reading time of matching and mismatching pronouns, and no priming effect was found. This could point to the fact that gender stereotypical mismatch does not affect the informants; and that gendered language is so well-integrated in our language that it cannot prime for gender stereotypes. Moreover, the pronoun han was read significantly faster than the pronoun hun, which could indicate that the masculine is expected as a linguistic norm. This is considered in relation to the reflections on the missing priming effect as a result of gendered language being the norm.
The article refers to the concept of intelligentsia as a social group which exerts significant influence on Polish standard patterns. Although the term intelligentsia is vague and questionable, it is well-established term in Polish linguistics, especially in sociolinguistics. Author argues that science communicators (young professional researchers, science journalists, PhD students) represent the young intelligentsia, because these well-educated people pursue their intellectual development and they have sense of public duty. The article examines standard of popular science texts in Internet, new tendencies in written Polish and attitude of young intelligentsia toward traditional linguistic norm. The errors (esp. punctuation and syntax) exemplify impact of technological changes and phenomenon of secondary orality. It would be useful for science communicators to edit carefully their texts. Both researchers and journalists need to improve their writing skills permanently. Nevertheless it must be emphasized that school education and competent teachers seem to have important influence on the linguistic patterns.
Abstract This article presents results from a study on hybrid linguistic norms in translated articles from New York Times made available on UOL website. According to Faraco (2008) and Bagno (2012), there is a difference between norma padrão (a prescriptive norm, but not based on usage) and norma culta (an alternative, usage-based norm). The first one combines normative rules that determine correct linguistic forms, but generally hard to follow by most users, while the second one brings together a set of linguistic forms frequently employed by users, because they are more intuitively accessible, although not subscribed by the conservative standard norm (norma padrão) commonly taught in grammar books and in writing style manuals. The research was meant to verify if journalistic texts translated from English have been as permeable to linguistic forms not subscribed by the prescriptive standard norm, as the ones originally written in Portuguese have proven to be.
Location: Dewberry Hall With over 34,000 students representing 123 countries, George Mason University is a vastly diverse university with students bringing different learning experiences and skill sets with them into the classroom. The study that our team conducted analyzes the essays of native English speakers, as well as the essays of students for whom English is their second language. Our objective when conducting this research was to observe the essays for signs of syntactic complexity and patterns of language errors. Specifically, we looked for subordinating clauses, transitions, subject/verb agreement, article usage, run-on sentences, and fragments. We found that some errors in L1 and L2 populations were consistent with our expectations, but others reveled a more complex understanding of the linguistic norms of the groups studied. The results of these findings will give professors of all disciplines and modalities insight to the challenges that first-year L1 and L2 students confront when faced with a writing assignment.