Papers reviewed and determined not to be word norm studies. Use the flag icon to report errors or suggest re-inclusion.
16504 papers
The communicative role of nonlinear vocal phenomena remains poorly understood since they are difficult to manipulate or even measure with conventional tools. In this study parametric voice synthesis was employed to add pitch jumps, subharmonics/sidebands, and chaos to synthetic human nonverbal vocalizations. In Experiment 1 (86 participants, 144 sounds), chaos was associated with lower valence, and subharmonics with higher dominance. Arousal ratings were not noticeably affected by any nonlinear effects, except for a marginal effect of subharmonics. These findings were extended in Experiment 2 (83 participants, 212 sounds) using ratings on discrete emotions. Listeners associated pitch jumps, subharmonics, and especially chaos with aversive states such as fear and pain. The effects of manipulations in both experiments were particularly strong for ambiguous vocalizations, such as moans and gasps, and could not be explained by a non-specific measure of spectral noise (harmonics-to-noise ratio) – that is, they would be missed by a conventional acoustic analysis. In conclusion, listeners interpret nonlinear vocal phenomena quite flexibly, depending on their type and the kind of vocalization in which they occur. These results showcase the utility of parametric voice synthesis and highlight the need for a more fine-grained analysis of voice quality in acoustic research.
We present our CHARLES-SAARLAND system for the SIGMORPHON 2019 Shared Task on Crosslinguality and Context in Morphology, in task 2, Morphological Analysis and Lemmatization in Context. We leverage the multilingual BERT model and apply several fine-tuning strategies introduced by UDify demonstrating exceptional evaluation performance on morpho-syntactic tasks. Our results show that fine-tuning multilingual BERT on the concatenation of all available treebanks allows the model to learn cross-lingual information that is able to boost lemmatization and morphology tagging accuracy over fine-tuning it purely monolingually. Unlike UDify, however, we show that when paired with additional character-level and word-level LSTM layers, a second stage of fine-tuning on each treebank individually can improve evaluation even further. Out of all submissions for this shared task, our system achieves the highest average accuracy and f1 score in morphology tagging and places second in average lemmatization accuracy.
* Introduction This is the Myanmar ALT of the Asian Language Treebank (ALT) Corpus. Please refer to<br> http://www2.nict.go.jp/astrec-att/member/mutiyama/ALT/index.html<br> for an introduction of the ALT project. The process of building the Myanmar ALT began with sampling about 20,000 sentences from English Wikinews, and then these sentences were translated into Myanmar language.<br> <br> The English Wikinews<br> https://en.wikinews.org/wiki/Main_Page<br> is available under the terms of the Creative Commons Attribution 2.5 License.<br> https://creativecommons.org/licenses/by/2.5/ Myanmar ALT has been developed by NICT and UCSY. The license of Myanmar ALT is Creative Commons Attribution-NonCommercial-ShareAlike 4.0 International (CC BY-NC-SA 4.0) License<br> https://creativecommons.org/licenses/by-nc-sa/4.0/ <br> * Contents - data: Myanmar ALT treebank
Tree-LSTMs have been used for tree-based sentiment analysis over Stanford Sentiment Treebank, which allows the sentiment signals over hierarchical phrase structures to be calculated simultaneously. However, traditional tree-LSTMs capture only the bottom-up dependencies between constituents. In this paper, we propose a tree communication model using graph convolutional neural network and graph recurrent neural network, which allows rich information exchange between phrases constituent tree. Experiments show that our model outperforms existing work on bidirectional tree-LSTMs in both accuracy and efficiency, providing more consistent predictions on phrase-level sentiments.
Animal phobias are one of the most prevalent mental disorders. We analysed how fear and disgust, two emotions involved in their onset and maintenance, are elicited by common phobic animals. In an online survey, the subjects rated 25 animal images according to elicited fear and disgust. Additionally, they completed four psychometrics, the Fear Survey Schedule II (FSS), Disgust Scale - Revised (DS-R), Snake Questionnaire (SNAQ), and Spider Questionnaire (SPQ). Based on a redundancy analysis, fear and disgust image ratings could be described by two axes, one reflecting a general negative perception of animals associated with higher FSS and DS-R scores and the second one describing a specific aversion to snakes and spiders associated with higher SNAQ and SPQ scores. The animals can be separated into five distinct clusters: (1) non-slimy invertebrates; (2) snakes; (3) mice, rats, and bats; (4) human endo- and exoparasites (intestinal helminths and louse); and (5) farm/pet animals. However, only snakes, spiders, and parasites evoke intense fear and disgust in the non-clinical population. In conclusion, rating animal images according to fear and disgust can be an alternative and reliable method to standard scales. Moreover, tendencies to overgeneralize irrational fears onto other harmless species from the same category can be used for quick animal phobia detection.
This paper suggests one way to enhance the ability of Chinese learners to analyze sentences. It is building a treebank and visualizing it as a syntactic tree(or parsed tree) and providing it to learners. The process of building a treebank, which is the most important key in this method, is divided into three parts and described in detail.
An experiment was conducted to examine the independent and interactive influence of the audio and visual channels of information in television on viewers’ emotional experience. Audio-only, video-only, and audiovisual television content was presented as psychological stimuli, while participants completed continuous-response measures (CRMs) to index over-time changes in emotional experience of positive valence, negative valence, and arousal. Positive valence and arousal means were significantly influenced by channel over time. Participants reported the most positive emotional experience during audiovisual exposure. The channel and time interaction did not significantly affect negative valence ratings. However, positive valence, negative valence, and arousal ratings were significantly influenced by the interaction of channel and specific message content.
BACKGROUND: The tendency to inhibit anger (anger-in) is associated with increased pain. This relationship may be explained by the negative affectivity hypothesis (anger-in increases negative affect that increases pain). Alternatively, it may be explained by the cognitive resource hypothesis (inhibiting anger limits attentional resources for pain modulation). METHODS: A well-validated picture-viewing paradigm was used in 98 healthy, pain-free individuals who were low or high on anger-in to study the effects of anger-in on emotional modulation of pain and attentional modulation of pain. Painful electrocutaneous stimulations were delivered during and in between pictures to evoke pain and the nociceptive flexion reflex (NFR; a physiological correlate of spinal nociception). Subjective and physiological measures of valence (ratings, facial/corrugator electromyogram) and arousal (ratings, skin conductance) were used to assess reactivity to pictures and emotional inhibition in the high anger-in group. RESULTS: The high anger-in group reported less unpleasantness, showed less facial displays of negative affect in response to unpleasant pictures, and reported greater arousal to the pleasant pictures. Despite this, both groups experienced similar emotional modulation of pain/NFR. By contrast, the high anger-in group did not show attentional modulation of pain. CONCLUSIONS: These findings support the cognitive resource hypothesis and suggest that overuse of emotional inhibition in high anger-in individuals could contribute to cognitive resource deficits that in turn contribute to pain risk. Moreover, anger-in likely influenced pain processing predominantly via supraspinal (e.g., cortico-cortical) mechanisms because only pain, but not NFR, was associated with anger-in.
In the present article, a novel emotional complexity marker is proposed for classification of discrete emotions induced by affective video film clips. Principal Component Analysis (PCA) is applied to full-band specific phase space trajectory matrix (PSTM) extracted from short emotional EEG segment of 6 s, then the first principal component is used to measure the level of local neuronal complexity. As well, Phase Locking Value (PLV) between right and left hemispheres is estimated for in order to observe the superiority of local neuronal complexity estimation to regional neuro-cortical connectivity measurements in clustering nine discrete emotions (fear, anger, happiness, sadness, amusement, surprise, excitement, calmness, disgust) by using Long-Short-Term-Memory Networks as deep learning applications. In tests, two groups (healthy females and males aged between 22 and 33 years old) are classified with the accuracy levels of [Formula: see text] and [Formula: see text] through the proposed emotional complexity markers and and connectivity levels in terms of PLV in amusement. The groups are found to be statistically different ( p << 0.5) in amusement with respect to both metrics, even if gender difference does not lead to different neuro-cortical functions in any of the other discrete emotional states. The high deep learning classification accuracy of [Formula: see text] is commonly obtained for discrimination of positive emotions from negative emotions through the proposed new complexity markers. Besides, considerable useful classification performance is obtained in discriminating mixed emotions from each other through full-band connectivity features. The results reveal that emotion formation is mostly influenced by individual experiences rather than gender. In detail, local neuronal complexity is mostly sensitive to the affective valance rating, while regional neuro-cortical connectivity levels are mostly sensitive to the affective arousal ratings.
This chapter aims to provide a large-scale collection of digital tape recordings of Cantonese speech and establishing an archive of Cantonese texts based on transcriptions of these recordings. It explains a corpus of Cantonese syllables and words together with other polysyllabic Chinese expressions. The chapter considers generation of relevant lexical information of Cantonese Chinese speech, determination of the processing and production unit of Cantonese speech, and estimation of the code-switched situation in Hong Kong. Sources of the natural Cantonese speech include dialogues of Radio call-in programs, conversations of TV programs, casual chatting among the students in canteen. The chapter also provide useful information of the pervasive code-switching situation in Hong Kong that clearly confounded the traditional language teaching methods in Hong Kong education sector.
Neural parsers obtain state-of-the-art results on benchmark treebanks for constituency parsing-but to what degree do they generalize to other domains? We present three results about the generalization of neural parsers in a zero-shot setting: training on trees from one corpus and evaluating on out-of-domain corpora. First, neural and non-neural parsers generalize comparably to new domains. Second, incorporating pre-trained encoder representations into neural parsers substantially improves their performance across all domains, but does not give a larger relative improvement for out-of-domain treebanks. Finally, despite the rich input representations they learn, neural parsers still benefit from structured output prediction of output trees, yielding higher exact match accuracy and stronger generalization both to larger text spans and to out-of-domain corpora. We analyze generalization on English and Chinese corpora, and in the process obtain state-of-the-art parsing results for the Brown, Genia, and English Web treebanks.
Abstract This chapter summarises the contributions to the volume The Normative Animal? On the Anthropological Significance of Social, Moral and Linguistic Norms. The contributions are divided into three sections in line with the tripartite division of the types of norms discussed in the volume. The key claims of the individual chapters are presented and set into relation to one another, and a number of issues raised by competition between the claims are highlighted. This prepares the ground for an assessment of the normative animal thesis in the light of the varying accounts both of specific deontic phenomena and of normativity in general. Central issues concern the concepts of social norms and conventions, the relative importance of coordination and cooperation, the nature and role of collective intentionality, the place of norms in evolutionary explanations, and the structure of normative action guidance. Decisive for the normative animal thesis are the questions as to whether moral principles and linguistic rules are correctly characterised as both real and deontic in the same senses in which these characterisations apply to social norms.
The article examines the Universal Dependencies (UD) annotation scheme. The UD project is an international initiative to produce treebanks of the world’s languages, whereby the treebanks have been annotated in a cross-linguistically consistent manner. A central aspect of the UD annotation scheme is its analysis of function words. The scheme advocates subordinating function words to content words. This article discusses linguistic and practical motivations behind the UD decision to subordinate function words to content words. It demonstrates that UD choices in this area are not supported linguistically. At the same time, the near convertibility of the UD treebanks to a more linguistically motivated annotation format means that the UD initiative remains of great value to linguistics in general.
The vast majority of successful deep neural networks are trained using variants of stochastic gradient descent (SGD) algorithms. Recent attempts to improve SGD can be broadly categorized into two approaches: (1) adaptive learning rate schemes, such as AdaGrad and Adam and (2) accelerated schemes, such as heavy-ball and Nesterov momentum. In this paper, we propose a new optimization algorithm, Lookahead, that is orthogonal to these previous approaches and iteratively updates two sets of weights. Intuitively, the algorithm chooses a search direction by looking ahead at the sequence of ``fast weights generated by another optimizer. We show that Lookahead improves the learning stability and lowers the variance of its inner optimizer with negligible computation and memory cost. We empirically demonstrate Lookahead can significantly improve the performance of SGD and Adam, even with their default hyperparameter settings on ImageNet, CIFAR-10/100, neural machine translation, and Penn Treebank.
A recently proposed balanced-bracket encoding (Yli-Jyrä and GómezRodríguez 2017) has given us a way to embed all noncrossing dependency graphs into the string space and to formulate their exact arcfactored inference problem (Kuhlmann and Johnsson 2015) as the best string problem in a dynamically constructed and weighted unambiguous context-free grammar. The current work improves the encoding and makes it shallower by omitting redundant brackets from it. The streamlined encoding gives rise to a bounded-depth subset approximation that is represented by a small finite-state automaton. When bounded to 7 levels of balanced brackets, the automaton has 762 states and represents a strict superset of more than 99.9999% of the noncrossing trees available in Universal Dependencies 2.4 (Nivre et al. 2019). In addition, it strictly contains all 15-vertex noncrossing digraphs. When bounded to 4 levels and 90 states, the automaton still captures 99.2% of all noncrossing trees in the reference dataset. The approach is flexible and extensible towards unrestricted graphs, and it suggests tight finite-state bounds for dependency parsing, and for the main existing parsing methods.
The article discusses some problems that arose while implementing two sets of annotation rules for treebanking Ancient Greek (the first set by Bamman and Crane, the second by Celano). Educational uses of treebanking are discussed, together with the most common problems that occurred while treebanking the Iudicium vocalium by 2nd-century writer Lucian of Samosata.
There are different types of discourses and each of them possesses its own particular tools and wars of their linguistic implementation. These phenomena are in the focus of close attention of linguists. The relevance of the chosen study is determined by the importance of high-rate spreading scientific dental knowledge. The effect of such exchange and the quality of the scientific text depends on language skills, basic grammatical norms and communicative qualities of scientific speech. Scientists, who deliberately and consciously have mastered the normative basis of the Ukrainian language, are able to formulate his opinion correctly. It is necessary to know and to use the basic grammatical norms for the correct presentation of scientific thoughts. The purpose of this article is to search for and analyze the frequency of grammatical mistakes recorded in fragments of different genres of scientific dentistry texts: original articles, abstracts, manuals, textbooks, monographs. General scientific methods (observation, comparison, generalization, synthesis, description) and linguistic methods (functional-stylistic, semantic-stylistic, discourse analysis, etc.) were used. In this work the emphasis is placed on the importance of common rules of effective using dictionaries, and the need for an analysis of scientific dental discourse in the normative aspect has been determined. We have pointed out that the grammatical skills of doctors, their linguistic senses and skills to produce high-quality scientific texts are essential components of professional communicative competence. The article provides the details for correct use of verbal nouns and forms of active adjectives as based on our study these are the weak grammar point for young dental academic writers. The frequency of their representation in the analyzed analytical sources has been analyzed. In addition, we offer practical language recommendations that can be helpful for medical and dental professionals.
The guidelines of the Ancient Greek Dependency Treebank 2.0 have been written to annotate Ancient Greek texts.The epigraphic texts, however, pose a challenge for those carrying out morphosyntactic annotation: should we remain as close as possible to the actual epigraphic text, or represent it in an interpreted and normalized version?How should all epigraphic peculiarities which do not have standard editorial representation, such as, for example, punctuation marks, be treated?A small corpus such as that of the inscriptions of the Euboean colonies of Sicily of the archaic and classical period has allowed us to test different options and evaluate the annotation challenges.This contribution is the result of a discussion about the advantages and disadvantages of often opposed annotation possibilities.We present here our first proposal for an adaptation of the guidelines to analyse the morphosyntax of inscriptions, which we hope will stimulate discussion between epigraphists and linguists.In particular, we propose to try to stick to the epigraphic evidence as far as possible and therefore render its complexities (e.g., local alphabets, dialect variants not attested in literary texts, ellipsis, punctuation marks, and word forms which can be linguistically interpreted differently), while trying to preserve consistency with the annotation of literary texts.
In order to automatically extend a treebank of Old French (9 th -13 th c.) with new texts in Old and Middle French (14 th -15 th c.), we need to adapt tools for syntactic annotation. However, these stages of French are subjected to great variation, and parsing historical texts remains an issue. We chose to adapt a symbolic system, the French Metagrammar (FRMG), and develop a lexicon comparable to the Lefff lexicon for Old and Middle French. The final goal of our project is to model the evolution of language through the whole period of Medieval French (9 th -15 th c.).
The development of code-mixing (CM) NLP systems has significantly gained importance in recent times due to an upsurge in the usage of CM data by multilingual speakers. However, this proves to be a challenging task due to the complexities created by the presence of multiple languages together. The complexities get further compounded by the inconsistencies present in the raw data on social media and other platforms. In this paper, we present a neural stack based dependency parser for CM data of Bengali and English by utilizing pre-existing resources for closely related Hindi and English CM treebank as well as monolingual treebanks for Bengali, Hindi and English. To address the issue of scarcity of annotated resources for Bengali-English CM pair, we present a rule based system to computationally generate a synthetic code-mixing treebank for Bengali and English (Syn-BE) which is used to further improve the accuracy of our dependency parser. For evaluation purpose, we present a dataset of 500 Bengali-English tweets annotated under Universal Dependencies scheme.
The central object of computer lexicography is a computer or electronic dictionary, which must have a sufficiently large vocabulary, provide the consistent extraction of information depending on the user’s need and provide complete grammatical information about the words of input and output languages. Taking into account the current trend in the development of special terminological dictionaries, the authors propose an English-Belarusian-Russian dictionary of technical terms. At the initial stage of the work the dictionary was named TechLex and covers the following subject areas: architecture and construction, water supply, information technology, pedagogy, transport communications, economics, energy-supply. Currently, each subject area of the dictionary is located in the Internet GoogleTable and contains about 1000 terms. It has the possibility to be simultaneously filled by several teachers. The linguistic database of the dictionary is not created by the traditional way of processing a large number of paper dictionaries and combining the received translations. Lexis from sequential processing of scientific and technical English periodicals of particular subject areas is the base of it. The software of the proposed electronic dictionary is designed taking into account the analysis of modern electronic multilingual translation dictionaries and is a client-server application in Java programming language. The client part of the system contains a mobile application for the Android operating system, which was tested on tablets and smartphones with different screen diagonals. The interface of TechLex dictionary is designed in such a way that only a single zone is activated according to the query, so there is no need to view all the subject areas of the dictionary. The proposed TechLex dictionary is the first technical multilingual electronic dictionary with an English-Belarusian-Russian version.
Resources for syntactic parsing for Indonesian are very limited, as there are only two dependency treebanks publicly available and both are small in size.Not only that, we found out that the word segmentation method used by both treebanks needs improvement.Therefore, in this work we proposed a revision for one of these treebanks, Indonesian Parallel Universal Dependencies treebank.Besides improving word segmentation, we also improved POS tagging and syntactic annotations.Because in Indonesian grammar there are some special structures, we also proposed how to adjust UDv2 annotation guidelines with those Indonesian grammar rules.To evaluate the quality of the new treebank, we built Indonesian dependency parser model using Parsito (UDPipe) parser.Using ten-fold cross-validation, the model that built using the revised treebank had UAS of 83.33% and LAS of 79.39%, over the original treebank with UAS of 73.32% and LAS of 65.98%.
Transformers have a potential of learning longer-term dependency, but are limited by a fixed-length context in the setting of language modeling. We propose a novel neural architecture Transformer-XL that enables learning dependency beyond a fixed length without disrupting temporal coherence. It consists of a segment-level recurrence mechanism and a novel positional encoding scheme. Our method not only enables capturing longer-term dependency, but also resolves the context fragmentation problem. As a result, Transformer-XL learns dependency that is 80% longer than RNNs and 450% longer than vanilla Transformers, achieves better performance on both short and long sequences, and is up to 1,800+ times faster than vanilla Transformers during evaluation. Notably, we improve the state-ofthe-art results of bpc/perplexity to 0.99 on en-wiki8, 1.08 on text8, 18.3 on WikiText-103, 21.8 on One Billion Word, and 54.5 on Penn Treebank (without finetuning). When trained only on WikiText-103, Transformer-XL manages to generate reasonably coherent, novel text articles with thousands of tokens. Our code, pretrained models, and hyperparameters are available in both Tensorflow and PyTorch 1.
Study Context: The question of whether relationships between valence and arousal might differ among older and younger adults has not yet been totally clarified. Previous studies focused on only age-related mean-differences, but in the current study mean differences and variance in emotional ratings for the International Affective Picture Systems (IAPS) were both examined in Japanese older and younger adults. METHODS: Participants were 31 older adults (69 ± 5.17 years) and 31 younger adults (19 ± 0.77 years). Each picture was projected on the screen for about 5 s in random order and participants subsequently rated its valence from "unhappy" to "happy" and its arousal from "calm" to "exciting," using 9-point scales. RESULTS: Pearson's correlation analysis showed that positive and negative valence tended to be negatively correlated with arousal in both age groups. The 95% Confidence Intervals for positive arousal in older adults included those of younger adults. Arousal ratings for negative pictures were higher than those for neutral pictures, and those for neutral pictures were higher than those for positive pictures in older adults. There were no significant differences between arousal ratings for neutral pictures and positive pictures in younger adults. CONCLUSION: Older adults tended to rate the pictures as more arousing, with higher arousal ratings for negative pictures than for positive pictures, and the variance of positive arousal in older adults was the highest variance for all conditions. The results of this study suggest that older adults may be sensitive to harmful negative experiences in order to make them less aversive.
In this paper we present a dependency treebank of travel domain sentences in Modern Standard Arabic. The text comes from a translation of the English equivalent sentences in the Basic Traveling Expressions Corpus. The treebank dependency representation is in the style of the Columbia Arabic Treebank. The paper motivates the effort and discusses the construction process and guidelines. We also present parsing results and discuss the effect of domain and genre difference on parsing.
A constituency treebank is a key component for deep syntactic parsing of natural language sentences. For Indonesian, this task is unfortunately hindered by the fact that the only one constituency treebank publicly available is rather small with just over 1000 sentences, and not only that, it employs a format incompatible with readily available constituency treebank processing tools. In this work, we present a conversion of the existing Indonesian constituency treebank to the widely accepted Penn Treebank format. Specifically, the conversion adjusts the bracketing format for compound words as well as the POS tagset according to the Penn Treebank format. In addition, we revised the word segmentation and POS tagging of a number of tokens. Finally, we performed an evaluation on the treebank quality by employing the Shift-Reduce parser from Stanford CoreNLP to create a parser model. A 10-fold cross-validated experiment on the parser model yields an F1-score of 70.90%.
This paper shows the extent to which treebanks of Ancient Greek play a central role in the ongoing Pedalion project at the University of Leuven. Building on diverse treebanks readily available today, the project aims to make progress in the automated parsing of classical and postclassical Greek texts. Rather than developing new technology as such, our project endeavours to make deliberate and methodical use of the technology that already exists, essentially by combining and adapting both technology and data. This contribution offers a 'roadmap' of our project, surveying (a) the existing work on which we can rely, (b) the strategies which we adopt to reach better results in the automated processing of Ancient Greek and (c) the deliverables that have already been realised or are forthcoming.
This paper presents challenges and observations on creating a code-switching treebank based on ongoing annotation efforts of a Turkish-German spoken corpus following the Universal Dependencies annotation scheme. We present and discuss a number of issues that arise because of the need for consistent multilingual annotation within a single treebank, as well as the informal language which is where code-switching is observed most. Besides proposing solutions to these issues, our aim in this paper is to stimulate discussion and facilitate consistency over upcoming code-switching annotation projects.
Head-driven phrase structure grammar (HPSG) enjoys a uniform formalism representing rich contextual syntactic and even semantic meanings. This paper makes the first attempt to formulate a simplified HPSG by integrating constituent and dependency formal representations into head-driven phrase structure. Then two parsing algorithms are respectively proposed for two converted tree representations, division span and joint span. As HPSG encodes both constituent and dependency structure information, the proposed HPSG parsers may be regarded as a sort of joint decoder for both types of structures and thus are evaluated in terms of extracted or converted constituent and dependency parsing trees. Our parser achieves new state-of-the-art performance for both parsing tasks on Penn Treebank (PTB) and Chinese Penn Treebank, verifying the effectiveness of joint learning constituent and dependency structures. In details, we report 96.33 F1 of constituent parsing and 97.20\% UAS of dependency parsing on PTB.
Building a treebank from scratch can easily be an elaborate, highly time consuming task, especially when working with a minority language with moderately complex morphology and no existing resources. It is also then typically true that language experts and informants with suitable skill sets are a very scarce resource. In this experiment I have attempted to work in parallel on building NLP resources while gathering and annotating the treebank. In particular, I aim to build a decent coverage morphologically annotated lexicon suitable for rule-based morphological analysis as well as accompanying rules for basic morphosyntactic analysis. I propose here a workflow, that I have found useful in avoiding redoing same work with related NLP resource construction.
We introduce the first German treebank for Twitter microtext, annotated within the framework of Universal Dependencies. The new treebank includes over 12,000 tokens from over 500 tweets, independently annotated by two human coders. In the paper, we describe the data selection and annotation process and present baseline parsing results for the new testsuite.
The article aims to be an introduction to the dependency treebanks currently available for Ancient Greek and Latin, i.e., the Ancient Greek and Latin Dependency Treebank (AGLDT), the Index Thomisticus Treebank (IT-TB), the PROIEL Treebank, and the SEMATIA Treebank.Their pipelines for creation of morphosyntactic annotations are presented so as to highlight major commonalities and differences.All treebanks share the same basic underlying formalism, whereby syntactic words are connected to each other to form labeled directed acyclic graphs, and their annotation schemes, although different, are comparable to a very large extent. An introduction to the dependency treebank formalismA dependency treebank is a corpus containing a symbolic representation of the syntax of one or more texts.It can be defined as a set of sentences parsed according to the linguistic formalism of dependency grammar.Most treebanks for Ancient Greek and Latin, i.e., the Ancient Greek and Latin Dependency Treebank (AGLDT), the Index Thomisticus Treebank (IT-TB), the PROIEL Treebank, and the SEMATIA Treebank, are dependency treebanks.Even if, in the present article, I deal only with dependency treebanks, most of what follows in the present section could also be applied, mutatis mutandis, to describe constituency treebanks, such as the Nestle 1904 and SBNLGT Treebanks, 1 the major difference being that in dependency treebanks all nodes except the ROOT node are paired with tokens, 2 while in constituency treebanks nonterminal nodes, which represent phrases, such as VPs or PPs, are also licensed.The parsed sentences in a treebank are formally represented as labeled directed acyclic graphs, where each token, excluding the ROOT node, is annotated
We report on the conversion of the Hamburg Dependency Treebank The HDT consists of more than 200.000 sentences annotated with dependency structure, making every attempt at manual conversion or manual post-processing extremely costly. The conversion employs an unranked tree transducer. This formalism allows to express transformation rules in a concise way, guarantees the well-formedness of the output and is predictable to the rule writers. Together with the release of a converted subset of the HDT spanning 3 million tokens, we release an interactive workbench for writing and refining tree transducer rules. Our conversion achieves a very high labeled accuracy with respect to a manually converted gold standard of 97.3%. Up to now, the conversion effort took about 1000 hours of work.
In this paper we describe the early stage application of the Universal Dependencies to an Italian corpus from social media developed for shared tasks related to irony and stance detection.The development of this novel resource (TWITTIR Ò-UD) serves a twofold goal: it enriches the scenario of treebanks for social media and for Italian, and it paves the way for a more reliable extraction of a larger variety of morphological and syntactic features to be used by sentiment analysis tools.On the one hand, social media texts are especially hard to parse and the limited amount of resources for training and testing NLP tools further damages the situation.On the other hand, we thought that adding the Universal Dependencies format to the fine-grained annotation for irony, that was previously applied on TWITTIR Ò, might meaningfully help in the investigation of possible relationships between syntax and semantics of the uses of figurative language, irony in particular.
Dependency parser is one of the most important fundamental tools in the natural language processing, which extracts structure of sentences and determines the relations between words based on the dependency grammar. The dependency parser is proper for free order languages, such as Persian. In this paper, data-driven dependency parser has been developed with the help of phrase-structure parser for Persian. The defined feature space in each parser is one of the important factors in its success. Our goal is to generate and extract appropriate features to dependency parsing of Persian sentences. To achieve this goal, new semantic and syntactic features have been defined and added to the MSTParser by stacking method. Semantic features are obtained by using word clustering algorithms based on syntagmatic analysis and syntactic features are obtained by using the Persian phrase-structure parser and have been used as bit-string. Experiments have been done on the Persian Dependency Treebank (PerDT) and the Uppsala Persian Dependency Treebank (UPDT). The results indicate that the definition of new features improves the performance of the dependency parser for the Persian. The achieved unlabeled attachment score for PerDT and UPDT are 89.17% and 88.96% respectively.
Background \n \nBack pain originating from multiple spinal disorders is the leading cause of disability worldwide. Adult spinal deformity (ASD) is a complex three-dimensional entity of multiple anatomical and functional disorders and predisposes to decreased health-related quality of life (HRQoL). With rising life expectancy and a growing elderly population it is predicted that an increasing number of symptomatic ASD patients will need surgical treatment. It is noteworthy that the prevalence of different grades of ASD, impact on HRQoL and patient-reported outcome (PRO) of ASD surgery has not earlier been reported in the Finnish population. \n \nThe aims of this study were to assess the reliability and repeatability of radiographic diagnostic imaging of ASD and to produce a culturally adapted and valid Finnish version of the Scoliosis Research Society (SRS) Questionnaire version 30 for spinal deformities. Thereafter the prevalence of ASD, applicability of a simplified version of the SRS-Schwab ASD classification and the Finnish SRS-30 were evaluated in a symptomatic adult patient cohort with prolonged degenerative spinal disorders. Finally, the long-term outcome, complications, patient satisfaction and predictive factors for poor outcome of ASD surgery, were investigated. \n \nMethods \n \nOver one year, a consecutive cohort of adult patients was recruited to Studies I-III after referral to the Central Hospital of Central Finland spine clinic due to prolonged degenerative spinal disease. 637 patients returned the completed HRQoL questionnaires and digital full spine radiographs were obtained. The radiographs of 49 patients were randomly selected for the reliability assessment, and a repeatability study of sagittal spinopelvic measurements with basic software tools was performed by three raters differing in their experience of image rating. The SRS-30 underwent translation and cross-cultural adaptation into Finnish and was subsequently validated and psychometrically tested among 274 patients. The SRS-Schwab ASD classification was graded and simplified dividing into mild, moderate and marked groups. The division was tested along with the Finnish SRS-30 questionnaire during evaluation of the prevalence and HRQoL of patients with sagittal malalignment among symptomatic adult patients with spinal degenerative disease but no pre-known deformity. The 79 patients in Study IV were operated during 2007-2016 in our clinic. The clinical and radiographic outcome, patient satisfaction, predictive factors for poor outcome and complications were analysed using the diagnostic tools renovated and tested in Studies I-III. \n \nResults \n \nThe intra-and interrater reliability of the sagittal spinopelvic measurements proved reliable and repeatable with intraclass correlation coefficients (ICC) between 0.78-0.99 and standard error of measurement (SEM) of 0.80-6.2° or 2.2-5.8mm. Greater rater experience in performing the radiographic measurements decreases, and greater complexity of the measurement landmarks increases, intra- and inter-rater bias. \n \nThe reproducibility and internal consistency (ICC 0.905, SEM 0.17, Cronbach α 0.885) of the Finnish version of the SRS-30 was good. The SRS-30 had discriminative validity in the pain, self-image and satisfaction with management domains compared with other questionnaires. A statistically significant difference between the moderate and marked deformity groups in the SRS-30 domains of function/activity (p=0.022) and self-image/appearance (p=0.016) was found. \n \nOf the 637 patients in the consecutive cohort, 25% had moderate and 11% marked spinal deformities. The patients with marked deformity were significantly older, more overweight and more physically inactive than the others in the study population. The 3-class categorization of the SRS-Schwab ASD classification determined well the severity of sagittal deformity and concomitant loss of function, activity (p=0.004), and self-image/appearance (p=0.030) measured with the SRS-30, and disability with the ODI (p=0.033). \n \nASD operation decreased disability (ODI) and pain (VAS) significantly (p=0.001). Postoperative improvement in radiographic sagittal parameters was significant and maintained at 4-5 years of follow-up (p≤0.001). The mechanical failure of instrumentation of bone resulted in reoperation risk of 13.9% within the first and 29.8% during the 5-year follow-up. According to SRS-30, 49 (62.0%) patients were satisfied or very satisfied with the treatment and 57 (72.1%) would have the same operation again. Depression predicted poor outcome with an odds ratio of 6.97 (p=0.018). \n \nConclusions \n \nThe study comprised an unselected consecutive cohort of adult patients with prolonged degenerative spinal diseases, and thus the results can be generalized. Rater experience had a positive influence on the otherwise good reliability and repeatability of the spinopelvic measurements taken from full spine radiographs. The deformity-specific Finnish SRS-30 translation proved reliable and valid among the study cohort. The simplified categories of the SRS-Schwab ASD classification can detect different grades of deformity and related loss of HRQoL. Long-term radiographic and patient-reported clinical outcomes after the ASD surgery remained significantly better than preoperative scores. Risk for reoperation was highest during the first postoperative year. However good patient satisfaction and outcomes could be achieved irrespective of adverse effects. Depression was the only significant predictive factor for poor outcome after ASD surgery. \n \n \n \nKeywords: adult spinal deformity, ASD, scoliosis, kyphosis, full-spine radiograph, reliability, repeatability, validation, outcome, health-related quality of life, Scoliosis Research Society questionnaire 30, SRS-30, SRS-Schwab ASD classification, spine surgery, pelvic incidence, pelvic tilt, sagittal vertical axis, lumbar lordosis, thoracic kyphosis
The article considers the modernstate of computer (electronic) lexicography, which is traditionally divided into corpus and electronic ones. At first glance theincreasing global computerization has greatly facilitated the work of lexicographers and linguists, however, a number ofproblems have arisen at once. Among the first ones are the principles of corpora compilation, the requirements of which haveconsiderably expanded since the appearance of the first ones, and at present stage of corpus technologies development corporashould be able to answer a wide range of possible inquiries and meet the needs of different users. Most of the principlesdeveloped up till now are standardized and unified for the convenience of both scientific and non-scientific research. The nextkey issue that exists nowadays is not only the inclusion of the main lexicographic postulates, but also consideringtechnological and operational features of the new media (smartphones and tablets). Provided practical lexicography isprimarily intended to meet the needs of the end user, the procedures for determining the needs of users are becomingincreasingly relevant today by involving them in surveys and analysis of logs. It is noted that more and more non-specialists ofthe field are involved in the actual development of lexicographic projects. They perform simple tasks, mainly as volunteers.Such practices also require well-balanced and clear rules in order to avoid creating a low-quality final product. There exist afew online platforms where you can “hire” a number of volunteers to perform such tasks. Another aspect of the developmentof new lexicography, mainly in the English-speaking world, is a special type of lexicography - lexicography for fun. The mainobject of research here is the vocabulary associated with certain hobbies, books, films or computer games. The articleemphasizes the importance and relevance of this subject with anthropocentric cognitive paradigm of modern linguistics. Thestate of research of this phenomenon abroad and in Ukraine is clarified and disclosed. The novelty of the work is to analyzethe current state of research, development and methods of organization of lexicographic work, as well as to indicate theperspectives of its development in Ukraine. References Burkhanov, Igor. 1999. Linguistic Foundations of Ideography: Semantic Analysis and IdeographicDictionaries. Rzeszow: WSP. Cibej, Jaka, Darja Fise, Kosem Iztok. “The role of crowdsourcing in lexicography” (paper presented at eLEX 2015 conference in Herstmonceux Castle (UK) from 11 to 13 August 2015). Down, Ellie. 2005. The Unofficial Guide to Harry Potter. Chichester: Summersdale Publishers Ltd. Gao, Yongwei “The Appification of Dictionaries: From a Chinese Perspective”. (paper presented at eLex 2013 conference in Tallinn (Estonia) from 17 to 19 October 2013). Garside Roger, Geoffrey Leech, and Tony McEnery. 1997. Corpus Annotation: Linguistic Information from Computer Text Corpora. London: Longman. Gouws, Rufus Hjalmar, Ulrich Heid, Wolfgang Schweickard, et al. 2013. An International Encyclopedia ofLexicography. Supplementary. Volume: Recent Developments with Focus on Electronic and Computational Lexicography. Berlin; Boston: De Gruyter Mouton. Granger, Sylviane. 2012. “Introduction: Electronic lexicography-from challenge to opportunity”. Electronic Lexicography, edited by Sylviane Granger, and Magali Paquot. Oxford: Oxford University Press. Hidalgo, Pablo. 2017. Star Wars: The Last Jedi. The Visual Dictionary. N. Y.: DK Publishing. Holmer Louise, Martens von, Monica, and Skoldberg Emma. Making a dictionary app from a lexical database: the case of the Contemporary Dictionary of the Swedish Academy. (paper presented at eLEX 2015 conference in Herstmonceux Castle (UK) from 11 to 13 August 2015) Horot, Yevheniya, Lesya Malimon. 2015. “Suchasna leksykohrafiya: problemy y perspektyvy”. Aktualnipytannya inozemnoyi filolohiyi 3: 42–48. Danchevska, Yuliya. 2014. “Korpusy tekstiv u linhvodydaktytsi: zdobutky ta perspektyvy”. Novapedahohichna dumka 1: 58–60. Dubichynskyy, Volodymyr. 2013. “Ukraynskaya leksykohrafyya: istoriya i sovremennost”. Slavyanskayaleksykohrafyya. Moskva: Azbukovnik. Kulchytska, Tetyana. 1999. Ukrayinska leksykohrafiya XIX – XX st.: bibliohraf. pokazhchyk. NAN Ukrayiny: Lviv. nauk. b-ka im. V. Stefanyka. Kulchytskyy Ihor, Yulia Danchevska, and Ihor Likhnyakevych. 2013. “Deyaki aspekty stvorennya tavykorystannya paralelnykh korpusiv”. Naukovyy visnyk VNU im. Lesi Ukrayinky 20: 48–52. Kulchytskyy Ihor. 2017. “Informatsiyna tekhnolohiya poperedn’oho opratsyuvannya pryrodomovnykh tekstiv. Informatsiyni tekhnolohiyi ta vzayemodiyi”. Kulchytskyy, Ihor. 2002. Kompyuterno-tekhnolohichni aspekty stvorennya suchasnykh leksykohrafichnykh system. – K.: Nats. b-ka Ukrayiny im. V. I. Vernadskoho NAN Ukrayiny Kulchytskyy, Ihor. 2015. “Tekhnolohichni aspekty ukladannya korpusiv tekstiv”. Dani tekstovykh korpusiv u linhvistychnykh doslidzhennyakh. Lviv: Vydavnytstvo Lvivskoi politekhniky. Kupriyanov, Yevhen. 2008. “Kompyuterna leksykohrafiya yak problema suchasnoho movoznavstva(istorychnyy aspekt)”. Visnyk Kharkivskoho natsionalnoho universytetu im. V.N. Karazina 53: 12–16. Levchenko Olena, Ihor Kulchytskyy. 2013. “Tekhnolohiya peretvorennya p’atymovnoho slovnyka porivnyan u elektronnu formu”. Visnyk Natsionalnoho universytetu Lvivska politekhnika 770: 129–138. Lew, Robert. Space restrictions in paper and electronic dictionaries and their implications for the design of production dictionaries. Accessed February 25, 2019. https://repozytorium.amu.edu.pl/bitstream/10593/799/1/Lew_space_restrictions_in_paper_and_electronic_dictionaries.pdf McArthur, Tom. 1986. Worlds of Reference. Cambridge: Cambridge University Press. Meyer, Christian, Andrea Abel. 2017. “User participation in the Internet era”. The Routledge Handbook of Lexicography. Abingdon; New York: Routledge. Perebyynis, Valentyna, Viktor Sorokin. 2009. Tradytsiyna ta kompyuterna leksykohrafiya. Kyiv: Vyd. tsentr KNLU. Polyuha, Lev. 2006. “Ukrayinske slovnytstvo na perelomi tysyacholit”. Ukrayinoznavchi studiyi 6-7:17–25. Reynolds, David. 2000. Star Wars: the Visual Dictionary. The Ultimate Guide to Star Wars Characters and Creatures. N.Y.: Dorling Kindersley. Rundell, Michael. Redefining the dictionary: From print to digital. Accessed February 20, 2019.https://www.kdictionaries.com/kdn/kdn21_2013.pdf Rusanivskyy, Vitaliy, Volodymyr Shyrokov. 2002. “Informatsiyno-linhvistychni osnovy suchasnoyitlumachnoyi leksykohrafiyi”. Movoznavstvo 6: 7–48. Shyrokov, Volodymyr. 2011. Kompyuterna leksykohrafiya. Kyiv: Nauk. dumka. Syvokozova, Viktoriya. 2013. “Do pytannya periodyzatsiyi ukrayinskoyi tlumachnoyi leksykohrafiyi”. Movni i kontseptualni kartyny svitu 43: 61–62. Snizhko, Nataliya. 2017. “Nova dzherelna baza ukrayinskoyi leksykohrafiyi v systemi intehralnykhlinhvistychnykh doslidzhen”. Lyudyna. Kompyuter. Komunikatsiya. Lviv: Vydavnytstvo Lvivskoyipolitekhniky 47–51. Starko, Vasyl. 2017. “Kompyuterni linhvistychni proekty hurtu r2u: stan ta zastosuvannya”. Ukrayinka mova. 3: 86–97. Svensen, Bo. 1993. Practical Lexicography: Principles and Methods of Dictionary-Making. Translated by John Sykes and Kerstin Schofield. Oxford: Oxford University Press. Tarp, Sven. 2012. “Online dictionaries: today and tomorrow”. Lexicographica. De Gruyter. 28: 253–268. Trap-Jensen, Lars. Lexicography between NLP and Linguistics: Aspects of Theory and Practice (paperpresented at eLEX 2017 conference in Leiden (the Netherlands) from 19 to 21 September 2017.
The battle against HIV is one of the objectives in our century. Along these lines, among the HIV-contaminated patients, one of the most perilous and remarkable with its entanglements is those with lung pathologies. Concurring clinical arranging of the illness, such patients may introduce Tuberculosis, Pneumocystis jirovecii, Cytomegaloviruses, Candidiasis, Toxoplasmosis and so on. The exploration by logical examination foundation of lung illness was done among the inpatient people in measure of 48.37 (77%) of them were given tuberculosis and 11 (23%) with Interstitial Lung Disease (ILD). Studies were introduced on HIV-positive patients who were isolated by the randomization methods. Among 37 patients with tuberculosis, 29 (78%) had AFB (corrosive quick bacillius) with Gexpert, HAIN strategies, 6 (22%) were analyzed by imaging techniques (HRCT, chest X-beam) and serum ADA level. As per past investigations, there were no relationships between's serum ADA level rises at HIV-positive patients (p esteem 0.05). Among 11 patients gave ILD Pneumocystis jirovecii were distinguished at 5 (45%), 3 (27.5%) were given every day mortality, 3 took a Co-Trimaxozole treatment analyzed by imaging strategies. Clinical viability was endorsed by the nearness of pneumocystis starting point. At the second phase of the examination was discovered a relationship between's various Cd4 cell check and imaging rating. In this way, among absolute number of 119 HIV-positive patients, 38 (32%) had penetration zones, 53 (44%) had an obliteration, 20 (17%) dispersal, 8 (7%) mediastinal lymphadenopathy. Measurement results p esteem 0.000424, hence there is immediate relationship. There are heap aspiratory conditions related with HIV, extending from intense contaminations to constant noncommunicable infections. The study of disease transmission of these illnesses has changed essentially in the time of far reaching antiretroviral treatment. Assessment of the HIV-tainted patient includes evaluation of the seriousness of ailment and an intensive yet effective quest for authoritative conclusion, which may include various etiologies at the same time. Significant pieces of information to an analysis incorporate clinical and social history, segment subtleties, for example, travel and geology of habitation, substance use, sexual practices, and domiciliary and detainment status. CD4 cell tally is a colossally valuable proportion of resistant capacity and hazard for HIV-related infections, and assists slender with bringing down the differential. Cautious history of current side effects and physical assessment with specific thoughtfulness regarding extrapulmonary signs are vital early advances. Numerous adjunctive research facility studies can recommend or preclude specific findings. Pneumonic capacity testing (PFT) may help in portrayal of a few incessant noninfectious sicknesses quickened by HIV. Chest radiograph and registered tomography (CT) check take into account arrangement of sicknesses by pathognomonic imaging designs, albeit numerous irresistible conditions present atypically, especially with lower CD4 tallies. At last, conclusive finding with sputum, bronchoscopy with bronchoalveolar lavage, or lung tissue is regularly required. It is of most extreme significance to keep up a serious extent of doubt for HIV in any case undiscovered patients, as the principal introduction of HIV might be through an intense pneumonic disease. The assessment of respiratory side effects in HIV-contaminated patients can be trying for various reasons. Respiratory manifestations are a continuous objection among HIV-contaminated people and might be brought about by a wide range of ailments. The range of pneumonic sicknesses in HIV-tainted patients incorporates both HIV-related and non-HIV-related conditions. The HIV-related aspiratory conditions incorporate both entrepreneurial contaminations (OIs) and neoplasms. The OIs include bacterial, mycobacterial, contagious, viral, and parasitic pathogens. Every one of these OIs and neoplasms has a trademark clinical and radiographic introduction. Nonetheless, there can be impressive variety and cover in these introductions. Thusly, no group of stars of manifestations, physical assessment discoveries, research facility anomalies, and chest radiographic discoveries is pathognomonic or explicit for a specific malady. Thus, a complete microbiologic or pathologic determination is desirable over empiric treatment at whatever point conceivable. Symptomatic tests incorporate societies from sputum and blood and from respiratory examples acquired by obtrusive systems, for example, bronchoscopy, thoracentesis, figured tomography (CT)- guided transthoracic needle goal, thoracoscopy, mediastinoscopy, and open-lung biopsy. This section portrays the recurrence of respiratory side effects, the range of pneumonic diseases that can influence HIV-tainted patients, and an indicative way to deal with the assessment of respiratory manifestations in HIV-contaminated patients, featuring certain parts of the clinical introduction that might be valuable in separating the most widely recognized OIs and neoplasms. The trademark chest radiographic introductions of the most well-known OIs and neoplasms are portrayed and blueprints of 3 case situations are introduced to outline differential judgments and significant analytic and restorative choices for an assortment of clinical and radiographic introductions. For subtleties on explicit indicative tests and treatment regimens for every one of the OIs and neoplasms, see explicit parts inside the HIV InSite Knowledge Base.
На иранских языках говорили многочисленные племена и народности, сыгравшие важную роль в мировой истории. К основным иранским языкам относятся персидский, таджикский, дари, афганский (пушту), осетинский, курдский, белуджский и др. Наиболее распространенным и статусным иранским языком в настоящее время является персидский. Предок современного персидского языка древнеперсидский сформировался еще в середине I тыс. до н.э. на территории западной части Иранского нагорья в области Фарс. После подчинения Александром Македонским царства Ахеменидов в нем официальным языком стал греческий, функционировавший долгие столетия, и лишь в III в. н.э. с установлением гегемонии Сасанидов официальным языком в государстве стал персидский. В результате завоевания Ирана арабами в 637-652 гг. н.э. официальное функционирование среднеперсидского языка надолго прекратилось. Официальным языком арабского Халифата стал арабский. Это продолжалось до IX в. С начала X в. началось бурное развитие новоперсидского языка и персидской литературы. В настоящее время персидский является государственным языком большого и многонационального государства Иран. Персы доминирующий этнос в государстве. На персидском происходит обучение и в школах, начиная с 1-го класса. Делопроизводство также осуществляется исключительно на персидском. Другие языки в официальной сфере не используются. Исторически персидский язык оказал огромное влияние не только на иранские, но и на многие тюркские и индийские языки. На базе классического персидского языка сформировались современный персидский, таджикский и дари. Высоким является и статус таджикского языка. Он полностью используется во всех сферах деятельности. Объем научных исследований таджикского языка не уступает персидскому. Дариязычное население проживает в Афганистане. Оно составляет около 40 жителей страны. Афганцы (пуштуны) являются одним из крупнейших ираноязычных этносов. Они живут в Афганистане и Пакистане. В Афганистане язык пушту является официальным наряду с дари. Другой иранский народ осетины живет в центральной части Кавказа по обеим сторонам Главного Кавказского хребта. В результате ассимиляционных процессов численность осетиноговорящего населения имеет тенденцию к сокращению. Крупными иранскими языками являются также курдский и белуджский. Несмотря на многочисленность курдоговорящего населения, этот язык не имеет высокого официального статуса. Лишь в Ираке в курдских районах курдский был объявлен официальным наряду с арабским. Другой крупный иранский язык белуджский ни в одном государстве не имеет официального статуса. Однако белуджи хорошо чувствуют и охраняют языковую норму. The Iranian languages were spoken by numerous tribes and nationalities, which played an important role in world history. The main Iranian languages include Persian, Tajik, Dari, Afghan (Pashto), Ossetian, Kurdish, Balochi, etc. The most common and status Iranian language is currently Persian. The ancestor of the modern Persian language ancient Persian was formed in the middle of the first millennium BC in the western part of the Iranian highlands in the region of Fars. After Alexander the Great had subjugated the Achaemenid kingdom, Greek became an official language there, functioning for centuries, and only in the 3rd c. AD with the establishment of hegemony of the Sassanids, Persian became the official language in the state. As a result of the conquest of Iran by the Arabs in 637-652 AD the official functioning of the Middle Persian language ceased for a long time. The official language of the Arabic Caliphate is Arabic. This continued until the 9th century. The rapid development of the New Persian language and Persian literature began in the early 10th century. It is currently the state language of the large and multinational state of Iran. Persians are the dominant nation in the state. It takes place in schools, starting from the 1st grade. Office work is also carried out exclusively in Persian. Other languages are not used in the official sphere. Historically, the Persian language has had a huge impact not only on Iranian, but also on many Turkic and Indian. On the basis of the classical Persian language, modern Persian, Tajik and Dari were formed. The status of the Tajik language is also high. It is fully used in all fields of activity. The volume of scientific studies of the Tajik language is not inferior to the Persian. Daria-speaking population lives in Afghanistan. It makes up about 40 of the countrys population. Afghans (Pashtuns) are one of the largest Iranian-speaking ethnic groups. They live in Afghanistan and Pakistan. In Afghanistan, the Pashto language is an official language alongside with Dari. Another Iranian people, the Ossetians, live in the central part of the Caucasus, on both sides of the Main Caucasian Range. As a result of assimilation processes, the Ossetian-speaking population tends to decline. The major Iranian languages are also Kurdish and Balochi. Despite the large Kurdish-speaking population, this language does not have a high official status. Only in Iraq in Kurdish areas Kurdish was declared official alongside with Arabic. The other major Iranian language, Balochi, has no official status in any state. However, the Baluchis feel and properly preserve the linguistic norm.