Papers reviewed and determined not to be word norm studies. Use the flag icon to report errors or suggest re-inclusion.
18265 papers
The article deals with the characteristic feature of the period of language instability, namely the conflict between the norm and codification in one of the Russian language areas that is subject to greatest variation. The author analyses establishment of the new norm in the process of assigning grammatical category of gender to borrowed indeclinable nouns. The study of this lexical group shows that approximately half of the new words form gender according to the principle of semantic analogy. This fact allows us to maintain that in the Russian of the XXI c. there is a process of broadening in the sphere of application of the rule about assigning gender to indeclinable inanimate onyms and abbreviations. Previously having a limited sphere of application, now this rule also extends to indeclinable inanimate appellatives. However the action of the “older” norm, according to which indeclinable nouns were assigned neuter gender, is also manifested, creating conditions for word variation.
This paper describes a set of Prolog modules with predicates to access the information stored in the lexical database WordNet. The aim is to use the defined predicates for empowering (fuzzy) logic programming systems for approximate reasoning tasks. To achieve this goal it is necessary to have means to calculate the degree of relationship between the words. Because WordNet relates words but does not give graded information of the relation between those words, it is necessary to implement standard methods to compute that gradation.
Machine translation (MT) is directly linked to its evaluation in order to both compare different MT system outputs and analyse system errors so that they can be addressed and corrected. As a consequence, MT evaluation has become increasingly important and popular in the last decade, leading to the development of MT evaluation metrics aiming at automatically assessing MT output. Most of these metrics use reference translations in order to compare system output, and the most well-known and widely spread work at lexical level. In this study we describe and present a linguistically-motivated metric, VERTa, which aims at using and combining a wide variety of linguistic features at lexical, morphological, syntactic and semantic level. Before designing and developing VERTa a qualitative linguistic analysis of data was performed so as to identify the linguistic phenomena that an MT metric must consider (Comelles et al. 2017). In the present study we introduce VERTa’s design and architecture and we report the experiments performed in order to develop the metric and to check the suitability and interaction of the linguistic information used. The experiments carried out go beyond traditional correlation scores and step towards a more qualitative approach based on linguistic analysis. Finally, in order to check the validity of the metric, an evaluation has been conducted comparing the metric’s performance to that of other well-known state-of-the-art MT metrics.
Part of speech (POS) tagging, the assignment of syntactic categories for words in running text, is significant to natural language processing as a preliminary task in applications such as speech processing, information extraction, and others. Urdu language processing presents a challenge due to the dual behaviour of various Urdu POS tags in differing situations (morphosyntactic ambiguity). This paper addresses this challenge by developing a novel tagging approach using linear-chain conditional random fields (CRF). Our work is the first instance of a CRF approach for Urdu POS tagging. The proposed model employs a strong, stable and balanced language-independent as well as language dependent feature set. The language-dependent feature considered includes part-of-speech tag of the previous word and suffix of the current word while the language-independent features includes the ‘context words window’. Our approach was evaluated against support vector machine techniques for Urdu POS—considered as state of the art—on two benchmark datasets. The results show our CRF approach to improve upon the F-measure of prior attempts by 8.3–8.5%.
Independent recollection-familiarity (RF) ratings are sometimes collected to measure subjective experiences of recollection and familiarity during recognition. Although the RF ratings task purports to measure the degree to which each recognition state is experienced, the rating scale has been worded in terms of confidence rather than amount. Given prior evidence that wording influences recognition and remember/know judgments, we compared RF rating scales worded in terms of amount versus confidence across 2 groups. A robust levels-of-processing effect occurred on both recollection and familiarity ratings, and its magnitude was similar across scale wording. Scale wording did not influence recognition, and, most importantly, it had little influence on ratings of recollection and familiarity. These findings suggest that participants may use confidence to rate amount, or vice versa. Regardless, researchers should align their task instructions and scale wording, and should publish them. Such alignment and transparency is crucial for interpreting measures of the memory states that arise during recognition memory. (PsycINFO Database Record (c) 2019 APA, all rights reserved).
Musicological texts about classical music frequently include detailed technical discussions concerning the works being analysed. These references can be specific (e.g. C sharp in the treble clef) or general (fugal passage, Thor’s Hammer). Experts can usually identify the features in question in music scores but a means of performing this task automatically could be very useful for experts and beginners alike. Following work on textual question answering over many years as co-organisers of the QA tasks at the Cross Language Evaluation Forum, we decided in 2013 to propose a new type of task where the input would be a natural language phrase, together with a music score in MusicXML, and the required output would be one or more matching passages in the score. We report here on 3 years of the C@merata task at MediaEval. We describe the design of the task, the evaluation methods we devised for it, the approaches adopted by participant systems and the results obtained. Finally, we assess the progress which has been made in aligning natural language text with music and map out the main steps for the future. The novel aspects of this work are: (1) the task itself, linking musical references to actual music scores, (2) the evaluation methods we devised, based on modified versions of precision and recall, applied to demarcated musical passages, and (3) the progress which has been made in analysing and interpreting detailed technical references to music within texts.
COMPASS-XP is a dataset of matched photographic and X-ray images of single objects, made available for use in Machine Learning & Computer Vision research, in particular in the context of transport security. Objects are imaged in multiple poses, and accompanied by metadata including labels for whether we consider the object to be dangerous in the context of aviation. Object classes overlap with those in the popular ImageNet Large Scale Visual Recognition Challenge class set and the WordNet lexical database, and identifiers for shared classes in both schemes are also provided.
The article deals with a gender-oriented experimental phonetic study of explicit male negative evaluative utterances in modern English everyday dialogical speech. The realization of the category of evaluation in speech is in the limelight of modern sociolinguistic research aimed at establishing the linguistic behaviour peculiarities and verbalization of roles, norms and values attributed to men and women by society. Evaluation is objectified in language units of lexical, phonetic, syntactic and discourse levels. Negative evaluative utterances verbalise the negative evaluation of the object that doesn’t meet the subject’s expectations according to certain criteria. On the basis of auditory and acoustic analyses prosodic markers of male negative rational and emotional evaluative utterances in communication between men and in communication with women of the same social status have been outlined: pitch pattern, loudness, tempo and pausation. Invariant prosodic features of male negative evaluative utterances in communication between representatives of the same and different sex but of the same social status have been established: mid level pre-head; descending stepping and mid level scales; mid, narrowed and wide ranges of intonation group; high and low falling terminal tones; falling and rising-falling pitch contour; moderate tempo; moderate loudness; sad and dramatic timbre; voice pitch frequency maximum on head and nucleus of the first or second intonation group; peak intensity on nucleus and the first rhythm group; minimal average syllable duration; small and average unfilled tentative reflective pauses. Variant prosodic features of the following male negative evaluative utterances have been distinguished: 1. Rational evaluative utterances in communication between men: level pitch contour; accelerated tempo; reduced loudness; peak intensity on another rhythm group than pre-head, head and nucleus; small unfilled and average filled tentative choice pauses; 2. Rational evaluative utterances in communication between men and women: ascending stepping scale; the Low Low-Rise and the Rise; falling-rising pitch contour; increased loudness; small unfilled tentative psychological pauses; short filled tentative reflective pauses; intemporal psychological pauses; 3. Emotional evaluative utterances in communication between men: increased and reduced loudness; peak intensity on another rhythm group than pre-head, head and nucleus; small unfilled tentative psychological pauses; intemporal psychological, choice and reflective pauses; 4. Emotional evaluative utterances in communication between men and women: ascending stepping scale; widened and narrow ranges of intonation group; the Low Low-Rise and the High Rise-Fall; short average syllable duration; small filled tentative choice pauses; average filled tentative reflective pauses.
The article explores the problems of lexical and grammatical and specially legal interpretation of certain criminal procedural norms, in particular, those that regulate the victim's right to procedural communication in criminal proceedings. It is noted that the guiding principle of lawful interpretation is the principle of dialogical communication, where dialogue is understood as a dynamic and constructive way of thinking, creating, interpreting, leading from the analysis of legal and technical errors of legal norms to effective law-making and enforcement and developing in a spiral, because every dialogue must continue the previous ones and prepare the next ones.The notion of a lawful interpretation of a sectoral (criminal-procedural) legal norm is defined asa special independent form of official and informal interpretation, an actual need in which it exists before, during and after the application of the criminal-procedural norm, and is intended to ensure its interpretative evolution for the sake of effective legalization.It is emphasized that the importance and necessity of lexical and grammatical and special-legal interpretation is dictated by: the presence of gaps in sectoral (criminal-procedural) legislation; the existence of conflicts in sectoral (criminal procedural) legislation; availability of valuation concepts in sectoral (criminal procedural) legislation; the presence of issues related to legal and technical errors in sectoral (criminal) law; the presence of problems of the degree of legal regulation of the compositional construction of criminal procedural norms; the presence of issues related to the appeal to the linguistic structural elements of the criminal procedural norm.
This paper introduces a new experimental protocol for studying mental representations of urban soundscapes through a simulation process. Subjects are asked to create a full soundscape by means of a dedicated software tool, coupled with a structured sound data set. This paradigm is used to characterize urban sound environment representations by analyzing the sound classes that were used to simulate the auditory scenes. A rating experiment of the soundscape pleasantness using a seven-point bipolar semantic scale is conducted to further refine the analysis of the simulated urban acoustic scenes. Results show that (1) a semantic characterization in terms of presence/absence of sound sources is an effective way to characterize urban soundscape pleasantness, and (2) acoustic pressure levels computed for specific sound sources better characterize the appraisal than the acoustic pressure level computed over the overall soundscape.
The compilation of dictionaries to be used as special textbooks when learning a native or a non-native language is a special area of lexicography. Educational dictionaries have their own purpose, audience, and didactic focus. These determine the corpus of the language units described, the parameters of their lexicographic description, the content and structure of the dictionary text and entries, and the inclusion of special appendixes. Educational dictionaries also differ from reference dictionaries in their selectivity, which is expressed in orientation exclusively towards the literary norm, a purposeful and rigorous selection of language sources and language units, along with the parameters of their description and illustrative material. The specificity of educational dictionaries explains the limited scope of their use and the relatively small number of users. The sphere of modern communication is expanding and becoming more diverse. In the process of communication, a person is confronted with a large number of foreign words and expressions unknown to him, special names and terms from different areas of human knowledge, nominations, which are outside the literary language. A modern user is interested in more detailed information about language units: their origin and peculiar features, the way a word functions or expresses a concept in a language; the possibilities of using it in different communicative spheres, its morphological links, etc. The educational dictionaries cannot satisfy the overall need for such information. Therefore, it is necessary to refer to other types of dictionaries, in particular to reference dictionaries. Observations suggest that explanatory reference dictionaries have a certain educational potential and can be used as a supplement for educational dictionaries. Among their peculiar features it is necessary to note an extended vocabulary, which provides information on the language units that are not described in educational dictionaries; more detailed information on the form, content and usage of a lexical unit; the specifics of borrowed units’ functions in one’s native language and the source language, etc. The possible use of explanatory reference dictionaries for educational purposes substantiates the need to further improve contemporary dictionaries in a number of areas, including the update rate, the fullness, availability and convenience and information, selective use of the data contained in the dictionary, etc.
This paper presents NorNE, a manually annotated corpus of named entities which extends the annotation of the existing Norwegian Dependency Treebank. Comprising both of the official standards of written Norwegian (Bokmål and Nynorsk), the corpus contains around 600,000 tokens and annotates a rich set of entity types including persons, organizations, locations, geo-political entities, products, and events, in addition to a class corresponding to nominals derived from names. We here present details on the annotation effort, guidelines, inter-annotator agreement and an experimental analysis of the corpus using a neural sequence labeling architecture.
Introduction: Magnetic resonance imaging (MRI) is gold-standard for investigating Degenerative Cervical Myelopathy (DCM), a disabling disease triggered by compression of the spinal cord following degenerative changes of adjacent structures. Quantifiable compression correlates poorly with disease and language describing compression in radiological reports is un-standardised. Study design: Retrospective chart review. Objectives: 1) Identify terminology in radiological reporting of cord compression and elucidate relationships between language and quantitative measures 2) Evaluate language’s ability to distinguish myelopathic from asymptomatic compression 3) Explore correlations between quantitative or qualitative features and symptom severity 4) Investigate the influence of quantitative and qualitative measures on surgical referrals. Methods: From all cervical spine MRIs conducted during one year at a tertiary centre (N = 1123), 166 patients had reported cord compression. For each sp)
The given article deals with the research of the examples of lexicalsemantic interference of English and Russian in the field of professional-oriented translation. The study of interference in the framework of technical translation remains relevant in connection with the fact that the translation of technical literature requires knowledge not only of the basics of translation, but also of professional knowledge in a particular branch of science and technology. The article is aimed to analyze the characteristics of lexical-semantic interference in teaching translation in a language university. The methods of the experiment (collection of practical research material) and analysis (quantitative and semantic analysis of the materials obtained) are used in the research. The materials of the linguistic experiment allowed us to select some examples of the most significant violations of the lexical-semantic norms of the Russian language when translating commonly used terms in an Englishlanguage technical text. The results of the study revealed the main characteristics of lexical-semantic interference in the translation of a technical text from the point of view of the divergence of semantics of commonly used vocabulary in the context of professional-oriented translation. The results of the study can be used in the study of English, in teaching various aspects of translation, in studying issues of general linguistics and theory of language.
People tend to like stimuli—ranging from human faces to text—that are prototypical, and thus easily processed. However, recent research has suggested that less typical stimuli may be preferred in creative contexts, such as fine art or music lyrics. In an archival sample of movie scripts, we tested whether genre-typicality predicted film ratings as a function of rater role (novice audience member or expert film critic). Genre-typicality was operationalized as the profile correlations between linguistic arcs (across five segments, or acts) for each script and within-genre averages. We predicted (1) that critics would prefer more disfluent (genre-atypical) films and general audiences would prefer fluent (genre-typical) films, and (2) that these differences would be most pronounced for genres expected to be more entertaining (e.g., action/adventure) than challenging (e.g., tragedy). Partly consistent with our hypotheses, the results showed that critics gave higher ratings to action/adventure films with less typical positive emotion arcs. However, regardless of audience-member or professional-critic status, higher ratings were attributed to films that were more genre-atypical (or disfluent), in terms of analytic thinking, narrative action, and emotional tone, across all genres except family/kids films. Such findings support the growing literature on the appeal of disfluency in the arts and have relevance for researchers in psychology and computer science who are interested in computational linguistic approaches to attitudes, film, and literature.
This paper investigates the lexical immigration of “arabismos” (Arabic loanwords) in the Spanish language. It compares the differences between old arabismos (integrated during the Middle Ages) and modern arabismos (integrated in contemporary time) in terms of their adaptation to the Spanish norms of spelling, phonetics, semantics and morphosyntactic. It also analyses the use of the arabismos and their immigration process during their first phase of presence in the written Spanish language (or predecessor language to Spanish). Through the study of a corpus collected from the databases of the Royal Spanish Academy (RAE), this paper has come to conclude that the anticipated differences between the two types of arabismos has been affirmed, and that the loss of adaption of the modern arabismos to the Spanish linguistic patterns has resulted in a maintenance of them as foreignisms.
In the light of the most recent critical debate, sixteenth-century Petrarchism has been divested of the simple dichotomy between norm and rejection, similarity and dissimilarity, imitation and deviation in relation to Petrarch's model or Bembo's codification, and qualified as a complex and composite movement in which both significant constants and equally significant variations should be identified. In the frame of this dialectic, we analyse Michelangelo Buonarroti's Rime in comparison to the original model of Rerum vulgarium fragmenta. The analysis highlights the fact that Michelangelo's genius distorts the reference model and deviates from it in a material, tragic and expressionist sense, rather than offering a harmonious result of a strict observance of Petrarchism. Michelangelo, however, achieves this effect by employing the same rhetorical and expressive tools as Petrarch. The paper presents a comparative analysis illustrated by numerous examples of the rhetorical figures of antithesis, oxymoron, synonymy and a wide array of metaphors and lexical choices.
⚠️ This version of spaCy requires downloading <strong>new models</strong>. You can use the <code>spacy validate</code> command to find out which models need updating, and print update instructions. If you've been training <strong>your own models</strong>, you'll need to <strong>retrain them</strong> with the new version. ✨ New features and improvements Tagger, Parser, NER and Text Categorizer <strong>NEW:</strong> Experimental ULMFit/BERT/Elmo-like pretraining (see #2931) via the new <code>spacy pretrain</code> command. This pre-trains the CNN using BERT's cloze task. A new trick we're calling <em>Language Modelling with Approximate Outputs</em> is used to apply the pre-training to smaller models. The pre-training outputs CNN and embedding weights that can be used in <code>spacy train</code>, using the new <code>-t2v</code> argument. <strong>NEW:</strong> Allow parser to do joint word segmentation and parsing. If you pass in data where the tokenizer over-segments, the parser now learns to merge the tokens. Make parser, tagger and NER faster, through better hyperparameters. Add simpler, GPU-friendly option to <code>TextCategorizer</code>, and allow setting <code>exclusive_classes</code> and <code>architecture</code> arguments on initialization. Add <code>EntityRecognizer.labels</code> property. Remove document length limit during training, by implementing faster Levenshtein alignment. Use Thinc v7.0, which defaults to single-thread with fast <code>blis</code> kernel for matrix multiplication. Parallelisation should be performed at the task level, e.g. by running more containers. Models & Language Data <strong>NEW:</strong> 2-3 times faster tokenization across all languages at the same accuracy! <strong>NEW:</strong> Small accuracy improvements for parsing, tagging and NER for 6+ languages. <strong>NEW:</strong> The English and German models are now available under the MIT license. <strong>NEW:</strong> Statistical models for Greek. <strong>NEW:</strong> Alpha support for Tamil, Ukrainian and Kannada, and base language classes for Afrikaans, Bulgarian, Czech, Icelandic, Lithuanian, Latvian, Slovak, Slovenian and Albanian. Improve loading time of <code>French</code> by ~30%. Add <code>Vocab.writing_system</code> (populated via the language data) to expose settings like writing direction. CLI <strong>NEW:</strong> <code>pretrain</code> command for ULMFit/BERT/Elmo-like pretraining (see #2931). <strong>NEW:</strong> New <code>ud-train</code> command, to train and evaluate using the CoNLL 2017 shared task data. Check if model is already installed before downloading it via <code>spacy download</code>. Pass additional arguments of <code>download</code> command to <code>pip</code> to customise installation. Improve <code>train</code> command by letting <code>GoldCorpus</code> stream data, instead of loading into memory. Improve <code>init-model</code> command, including support for lexical attributes and word-vectors, using a variety of formats. This replaces the <code>spacy vocab</code> command, which is now deprecated. Add support for multi-task objectives to <code>train</code> command. Add support for data-augmentation to <code>train</code> command. Other <strong>NEW:</strong> Enhanced pattern API for rule-based <code>Matcher</code> (see #1971). <strong>NEW:</strong> <code>Doc.retokenize</code> context manager for merging and splitting tokens more efficiently. <strong>NEW:</strong> Add support for custom pipeline component factories via entry points (#2348). <strong>NEW:</strong> Implement fastText vectors with subword features. <strong>NEW:</strong> Built-in rule-based NER component to add entities based on match patterns (see #2513). <strong>NEW:</strong> Allow <code>PhraseMatcher</code> to match on token attributes other than <code>ORTH</code>, e.g. <code>LOWER</code> (for case-insensitive matching) or even <code>POS</code> or <code>TAG</code>. <strong>NEW:</strong> Replace <code>ujson</code>, <code>msgpack</code>, <code>msgpack-numpy</code>, <code>pickle</code>, <code>cloudpickle</code> and <code>dill</code> with our own package <code>srsly</code> to centralise dependencies and allow binary wheels. <strong>NEW:</strong> <code>Doc.to_json()</code> method which outputs data in spaCy's training format. This will be the only place where the format is hard-coded (see #2932). <strong>NEW:</strong> Built-in <code>EntityRuler</code> component to make it easier to build rule-based NER and combinations of statistical and rule-based systems. <strong>NEW:</strong> <code>gold.spans_from_biluo_tags</code> helper that returns <code>Span</code> objects, e.g. to overwrite the <code>doc.ents</code>. Add warnings if <code>.similarity</code> method is called with empty vectors or without word vectors. Improve rule-based <code>Matcher</code> and add <code>return_matches</code> keyword argument to <code>Matcher.pipe</code> to yield <code>(doc, matches)</code> tuples instead of only <code>Doc</code> objects, and <code>as_tuples</code> to add context to the <code>Doc</code> objects. Make stop words via <code>Token.is_stop</code> and <code>Lexeme.is_stop</code> case-insensitive. Accept <code>"TEXT"</code> as an alternative to <code>"ORTH"</code> in <code>Matcher</code> patterns. Use <code>black</code> for auto-formatting <code>.py</code> source and optimse codebase using <code>flake8</code>. You can now run <code>flake8 spacy</code> and it should return no errors or warnings. See <code>CONTRIBUTING.md</code> for details. 🔴 Bug fixes Fix issue #795: Fix behaviour of <code>Token.conjuncts.</code> Fix issue #1487: Add <code>Doc.retokenize()</code> context manager. Fix issue #1537: Make <code>Span.as_doc</code> return a copy, not a view. Fix issue #1574: Make sure stop words are available in medium and large English models. Fix issue #1585: Prevent parser from predicting unseen classes. Fix issue #1642: Replace <code>regex</code> with <code>re</code> and speed up tokenization. Fix issue #1665: Correct typos in symbol <code>Animacy_inan</code> and add <code>Animacy_nhum</code>. Fix issue #1748, #1798, #2756, #2934: Add simpler GPU-friendly option to <code>TextCategorizer</code>. Fix issue #1773: Prevent tokenizer exceptions from setting <code>POS</code> but not <code>TAG</code>. Fix issue #1782, #2343: Fix training on GPU. Fix issue #1816: Allow custom <code>Language</code> subclasses via entry points. Fix issue #1865: Correct licensing of <code>it_core_news_sm</code> model. Fix issue #1889: Make stop words case-insensitive. Fix issue #1903: Add <code>relcl</code> dependency label to symbols. Fix issue #1963: Resize <code>Doc.tensor</code> when merging spans. Fix issue #1971: Update <code>Matcher</code> engine to support regex, extension attributes and rich comparison. Fix issue #2014: Make <code>Token.pos_</code> writeable. Fix issue #2091: Fix <code>displacy</code> support for RTL languages. Fix issue #2203, #3268: Prevent bad interaction of lemmatizer and tokenizer exceptions. Fix issue #2329: Correct <code>TextCategorizer</code> and <code>GoldParse</code> API docs. Fix issue #2369: Respect pre-defined warning filters. Fix issue #2390: Support setting lexical attributes during retokenization. Fix issue #2396: Fix <code>Doc.get_lca_matrix</code>. Fix issue #2464, #3009: Fix behaviour of <code>Matcher</code>'s <code>?</code> quantifier. Fix issue #2482: Fix serialization when parser model is empty. Fix issue #2512, #2153: Fix issue with deserialization into non-empty vocab. Fix issue #2603: Improve handling of missing NER tags. Fix issue #2644: Add table explaining training metrics to docs. Fix issue #2648: Fix <code>KeyError</code> in <code>Vectors.most_similar</code>. Fix issue #2671, #2675: Fix incorrect match ID on some patterns. Fix issue #2693: Only use <code>'sentencizer'</code> as built-in sentence boundary component name. Fix issue #2728: Fix HTML escaping in <code>displacy</code> NER visualization and correct API docs. Fix issue #2740: Add ability to pass additional arguments to pipeline components. Fix issue #2754, #3028: Make <code>NORM</code> a <code>Token</code> attribute instead of a <code>Lexeme</code> attribute to allow setting context-specific norms in tokenizer exceptions. Fix issue #2769: Fix issue that'd cause segmentation fault when calling <code>EntityRecognizer.add_label</code>. Fix issue #2772: Fix bug in sentence starts for non-projective parses. Fix issue #2779: Fix handling of pre-set entities. Fix issue #2782: Make <code>like_num</code> work with prefixed numbers. Fix issue #2833: Raise better error if <code>Token</code> or <code>Span</code> are pickled. Fix issue #2838: Add <code>Retokenizer.split</code> method to split one token into several. Fix issue #2869: Make <code>doc[0].is_sent_start == True</code>. Fix issue #2870: Make it illegal for the entity recognizer to predict whitespace tokens as <code>B</code>, <code>L</code> or <code>U</code>. Fix issue #2871: Fix vectors for reserved words. Fix issue #2901: Fix issue with first call of <code>nlp</code> in Japanese (MeCab). Fix issue #2924: Make IDs of displaCy arcs more unique to avoid clashes. Fix issue #3012: Fix clobber of <code>Doc.is_tagged</code> in <code>Doc.from_array</code>. Fix issue #3027: Allow <code>Span</code> to take unicode value for <code>label</code> argument. Fix issue #3036: Support mutable default arguments in extension attributes. Fix issue #3048: Raise better errors for uninitialized pipeline components. Fix issue #3064: Allow single string attributes in <code>Doc.to_array</code>. Fix issue #3093, #3067: Set <code>vectors.name</code> correctly when exporting model via CLI. Fix issue #3112: Make sure entity types are added correctly on GPU. Fix issue #3191: Fix pickling of <code>Japanese</code>. Fix issue #3122: Correct docs of <code>Token.subtree</code> and <code>Span.subtree</code>. Fix issue #3128: Improve error handling in converters. Fix issue #3248: Fix <code>PhraseMatcher</code> pickling and make <code>__len__</code> consist
This study aimed to develop affective norms for Persian attachment related words in three phases. In the first phase, 72 students (43 female, 29 male) of Shahid Beheshti University were selected to collect attachment related words. In the second phase, 35 attachment specialists (27 female, 8 male) in Tehran, were chosen to evaluate attachment relevance of words. In the last phase, 124 students (65 female, 59 male) of Shahid Beheshti University were chosen and they rated selected words respect to emotional dimensions through Self-Assessment Manikin (Bradley & Lang, 1994) and attachment relevance. Pearson correlation coefficient and mixed ANOVA were used for analysis. After several phases of screening, 107 attachment related words were collected and we presented their emotional dimensions in a descriptive manner. Results showed that attachment relevance has a positive correlation with valence and negative correlation with arousal. There was a negative relationship between valence and arousal. Although females had a positive and negative bias in valence rating of respected words, there was no significant gender difference in arousal. The set of words that was prepared in this study can make the replication of results and scientific communication easier and provides the possibility of comparing results of different studies especially in the field of attachment.
rent ways of their classifications and views of leading translationstudies scholars on their rendering into different languages. This paper alsoexamines stylistic value of metaphors in the literary text and concludes thatstylistic equivalence is one of the most important issues in translation. Themetaphor should correspond to the original grammatically, lexically, contextuallyand culturally. The study is based on the idea of possible applications of moderncognitive linguistics to the study of metaphors in the original text and itstranslation. The combination of cognitive and traditional methods of analysis oftranslational transformations opens up new prospects for the development of the162theory and practice of translation. It can characterize more fully the general lawsof the particular translation of the text and identify the individual characteristics ofthe works of individual translators. The metaphor can be translated either literallyor via paraphrasing or using the substitution techniques. Moreover, the translatorshould have special skills and competences which will allow them to render thetext in accordance with common rules, norms and principles. It is a key problemfor metaphor translation as soon as the interpreter is obliged to connect unusualimpression and emotional features of the source text to the perception of the readerof the target text. The most numerous category in Stephen King’s works is originalmetaphors. The evidence proves that, all metaphors were rendered withoutdestruction of meaning, even if functions d
The article is devoted to the question of selecting a special (terminological) vocabulary for a comprehensive explanatory dictionary and its lexicographic normalization (standardization). Of particular relevance, this problem occurs today, when there is an active replenishment of the terminology in various fields of science, technology, production, etc. The language policy regarding the dynamic nature of the norms of the literary language consists in its codification – the establishment of rules of use in dictionaries and grammar. The main principle of graphic design of words in the dictionary of an interpretative type is the strict observance of the rules of the spelling. Independent standardization (bringing terminology to a single system) from the necessity of the first priority solving a number of problems caused by violations of the lexical general literary norm. The most typical spelling mistakes that occur in dictionaries are outlined, and the ways of overcoming them are proposed. The combination of objective updating of the lexical structure with the consciously carried on by linguists subjective updating has become a specific feature of work on organization, codification and development of the Ukrainian vocabulary. In a chain of terms on the notation of the notion “bringing something into a uniform” there is a certain hierarchical subordination: systematization – normalization – unification – codification – standardization “observance of the only stable grammatical and stylistic norms on the national language”. The question of the place of special vocabulary in general dictionaries is one of the most difficult in modern lexicography, especially today, when there is an active replenishment of the vocabulary by terms of various branches of science, technology, production. Theoretical lexicography substantiates the necessity of introducing terms into the register of explanatory dictionaries, which is the core of the terminology and at the same time a part of the literary language. Practical work on the dictionary reveals various tendencies in fixing the new words of the special vocabulary: from the aspiration of its full coverage to the reasoned selection. There are several main criteria for selecting terminological units for a general-purpose dictionary: 1) the social significance of the terms of science and technology; 2) structural features contributing to the term adaption by the system of literary language; 3) the scope of the special vocabulary. The stable tendency of productivity growth of syntactic derivation in the Ukrainian language is noted – replenishment of the terminological fund by terms of phrases, and also the criteria of the selection of such multi-component phrases, their codification and standardization are outlined. The main tendencies, which are traced in the new dictionaries, are noted. One of the most important is devoted to the foreign words entry into Ukrainian and problems connected with their spelling. Violations of spelling norms caused by wrong adaptation of borrowings in modern dictionaries are pointed out, and the negative influence of spelling mismanagement of borrowings on the functioning of the Ukrainian lexemes is emphasized. As a result, a well-considered way of building and normalizing the Ukrainian literary lexicon is the principle of a reasonable balance between tradition and the need for renewal.
On the results of the linguistic analysis of the text of the Constitution of the Russian Federation and the requirements that are imposed by legal discourse to the language of legal documents the authors reveal a set of obligatory characteristics of the language of the legal document, including definiteness, accuracy, clearness, stylistic neutrality, and language correctness. Application of these characteristics to assessing the quality of the text of Russian Constitution lead to finding some contextual gaps which might cause difficulties in understanding of some provisions in the basic law of the country. Violations of the accuracy and clarity of statements include those provisions of the Constitution in which lexical units are implemented in different meanings. In such cases, the distortion of meaning is possible at the stage of understanding (interpretation) of the verbal form. The articles of the Constitution also reveal a violation of clarity, which is accompanied by legal and linguistic uncertainty. The article provides examples of formulations that do not comply with the rules and norms of the use of language means in legal acts, as well as proposed ways to overcome these inconsistencies. It is argued that the development and application of a special linguistic methodology focused on communicative and rhetorical analysis of legally significant texts will allow more effective linguistic interpretation of normative legal acts, including the Constitution of the Russian Federation, as well as ensure uniformity of their application.
Abstract Gothic is a null subject language. The binder of an anaphor can be a null subject. Binding requires asymmetrical c-command. Possessive sein- can be a syntactic or discourse anaphor. Gothic may attest the beginning of the Germanic two-reflexive system. The simple reflexive, without silba (self), is productive in anticausative structures. Verbal prefixes alter meaning, lexical, or grammatical aspect. Ga- has numerous other functions, including definiteness and temporal completion. The nonpast participle functions as a relative clause substitute and in absolute constructions. In the absence of switch reference, infinitives are the norm with modal and control verbs and purposives after verbs of motion (otherwise + du). The accusative with a participle or infinitive can be a matrix object or embedded subject. Accusative and infinitive depends on case from the matrix verb. The infinitive is usually wisan (to be) as an expansion of a small clause. Relative clauses require the complementizer ei (that). Verbs whose complements are factual or realizable are typically in the indicative. Those that do not allow a full range of independent tenses in the complement clause, or whose complements are not realized, are only potentially realized, or deal with possible worlds or alternate states of reality, trigger a shift to the optative, which has a number of independent uses as well.
The paper raises the question of the normative representation of socio-political terms in modern explanatory dictionaries, the elaboration of a correct definition that would reflect the key characteristics of the latest lexemes. Among the special features of socio-political nomination we can emphasize the following: high instability, if to compare with other groups of vocabulary; specifics of communicative influence; names emergence before the establishment of the phenomena itself; desemantization of frequently used words and phrases. Among the most important principles of dictionary compiling we point out: 1) the team of compilers formation (specialists in the field of knowledge to be covered by dictionary, who is responsible for selecting terms for the dictionary, and for the professional correctness and objectivity of the definitions; a linguist-terminologist who recommends which term to use by agreeing it with a definition and terms of other languages; Ukrainian philologist, who edits terms and their definition according to the norms of modern Ukrainian language; a programmer who provides the implementation of the project at a computer level that makes it possible substantially to improve the work of the group); 2) the creation of a dictionary database (the use of all available branch dictionaries (translated, interpretative, encyclopaedic); the use of scientific literature and periodicals (for example, there are many terms in the periodicity that have already been distributed but have not yet found their place in dictionaries), direct work with informants – specialists of the relevant branch, who will complement the existing bank with the frequently used terms, or with the terms difficult to translate into Ukrainian, the latest in the branch, do not have a unique interpretation among specialists etc.; 3) working out definitions; 4) choosing appropriate Ukrainian term or translating into Ukrainian. The analysis of the latest lexical material in comparison with the data of the earlier dictionaries makes it possible to talk about changes in the meaning of a number of words and phrases (including Sovietisms), about the emergence or actualization of political meanings in words from outside of the political sphere, about the complication of systemic links between the words of one nest, the blurring of the semantics of some units. The absolute unification of definition is impossible because of certain restrictions, such as: 1. the semantics of some words is not limited to the certain standard interpretation, it requires the introduction of a new generic concept into the definition; 2. the word with political semantics should organically fit into the vocabulary system, which includes the vocabulary of other thematic groups. In a language, a term may belong to several different fields of knowledge or have several meanings, only one of which belongs to the group under consideration. The standard definition can only be used unchanged if it is potentially can be used in different thematic groups and covers the semantics of a word that goes beyond the scope of politics, or if the political and non-political meanings can be clearly deduced in the dictionary article. It can be argued that the problem of unifying definitions at the level of the main part of the term-fund remains open in terminology, despite the fact that possible solutions to this problem have been repeatedly violated in the scientific literature. In addition, the study and practical application of extra-intronging factors of creation and use of certain terms helps to unify the terminology system itself.
Morphology is the study of the systematic correspondence between the form and meaning of words and their components. It can be divided into three areas: inflection, derivation and compounding. Compounding consists of word formation by the combination of lexical units, which have the same autonomy as words. In order to establish norms that can be used in the assessment of patients with language disorders, this study focuses on morphological compounding in healthy subjects. The objective of this study is to characterize the performance of healthy subjects in a picture naming task that includes simple words and compounds. All the words included in the task were nouns. Simple words were manipulated for length (between one and four syllables). Compounds were manipulated for internal structure (e.g., noun-noun: chou-fleur - cauliflower; adjective-noun: ouvre-boîte – can opener) and for transparency, that is the ease with which a compound can be interpreted based on its components. A transparency judgement was obtained prior to the main experiment with another group of participants. 37 participants (18 women) aged between 45 and 75 and divided into three education levels completed the naming task. Results do not show a difference between naming accuracy of simple words compared to compounds. However, performance was influenced by the structure and transparency of compounds. Overall, compounds formed with a preposition and transparent compounds were named more accurately than other stimuli. These two factors seem related, as the preposition provides compounds with a more transparent interpretation. These findings can guide the interpretation of performance following the assessment of patients with acquired language disorders.
The transfer of research data management from one institution to another infrastructural partner is all but trivial, but can be required, for instance, when an institution faces reorganization or closure. In a case study, we describe the migration of all research data, identify the challenges we encountered, and discuss how we addressed them. It shows that the moving of research data management to another institution is a feasible, but potentially costly enterprise. Being able to demonstrate the feasibility of research data migration supports the stance of data archives that users can expect high levels of trust and reliability when it comes to data safety and sustainability.
In recent years, (retro-)digitizing paper-based files became a major undertaking for private and public archives as well as an important task in electronic mailroom applications. As first steps, the workflow usually involves batch scanning and optical character recognition (OCR) of documents. In the case of multi-page documents, the preservation of document contexts is a major requirement. To facilitate workflows involving very large amounts of paper scans, page stream segmentation (PSS) is the task to automatically separate a stream of scanned images into coherent multi-page documents. In a digitization project together with a German federal archive, we developed a novel approach for PSS based on convolutional neural networks (CNN). As a first project, we combine visual information from scanned images with semantic information from OCR-ed texts for this task. The multi-modal combination of features in a single classification architecture allows for major improvements towards optimal document separation. Further to multimodality, our PSS approach profits from transfer-learning and sequential page modeling. We achieve accuracy up to 95% on multi-page documents on our in-house dataset and up to 93% on a publicly available dataset.
The determination of the semantic extent of lexical units which refer to the nomination of an individual has led to a more precise lexical image of the speech of the Svrljig area, but also to the understanding of the cultural identity of a single linguistic community. In a word, the formation of these lexical units is primarily based on visual perception (color, size, shape and the like). At the same time, their final form was also influenced by the socio-cultural specificities of the speech community which rested on the traditional human life values of people living in a rural environment. The volume and diversity of the nominations used to denote physical characteristics reflect a clear cultural specificity of the speech community in which everything that disrupts the ideal of a natural image and is not in accordance with the social norms is deemed negative.
Issues surrounding English for Academic Purposes and its use by non-native English speakers in higher education have become increasingly significant in recent years, fueled both by increased international student mobility and increased linguistic and cultural diversity within and outside of the student body. As well as posing language-related challenges, the transfer of non-native English speakers to an English speaking foreign university also demands the negotiation of new university expectations, channeled through a new cultural environment. While academic literacies research has identified that concepts such as power, identity, and culture play a role in academic writing, the navigation of these aspects in academic writing have not yet been studied thoroughly. Consequently, this study analyzes the ways in which non-native English speaking (NNES) students articulate their navigation of power, identity, and culture within their own academic writing at a tertiary institution in Ireland. Data informing this study was gathered through questionnaires and followed by in-depth case studies of students interview responses analyzed through discourse analysis. The findings suggest that while participants generally positively reflect on their ability to negotiate academic writing through the English language, there is nonetheless a high level of conflict between dominant linguistic norms and the students’ expression of their identity and culture. These findings suggest a need to increase focus on academic literacies in tertiary institutions in order to aid the negotiation of these aspects and to increase the academic success of non-native English speakers.
Gamedesire (GD), a free online gaming website, is a rich resource for language research on Computer-Mediated Communication (CMC). GD raises a number of linguistic inquiries on written English. This paper analyzes the morphosemantic mechanisms of forming euphemistic GD usernames. A dataset of two hundred usernames has randomly been selected and tested against Warren&rsquo;s (1992) model. The study demonstrates that a plethora of GD usernames carry dysphemistic connotations that are denotatively euphemized with linguistic and paralinguistic mechanisms, including word formation, orthographic modification, borrowing and semantic innovation. Some of the dataset usernames could not be subsumed under the selected model, necessitating the addition of new devices and the development of a new rendition of the model. The study reveals that GD users employ several processes for creating their usernames, which are characterized by grammatical, lexical, phonological, graphological, and semantic deviations from language norms.
Is phonetic information encoded by distributional biases in the lexicon? Are phonotactic constraints robust enough to help a learner infer the phonic pattern of a language? Our work in progress attempts to shed light on these questions via: (i) statistical description of the distributional biases in phone sequences in a lexical database and (ii) connectionist simulation. The simulation focuses on V-to-V relations in V(C)’C(C)V phone strings since both harmony and contour constraints (the tendency for the vowels to share or avoid repetition of phonic properties, respectively) have been found in the distributional study (Albano, 2002).
Objective. The article presents the results of the empirical study of the impact of the Internet using experience on the process of the Internet texts understanding.
 Materials & Methods. Different theoretical methods and techniques were used for this purpose: deductive and inductive methods, analysis and synthesis, generalization, systematization. Empirical methods were used for this purpose: experiment (semantic and receptive), method of semantic and pragmatic interpretations, content analysis, subjective scaling procedure. Mathematical methods were used: primary statistics, checks on the normal nature of the data distribution, statistical output, taking into account statistical indicators of fashion and the scope of variation. As well as some interpretive methods that are based on specific principles of systemic, activity, cognitive and organizational approaches.
 Results. The author notes that Internet texts understanding is significantly different from the understanding of oral or written texts, since the Internet text is a pragmatically integral electronic document that constructing of conditionally completed text blocks in the form of «windows», that are opened in separate tabs of the browser, the order that depends on hyperlinks and user behavior. The peculiarities of the Internet texts include enhanced dialogue, divisibility, external informativity, reduced connectivity and comprehension, pragmatic and mostly formal integrity, conditional completeness, complicated structural, as well as hybrid and high degree of permeability, multimedia, presentation, inclination to speech game and collective authorship, saturation with neologisms, emoticons and abbreviations, fragmentation, non-compliance with linguistic norms, and the functioning of a special language etiquette.
 Conclusions. As a result of empirical research involving 716 respondents from different regions of Ukraine, it was determined that experience has the most significant effect on the understanding process at the reception stage, guiding users' activity and the accuracy of their expectations. Experience contributes to the accuracy of predicting the content of Internet texts by 15,0%. At the stage of interpretation, the adequacy and completeness of Internet texts interpretation with the accumulation of experience is improved by almost 10,0%. However, even experienced users were able to correctly interpret only a quarter of the dominant, while random – only a sixth part. Even less important is the experience and Internet activity at the stage of emotional identification, nor the assessment of comprehension, nor the coherence of emotional attitude is almost independent of the Internet using experience. With the accumulation of experience, users evaluate Internet texts more homogeneously, they are easier to realize their own attitude to the Internet texts, it becomes more consistent, however, they underestimate the complexity of Internet texts more than half of the cases, they are also inclined to share texts on the approval and critical like inexperienced readers.
Instructors at most career levels can agree there is irony in our expectations of students’ writing abilities. While we want our students to write well, and often bemoan their abilities, relatively few of us actually teach writing skills (Guilford 20012001, Robertson 20042004, Reynolds and Thompson 20112011). The real paradox, according to Reynolds and Thompson, is that while writing and associated communication skills are fundamental to most careers in ecology, “the teaching of writing is not central to science education” (Reynolds and Thompson 20112011). Few undergraduate biology courses “make explicit what most scientists agree […] that comprehension of primary scientific papers and communication of scientific concepts are two of the most important skills” students must learn (Brownell et al. 20132013). It would be easy to blame the lack of writing instruction in science courses on associated instructional challenges, but there are likely more straightforward reasons why ecologists do not emphasize writing skill development in their courses: Bean (20112011), an academic, consultant, and writing program administrator, adds to the list: In 2011, Bean published Engaging Ideas: The Professor's Guide to Integrating Writing, Critical Thinking, and Active Learning in the Classroom. In it, he productively counters these reasons for reluctance on the part of academic educators. The rest of this article will focus on some of the highlights from Engaging Ideas and associated scholarship that may be most immediately valuable to university science instructors. A rarity in a market flooded with books advising on how to become a better writer, Engaging Ideas emphasizes how to become a better writing instructor. In the early chapters, Bean outlines the research and theory underpinning current best practices in writing pedagogy and the scholarship of rhetoric and composition. The rest of Engaging Ideas is an immensely accessible, cross-curricular, writing instruction how-to book. It is also one specifically calibrated for the busy academic instructor. Consistent with current research on science writing instruction (and writing instruction more generally), Bean takes a “writing across the curriculum” approach which (1) discredits many of the myths obstructing science-related writing instruction efforts and (2) provides accessible and actionable recommendations for how we can enhance science learning through writing skill development in our own classrooms. Bean has taught and studied cross-disciplinary writing since 1976, and his tone is collegial and reasonable. He acknowledges, and is intimately familiar with, the challenges facing instructors throughout academia. His book is an explicit invitation to collaborate on the important work of developing students who can think and write critically. Bean's invitation is well-founded and timely. Results of a recent, informal, non-representative poll of readers of a well-regarded ecology blog bear this out (Merkle 20182018). Of 97 respondents, 77% (75/97) feel they should teach writing more than they do. Perhaps surprisingly, small class sizes, advanced courses, and TA availability were not the primary reasons people do not teach writing. However, other time-based reasons were as follows: Lack of time for grading time-intensive assignments (70%; 42/60) was the dominant reason, along with lack of time to support students (60%; 36/60). Lack of writing instruction (30%; 18/60) followed. The literature documents additional reasons. During a review of faculty perspectives on the importance of academic writing and instructional responsibility, Wei Zhu found, “much of what students need to write, particularly in upper division undergraduate and graduate level courses, is specifically tied to their disciplines” (2004). However, many professors perceive writing instruction responsibility as secondary to content knowledge (2004). Worse, Jackson et al. found that, among the faculty they surveyed, no academic science faculty felt they bore “any of the responsibility for developing discourse competence in students” (Jackson et al. 20062006). Essentially, Bean argues, we often pose our students reading and writing problems in a dialect they have not yet mastered (Bean 20112011), without feeling responsibility for helping them learn the dialect. Jackson et al. further highlight the issue of mismatched reading versus expected writing output. The undergraduate science courses Jackson et al. studied emphasized textbook readings, despite the primary written discourse of science taking place in academic peer-reviewed articles and the majority of assignments being article-esque laboratory reports (2006). This is a surmountable issue. Julie A. Reynolds and her colleagues hypothesize that the reluctance of STEM faculty “to incorporate writing in their courses derives largely from a lack of awareness of the research on the effectiveness of [writing teaching methods], since most published findings are in journals not regularly read by STEM faculty and the majority of studies use methods unfamiliar to most scientists” (Reynolds et al. 20122012). We can do something about this, by reading up on this literature. We can also reach across campus to our colleagues in English and Rhetoric and Composition Departments as well as the experts running our campus writing centers. Stepping out of our disciplinary box in this way can be mutually productive (e.g., Heard 20162016). Further, Melissa Kosinski-Collins and Susannah Gordon-Messer provide a case study for how to overcome misconceptions based on lack of familiarity with best practices in writing instruction. Cosinski-Collins and Gordon-Messer incorporated writing assignments of varying lengths throughout multi-session biology laboratories. They observed considerable improvement in both students’ comprehension of processes and their ability to articulate experiment design, execution, and results (20102010). Throughout Engaging Ideas, Bean uses examples akin to Cosinski-Collins and Gordon-Messer's efforts. And these three authors are not alone in asserting that “integrating writing and critical thinking components into a course can increase the amount of subject matter students actually learn.” However, Bean's particular contribution is in articulating both how these components contribute to subject matter mastery and how to get students to that point. For example, he discourages using quizzes and lecturing about assigned readings in favor of approaches such as (1) confirming that some texts are challenging, (2) assigning material not covered in class, (3) reducing the number of readings, and (4) developing and assigning reading guides. Commonly in some disciplines, reading guides are question sets that “define key terms with special disciplinary meanings, fill in needed cultural knowledge, explain the rhetorical context of the reading, illuminate the rhetorical purpose of genre conventions, and ask critical questions for students to consider as they progress through the text” (Bean 20112011). Furthermore, Bean admonishes us not to confuse the draft-like nature of examination essays with the potential polish of revised writing. In particular, excessive “worrying about spelling, grammar, and structure when you are trying to discover and clarify ideas can shut down any writer's creative energy.” To frame this point, Bean paraphrases P. Hartwell's insightful and useful grammar categories, which organize writing and speaking into the following: (1) native grammar learned in childhood, (2) academic and scientific study of grammar, (3) Standard Written English (which Bean points out is a prestige dialect), (4) parts-of-speech grammar, and (5) stylistic grammar (e.g., Strunk and White's Elements of Style). Importantly, Bean explains the documented detrimental effects of error-seeking and fixating on sentence-level errors. He provides insight into why we should not prioritize grammar. He couples this advice with a series of recommendations for how to shift grading, correction, and students’ own goals toward clarity, organization, and self-correction instead of “handbook grammar” perfection (Bean 20112011). Bean also frames the grammar issue in historical and socio-political lights, and gives particular emphasis to an argument shared by other writing instruction scholars (Bean 20112011). That is, ecology instructors (and indeed all writing-related instructors) should aim for writing skills that enable students to be actively intellectual citizens, rather than enshrining critical writing and thinking exclusively within the purview of academia (Harris 20122012). Bean asserts that it is how our language is structured, and a lack of complex reading experience, which impedes the production of writing that “is both a process of doing critical thinking and a product that communicates the results of critical thinking.” Make no mistake; Bean does not recommend teaching science writing as creative writing (although there are fine examples of how this synthesis can be powerful and productive, e.g., Skillen and Bowne 20142014). However, Bean is explicit: We can easily and inadvertently squelch students’ enthusiasm for thinking and writing about ecology by forcing them to adhere to rules and processes that we ourselves do not even follow (Bean 20112011). Bean further cautions us against assuming our simply need to be drilled on grammar, in order to execute an articulate and insightful laboratory report or research paper. He cites one study of American English speakers who, when asked to analyze new information by writing about it, “experienced partial or total linguistic collapse…Grammatical, lexical, and syntactic skills they seemed to have mastered disintegrated. Their papers were nearly incomprehensible” (Bean 20112011). “It may well be,” Bean hypothesizes, “that competence in editing and correctness is a late-developing skill that blossoms only after students begin taking pride in their writing and seeing themselves as having ideas important enough to communicate” (Bean 20112011). Other research suggests such skill development is also contingent upon students fully understanding the importance of writing for their own professional work outside of college (Guilford 20012001). Substantiating this premise, Bean cites cognitive research which indicates “that the frontal lobes of the brain, which seldom reach full maturity until age twenty-three to thirty, are needed for complex writing tasks that require writers first to wrestle with advanced, domain-specific knowledge and then to read their emerging texts from the audience's perspective” (Bean 20112011). In essence, we are asking students to do things their brains are not yet mature enough to do on their own. Throughout Engaging Ideas, Bean demonstrates both the pitfalls of leaving students to their own devices and the feasibility of us helping them meet our expectations. Shifting the focus from grammatical error to “writing as thinking” frees us and our students to focus on articulating the big ecological ideas we want them to appreciate, synthesize, and retain. Even so, Bean's lasting contribution to academic ecology instructors will not just be to emphasize students’ use of content versus merely acquiring content and using commas correctly. Engaging Ideas’ most essential contribution regards how students actually learn new information. He writes: “As cognitive research has shown […], to assimilate a new concept, learners must link it back to a structure of known material, determining how a new concept is both similar to and different from what the learner already knows. The more that unfamiliar material can be linked to the familiar ground of personal experience and already existing knowledge, the easier it is to learn” (Bean 20112011). So, our students will more readily connect to complex theories of speciation and conspecific competition if they are invited, through writing assignments, to make connections between these processes and the biological knowledge they already possess. Or better yet, students may learn to write and think more compellingly through making analogies between their own personal experiences and these foundational ecological principles. Bean suggests several ways to facilitate these past experience–new knowledge connections: (1) Students are assigned to explain the concept in writing to someone who has never heard of it (perhaps a relative or friend); (2) students are provided data and methods from a research article and directed to write the introduction and conclusion, according to disciplinary norms; (3) students are provided templates with fill-in-the-blank spaces designed to teach specific, formal writing structures expected in the discipline; (4) students are prompted to write short, summary-response essays about articles or course lectures. If organized in a sequential fashion, these assignments can also build students’ confidence in their capacity to write longer texts that meet disciplinary expectations (Bean 20112011). Herein lies Bean's book's overall utility; in it, he details how we can jumpstart developmental processes that bring students to this level of sophisticated writing. In short, Bean advocates two tools: (1) Use “scaffolding exercises that encourage students to take notes, generate ideas during prewriting, make an outline, or learn the structural features of different genres” and (2) link all new learning “to preexisting neural networks already in the learner's brain.” Do so by using “informal writing assignments aimed at helping students probe memory, connect new concepts to old networks, dismantle blocking assumptions, and understand the significance of the new concept” (Bean 20112011). Bean challenges us to pose writing problems that (1) compel students to develop research-substantiated argumentative positions and (2) require revision that results in refined expression and critical thinking. He provides a stand-out case study that could certainly be replicated in ecology classes: A business professor assigns numerous one- to two-page essays arguing for or against increasingly complex theses. These essays require research, data analysis, and reasonable argument development; revisions for higher grades are encouraged. Providing revision opportunities, Bean maintains, is of utmost necessity. He encourages requiring multiple drafts of even seemingly straightforward writing tasks (Bean 20112011). While this suggestion might seem obvious, subjects of a study exploring science faculty perceptions of their role in writing instruction reported providing few or zero revision opportunities (Zhu 20042004). However, research Bean cites indicates revision is central to students mastering discipline-specific writing (and writing more generally). Throughout Engaging Ideas, Bean provides guidance, in some cases step-by-step, for the following: (1) shifting the focus from grammar to critical thinking through writing, (2) designing writing assignments that are directly relevant to science course topics, (3) scaffolding writing assignments and revision to support student skill development, and (4) re-thinking our expectations to encompass the significant responsibility of instructors to model and mentor students toward enhanced writing skills. With Bean's book in hand, we no longer have an excuse to mutter into our cups about our students’ lack of skills. Bean and the bulk of the writing instruction literature make it clear—we can tap into a body of scholarship that will help us help our students meet mutual expectations. Engaging Ideas is an invaluable guide to doing so. Meghan Duffy, Jeremy Fox, and Kelly Kinney provided valuable comments on prior versions of this manuscript.
A test set that contains manually annotated sentences with gapping. The test set was compiled from SynTagRus (v. 2015) the dependency treebank for Russian that provides comprehensive manually-corrected morphological and syntactic annotation.
[In English below.] Este artigo traz um recorte semântico-lexical do projeto de atlas linguístico da região do Médio Tietê, no Estado de São Paulo, Brasil. Dentre as dez localidades que perfazem a rede de pontos de investigação, apresentamos aqui dados do município de Santana de Parnaíba (SdP). Trata-se de dez questões onomasiológicas respondidas por informantes de perfis sociais distintos. As perguntas foram extraídas do questionário semântico-lexical do Atlas linguístico do Brasil (CARDOSO, 2014) e aplicadas após modificações. Elas tematizam características pessoais, fenômenos e entes no mundo físico, cujas respostas mais frequentes em SdP foram ‹raiz›, ‹urubu›, ‹joão-de-barro›, ‹varejeira›, ‹terçol›, ‹catarro›, ‹tagarela›, ‹mão-de-vaca›, ‹balanço› e ‹interruptor›. Para a elicitação, aplicou-se a técnica de entrevista em três passos, que basicamente consiste em (i) perguntar, (ii) insistir e (iii) sugerir (THUN, 2000). O objetivo do presente texto é (a) sistematicamente expor e brevemente comentar as (primeiras) respostas espontâneas dos parnaibanos. Adicionalmente, apresentamos: (b) três critérios de invalidação de respostas; (c) correlações internas linguístico-sociais; (d) contrastes das normas lexicais e formas mais frequentes apuradas em SdP com resultados de seis atlas linguísticos fora do perímetro da região do Médio Tietê; e (e) observações linguísticas complementares acerca das lexias. No Brasil, a pesquisa é financiada pela Fundação de Amparo à Pesquisa do Estado de São Paulo (FAPESP), no âmbito do Projeto para a História do Português Paulista (PHPP), e, na Alemanha, pelo Serviço Alemão de Intercâmbio Acadêmico (DAAD). Palavras-chave: Médio Tietê. Dialeto Caipira. Dialetologia. Sincronia. Comparabilidade. Cite como (ABNT): FIGUEIREDO JR., S. R. Variação semântico-lexical em Santana de Parnaíba. LaborHistórico, v. 5, n. esp. 2, p. 58-82, 7 set. 2019. A semantic-lexical excerpt from our project of a linguistic atlas for the region of Médio Tietê, São Paulo State (Brazil), is presented here. The research field is formed by ten cities, each one defined as being a point of inquiry. Santana de Parnaíba (SdP), as one of them, is taken into account for this paper. Specifically, spontaneous answers to ten onomasiological questions provided by local residents with different social profiles. The questions were taken with modifications from the semantic-lexical questionnaire of the Linguistic Atlas of Brazil (CARDOSO, 2014). They are related to personal characteristics, as well as physical entities and phenomena, whose more frequent answers in SdP were ‹raiz›, ‹urubu›, ‹joão-de-barro›, ‹varejeira›, ‹terçol›, ‹catarro›, ‹tagarela›, ‹mão-de-vaca›, ‹balanço› e ‹interruptor›. For the elicitation, the three-step interview technique was employed, consisting basically in (i) asking; (ii) insisting; and (iii) suggesting (THUN, 2000). This paper aims to (a) expose systematically and offer brief comments on the (first) spontaneous answers of our SdP participants. Additionally, we (b) present three criteria for invalidating answers; (c) establish internal sociolinguistic correlations; (d) report a contrast between, on the one hand, lexical items that occurred as the most frequent or rather as norms in SdP; and, on the other hand, results of six other Brazilian linguistic atlases outside the Médio Tietê region; and finally (e) provide complementary linguistic observations on certain lexical occurrences. Within the framework of the Project for the History of Portuguese in São Paulo state (PHPP), this research receives funding from the São Paulo Research Foundation (FAPESP); in Germany, the German Academic Exchange Service (DAAD) also kindly finances this investigation. Keywords: Médio Tietê region; "Caipira" dialect; dialectology; synchrony; comparability. Cite as (APA): Figueiredo Jr., S. R. (2019). Variação semântico-lexical em Santana de Parnaíba. LaborHistórico, 5(spec. 2), 58-82.
The Argot language is one of the standard varieties of language that is formed among young people or a group of delinquents. Each social group has its own terms and expressions that must be learned in order to enter that group. Argot language is not separate from the language. Rather, it is one of its various varieties. This language represents a heterogeneous society, with each group having an impact on language. One can argue beyond this, claiming that the difference between each Argot language and the language commonly depends on the group attribute that uses this Argot language. The more different these groups are, the more they use different language forms to establish and maintain a relationship with the linguistic community. Young people form a large part of the active population of our community and the tendency towards peers and differences with adults are important features of this group. They have a particular language system that consists of the norms, values, behaviors and the core of their subculture. It is a secret code and a special communication symbol that transmits messages and creates certain rules and behaviors. Therefore, young people have their own subculture and use special vocabulary. The use of these vocabulary by young people and their influence on the youth subculture has created a certain verbal and non-verbal communication among young people that requires a scientific review. Groups and circles of friends, SMS, social networking messages, television movies, virtual social media, print media such as fictional characters, borrowing from other languages, and other forms of creativity are the means for dissemination of these words. The purpose of this paper is to identify semantic relations in the secret language within the framework of the theory of constructivism. Constructive semantics is one of the most important methods of achieving analysis using structuralism theory. In this view, the network language is one of the systematic relationships. Structural meanings refer to what the equivalent of a semantic unit is and how they are connected. The constructive semantic label is usually limited to lexical semantics. One of the most fundamental and general principles of constructivist linguistics is that languages, systems, and sub-systems or their constituent levels –grammatical, lexical, and phoneme levels– are interdependent. An important aspect of lexical semantics is how semantic relations of vocabulary are with each other; in particular, the discovery of the four classes of semantic relations, opposition, hyponymy, synonymy, and member-collection is important. Semantic Relationship is the relationship between lexical categories with other vocabulary, which confronts the speaker with the choice of different lexical categories. This term has different types. In other words, the meaning of the semantic relation is that the language has a semantic structure and the words are related in groups. Of course, these groups are formed on the basis of the semantic relations between the words. According to the tradition of studying meaning, these relations are in the semantic system of language between concepts that at first glance may seem independent, but have a close connection with each other, which is sometimes impossible to distinguish them from one another. Conceptual relationships have two types. Some of them are substitutions, and the others are synthetic, or, according to the famous statement, according to the Saussure’s attitude, is a function of substitution and conjunction. The substitution relationships between concepts arise among members of a grammatical category and with their replacement. This category of conceptual relationships, typically and not necessarily, consists of words from various grammatical categories that together create well-formedness. In order to achieve the purpose of the article, we first discuss the categories of semantic, synonymy, polysemy, opposition, hyponymy, meronymy, collocation, portion-mass, member-collection, homophony and homography. After processing and identifying these relationships in the Argot language vocabulary derived from an interview of the 15-30 year old young people in Tehran subway, as well as the Persian Dictionary of Argot of Samai’s work, the frequency of each semantic relationship was determined. The sample size is 1507 words; it has been extracted by two methods of documenting the Persian Dictionary of Argot and a researcher-made interview with a Snowball Sampling method. After collecting and deletion, the words were classified according to their nature and meaning in fourteen semantic areas. These classes include tools and objects, automobiles, moods, ethics and behavior, secret communication, the condition of organs, numbers, organs, eating and drinking, people, actions, opiates, clothing and places. The information of each word includes the semantic domain, the concept, the lexical entry and the Encyclopedia meaning. In addition to identifying semantic domains, the concept of each lexical category was also identified. Because the creators of the Argot language vocabulary use or create these words in an attempt to keep secrets hidden within a group of their inherent knowledge, the words reference may be different from these concepts. Then, the semantic relations of lexical data in each area were determined by qualitative content analysis method. The question of this research is if semantic relations exist in the secret language and what the relationship between the highest and lowest frequencies of semantic relations is. The results of the derivation of semantic analysis show that the highest and lowest lexical frequencies belong to the domains of people and clothing, respectively. Also, synonymy with the frequency of 63.05% has the highest semantic and homography with the frequency of 0.11% has the least semantic relation. The high frequency of synonymy relation in the vocabulary of this language represents the main reason for the use of Argot language; that is, to hide the meaning of these words. If these meanings are revealed to others, new terms replaces the previous words.
This article is devoted to the concept of norm definition in language communication, the establishing cases of deviation from different types of norms in political communication, and understanding of the language meanings that confirm the violation of the norm in the speeches of German politicians. The speeches of German politicians Sarah Wagenknecht, F.-M. Stritmayer were analysed on the basis of observance and violation of linguistic and communicative norms. Cases of violation of norms in the process of speech action "charges" were revealed. Language norms in German political communication are generally respected. The violation of the communicative norm is most often expressed by confrontational communicative strategy, conflict and speech action "blame", which are verbalized by means of various language tools - lexical, stylistic, syntactic, rhetorical. Speech action "blame" is explicated through negative emotional assessment, criticism of the opponents’ actions.
With the rapid development of the internet, social media has become an essential tool for getting information, and attracted a large number of people join the social media platforms because of its low cost, accessibility and amazing content. It greatly enriches our life. However, its rapid development and widespread also have provided an excellent convenience for the range of fake news, people are constantly exposed to fake news and suffer from it all the time. Fake news usually uses hyperbole to catch people’s eyes with dishonest intention. More importantly, it often misleads the reader and causes people to have wrong perceptions of society. It has the potential for negative impacts on society and individuals. Therefore, it is significative research on detecting fake news. In the paper, we built a model named SMHA-CNN (Self Multi-Head Attention-based Convolutional Neural Networks) that can judge the authenticity of news with high accuracy based only on content by using convolutional )
The MacArthur–Bates Communicative Development Inventories (CDIs) are among the most widely used evaluation tools for early language development. CDIs are filled in by the parents or caregivers of young children by indicating which of a prespecified list of words and/or sentences their child understands and/or produces. Despite the success of these instruments, their administration is time-consuming and can be of limited use in clinical settings, multilingual environments, or when parents possess low literacy skills. We present a new method through which an estimation of the full-CDI score can be obtained, by combining parental responses on a limited set of words sampled randomly from the full CDI with vocabulary information extracted from the WordBank database, sampled from age-, gender-, and language-matched participants. Real-data simulations using versions of the CDI-WS for American English, German, and Norwegian as examples revealed the high validity and reliability of the instrument, even for tests having just 25 words, effectively cutting administration time to a couple of minutes. Empirical validations with new German-speaking participants confirmed the robustness of the test.
Objective of the research is to find out structural-semantic and functionally stylistic features of the borrowed vocabulary in poetry by Lina Kostenko. Methods. Theoretical — analysis and synthesis of literature of the researched issue; practical — to obtain factual materials and based on them make conclusions regarding structural-semantic and functionally stylistic features of the borrowed vocabulary in oeuvre by Lina Kostenko. Results. Language (and therefore speech) is both stylistics, orthoepia, and compliance grammatical norms, etc., but above all it is vocabulary. And the vocabulary of the speaker is characterized first of all by its purity, synonymous and phraseological richness, normativity of pronunciation and, finally, correlation of native (Ukrainian) foreign words. Vocabulary of the modern Ukrainian language was formed in the process of its extended historical development and is a product of many epochs. Its shaping and development is closely connected with the history of Ukrainian people. Based on the factual material selected, we can conclude that Lina Kostenko most often uses words that are borrowed from Western European languages 48 % from 100 % (German, French, English, Italian); for the most part, it uses them without changing the value, but uses it as a comparison. Non-Slavic borrowings (from Latin, Ancient Greek, Turkic) make up 44 % of the studied words, the meaning of which the poet sometimes changes; uses for comparison. The borrowing from the Slavic languages (Old Slavic, Polonism, Russian) used by L. Kostenko is only 8 %. The author uses them without changing the values. The poet uses the borrowed elements of other languages with stylistic instruction with the purpose of creating the non-national diversity in the basis of Ukrainian language. Lina Kostenko also makes extensive use of terminology. After all, the potentialities of the corresponding lexical category create an artistic image, being part of a certain tropical figure, comparing and enriching the artistic color, expressive elevation of the artistic environment. Usage of vocabulary of foreign origin without its abuse and distortions, the way Lina Kostenko does, is one of the ways to enrich the vocabulary of language.
Syntactic structure of sentences obtained from Constituency Parsing is fundamental information in many Natural Language Processing tasks. However, due to the lack of available resources and the complex linguistic features of Vietnamese, the research into Constituency Parsing has not received enough attention in this language. To the best of our knowledge, the study presented in this paper is one of the first investigations to explore this task in Vietnamese. In this work, we present a Spanbased approach which focuses on representing spans through the use of contextualized pre-trained embeddings to obtain optimal parse trees for Vietnamese sentences. The conducted experiments indicate that our system achieved promising results on the VLSP Vietnamese Treebank dataset by significantly outperforming existing methods. The results of this study support the view that encoding context information into the representation of words is effective in improving the parsing performance of Vietnamese. Consequently, this idea can be generalized to apply to other tasks such as Dependency Parsing or other low-resource languages.
The human task-evoked pupillary response provides a sensitive physiological index of the intensity and online resource demands of numerous cognitive processes (e.g., memory retrieval, problem solving, or target detection). Cognitive pupillometry is a well-established technique that relies upon precise measurement of these subtle response functions. Baseline variability of pupil diameter is a complex artifact that typically necessitates mathematical correction. A methodological paradox within pupillometry is that linear and nonlinear forms of baseline scaling both remain accepted baseline correction techniques, despite yielding highly disparate results. The task-evoked pupillary response (TEPR) could potentially scale nonlinearly, similar to autonomic functions such as heart rate, in which the amplitude of an evoked response diminishes as the baseline rises. Alternatively, the TEPR could scale similarly to the cortical hemodynamic response, as a linear function that is independent of its baseline. However, the TEPR cannot scale both linearly and nonlinearly. Our aim was to adjudicate between linear and nonlinear scaling of human TEPR. We manipulated baseline pupil size by modulating the illuminance in the testing room as participants heard abrupt pure-tone transitions (Exp. 1) or visually monitored word lists (Exp. 2). Phasic pupillary responses scaled according to a linear function across all lighting (dark, mid, bright) and task (tones, words) conditions, demonstrating that the TEPR is independent of its baseline amplitude. We discuss methodological implications and identify a need to reevaluate past pupillometry studies.
The article deals with the problem of using a literary text in teaching Ukrainian as a foreign language, in particular, the need to take into account psychological factors when working with a literary text in a language class.The universal nature of fiction literature motivates and justifies the usage of such educational material. Fiction deals with universal and eternal problems: love, separation, hope, faith, struggle, death, betrayal, etc. Therefore, the literary text can excite people of different nationalities. The paper analyzes the psychological and psycholinguistic approaches to reading fiction and using it in teaching. A complete perception of a literary text is possible provided that the semantic fields of the text and semantic fields (primarily emotional) of the reader intersect. Stimulation of the emotional sphere contributes to the effective development of speech skills and facilitates the memorization of lexical units and grammatical and stylistic norms. Therefore, taking into account the interests of foreign students is a main factor for selecting literary texts and organizing effective work with them in the process of teaching Ukrainian as a foreign language. The author defines the specifics of communication “an author – a literary text in a foreign language – a reader”. It is carried out with the help of an intermediary. Such an intermediary is a teacher or an author of a book on literary reading. They select a literary text appealing to the audience, adapt it if necessary, offer comments and a task system, and draw the readers’ attention to important points that can inspire interest and cause discussion, help to establish a parallel between the literary text and the students’ life experience. Thus, the teacher acts as the moderator of the reading process helping to bring the artistic space of the literary text close to the foreign reader’s personality. The author gives examples from her own experience in teaching the Ukrainian language to foreigners using literary texts. Motivated reading and effective discussion of the texts are possible if the students are personally interested in the problems, the events, the heroes’ fate, and the author’s position.
Many of the criminal cases analysed by the Prosecution Office of the Federal District and Territories are repetitive and processing them can be streamlined by providing similar previous cases as template. We investigate the use of information retrieval techniques to enable automated identification of similar cases and evaluate if semantic search performs better than lexical search in the task of assisting legal opinion writing. As a proof of concept, syntactic indexing (TF-IDF and BM25) and semantic indexing (Latent Semantic Indexing - LSI and Latent Dirichlet Allocation - LDA) techniques were evaluated using document collections from two public prosecutors offices. In addition, we evaluate model enrichment with the use of recorded data about the cases, and also with the legal norm citations observed in documents. Baseline document collections sampled from full document collection from two public prosecutors offices were used for model evaluation utilizing Normalized Discounted Cumulated Gain (NDCG) as metric. We conclude that there is no significant performance difference between semantic and syntactic indexing techniques. In addition, we observe no significant performance gain with model enrichment. We chose the BM25 technique as more adequate because it has a good balance between performance and simplicity.
The article focuses on the formation of future Ukrainian language and literature teachers stylistic competence. Ukrainian and foreign linguodidactic researchers achievements are theoretical generalized. The attempts to determine the membership of students-philologists stylistic competence are analyzed. Based on the analysis and synthesis of scientific sources, the notion of stylistic competence of the future Ukrainian language and literature teacher as a multicomponent and dynamic personality ability is learned, which is the result of teaching the Ukrainian language Stylistics, in particular the mastering of knowledge in Stylistics, the mastery of stylistic skills and abilities, the presence of motivational and value attitudes, and behavioral norms, as well as their experience in designing stylistically completed text in oral and written forms, used I linguistic resources for the purpose, conditions and targeted communication guidelines, the implementation of research exploration and future educational activities. The components of stylistic competence according to psychological factors are determined by motivational, knowledge (cognitive), activity, value, emotional, and reflective components. According to the content it is necessary to speak about the generalistic (Stylistics general concepts and categories assimilation) and level-style phonetic-stylistic, lexical-stylistic (lexical-stylistic and phraseological-stylistic), grammatical-stylistic (word-and-stylistic, morphological-stylistic, syntactically-stylistic) components. According to the activity features stylistic competence components are characterized: linguistic, communicative, self-educational, scientific research, methodical.
This article proposes an optical measurement of movement applied to data from video recordings of facial expressions of emotion. The approach offers a way to capture motion adapted from the film industry in which markers placed on the skin of the face can be tracked with a pattern-matching algorithm. The method records and postprocesses raw facial movement data (coordinates per frame) of distinctly placed markers and is intended for use in facial expression research (e.g., microexpressions) in laboratory settings. Due to the explicit use of specifically placed, artificial markers, the procedure offers the simultaneous measurement of several emotionally relevant markers in a (psychometrically) objective and artifact-free way, even for facial regions without natural landmarks (e.g., the cheeks). In addition, the proposed procedure is fully based on open-source software and is transparent at every step of data processing. Two worked examples demonstrate the practicability of the proposed procedure: In Study 1(N= 39), the participants were instructed to show the emotions happiness, sadness, disgust, and anger, and in Study 2 (N= 113), they were asked to present both a neutral face and the emotions happiness, disgust, and fear. Study 2 involved the simultaneous tracking of 14 markers for approximately 12 min per participant with a time resolution of 33 ms. The measured facial movements corresponded closely to the assumptions of established measurement instruments (EMFACS, FACSAID, Friesen & Ekman, 1983; Ekman & Hager, 2002). In addition, the measurement was found to be very precise with sub-second, sub-pixel, and sub-millimeter accuracy.
It analyzes the literary and journalistic work of ngel de Campo (1868-1908) and the various phonetic, morphosyntactic and lexical records that appear in it both in dialogues and in the different narrative voices in order to see if its vitality continues in the Mexican dialect current or has followed other courses. The analyzed phenomena express the popular and cultured linguistic norms of the Spanish that was spoken in Mexico City at the end of the XIX century and the beginning of the XX, thanks to the record that the author makes especially of the marginalized social classes of the Mexican capital.