Papers reviewed and determined not to be word norm studies. Use the flag icon to report errors or suggest re-inclusion.
18265 papers
This paper describes the application of annotation engineering techniques for the construction of a corpus for Role and Reference Grammar (RRG). RRG is a semantics-oriented formalism for natural language syntax popular in comparative linguistics and linguistic typology, and predominantly applied for the description of non-European languages which are less-resourced in terms of natural language processing. Because of its cross-linguistic applicability and its conjoint treatment of syntax and semantics, RRG also represents a promising framework for research challenges within natural language processing. At the moment, however, these have not been explored as no RRG corpus data is publicly available. While RRG annotations cannot be easily derived from any single treebank in existence, we suggest that they can be reliably inferred from the intersection of syntactic and semantic annotations as represented by, for example, the Universal Dependencies (UD) and PropBank (PB), and we demonstrate this for the English Web Treebank, a 250,000 token corpus of various genres of English internet text. The resulting corpus is a gold corpus for future experiments in natural language processing in the sense that it is built on existing annotations which have been created manually. A technical challenge in this context is to align UD and PB annotations, to integrate them in a coherent manner, and to distribute and to combine their information on RRG constituent and operator projections. For this purpose, we describe a framework for flexible and scalable annotation engineering based on flexible, unconstrained graph transformations of sentence graphs by means of SPARQL Update.
Crowdsourcing is frequently employed to quickly and inexpensively obtain valuable linguistic annotations but is rarely used for parsing, likely due to the perceived difficulty of the task and the limited training of the available workers. This paper presents what is, to the best of our knowledge, the first published use of Mechanical Turk (or similar platform) to crowdsource parse trees. We pay Turkers to construct unlabeled dependency trees for 500 English sentences using an interactive graphical dependency tree editor, collecting 10 annotations per sentence. Despite not requiring any training, several of the more prolific workers meet or exceed 90% attachment agreement with the Penn Treebank (PTB) portion of our data, and, furthermore, for 72% of these PTB sentences, at least one Turker produces a perfect parse. Thus, we find that, supported with a simple graphical interface, people with presumably no prior experience can achieve surprisingly high degrees of accuracy on this task. To facilitate research into aggregation techniques for complex crowdsourced annotations, we publicly release our annotated corpus.
The current fundamental changes in science and education worldwide are having an impact on the use of languages, including in doctoral studies. English is dominant and this means student-researchers are often writing in a second or even third or fourth language. The languages of disciplines are also changing, as are the practices in different countries. This chapter analyses the considerable variation in the six case studies, from an Anglophone university to a multilingual university and variations in-between. Three distinctive themes emerged from the six cases: language of the discipline, the role of English, and multilingualism in research. Students explain how intensive reading and writing, often with the help of other students, is the main strategy for acquiring the language of the discipline. The role of English is evolving. Universities are striving for a comfortable balance between language diversity and unity, with English becoming the lingua franca. In many cases, research is also multilingual as student researchers collect data in a language other than the one they report in, for example. Being able to work in several languages is important and an advantage, but is at the same time very demanding. Language competences and linguistic norms and expectations are thus significant issues of which doctoral researchers are very aware, and their attitudes to English are sensitive to the dangers of the dominance of English. In sum, the question of language or multilingualism needs much more attention in policy development than is currently the case. Not only does the discourse on the internationalization of higher education neglect this topic, but also supervisors and supervisees feel more or less alone.
Progressive Supranuclear Palsy (PSP) patients present language disturbances in tasks like naming, repetition, reading, word comprehension and semantic association compared to Parkinson's disease (PD) and healthy controls (HC). In the present study we sought to validate a Screening for Aphasia in NeuroDegeneration (SAND) battery version specifically tailored on PSP patients and to describe language impairment in relation to PSP disease phenotype and cognitive status. Fifty-one PSP [23 with Richardson's syndrome (PSP-RS), 10 with predominant parkinsonism (PSP-P) and 18 with the other variant syndromes of PSP (vPSP)], 28 PD and 30 HC were enrolled in the present study. By excluding the tasks with poor acceptability (i.e., writing and picture description tasks) and increasing the items related to the remaining tasks, we showed that the PSP-tailored SAND Global Score is an acceptable, consistent and reliable tool to screen language disturbances in PSP. However, we failed to detect major di)
Psychologists have made substantial progress at developing empirically validated formal expressions of how people perceive, learn, remember, think, and know. In this article, we present an academic search engine for cognitive psychology that leverages computational expressions of human cognition (vector-space models of semantics) to represent and find articles in the psychological record. The method shows how psychological theory can be used to inform and aid the design of psychologically intuitive computer interfaces.
This study aims to explore cross-language intensification in affirmative sentences by examining the translation of standard amplifiers, words that scale upward towards an assumed norm to emphasize a quality of any entities, from Thai into English. The data comprises 602 parallel concordance lines with 17 intensifying patterns, which were drawn from a corpus of eight works of fiction in Thai and their English translations translated by qualified translators. The analysis of the data found that in the English translation, English amplifiers (e.g. very, really) were found with the highest frequency, followed by intensified lexemes and comparative and superlatives respectively. The findings suggest that the tendency to transfer standard amplifiers was through lexical (TL amplifiers, intensified lexemes, emphasizing adjectives) and syntactic means (comparatives and superlatives, exclamatory constructions, and metaphors), and that the selection was made in accordance with the context. Compared with the Thai standard amplifier maak2 ‘much-many’, the linguistic devices used in the English translations tend to reveal a stronger force of intensity. The findings can provide pedagogical implications in translations. They, for instance, can raise students’ awareness of the various linguistic forms used in transferring intensity expressed in the source text and also provide norms in translating amplifiers from Thai to English, which might be useful for students in translation programs. In addition, students may realize that if a literary work loses the expressivity of feelings or emotion, it becomes uninteresting and lacks vivacity, thus losing appeal to the TL reader.
The purpose of this work is to determine the role of the koine of the Bakhchisaray’s capital and its environs in the formation of the supra-dialect koine and the literary (standard) Crimean Tatar language in the seventeenth and eighteenth centuries. The period that we are considering in this article is quite indicative precisely in terms of the development of the Crimean Tatar literary language and its oral form – the supra-dialect koine. This time was the last stage of the fully functioning Crimean language before the Crimean Khanate lost its independence totally (1783). Phonological, lexical and grammatical norms which determined the vector of further development of the literary language of the Crimean Tatars, based mainly on the Bakhchisaray urban koine, had already crystallized in that epoch’s language. The material of this study consists of legal documents. They provide the best way to trace the processes of the formation of norms in the general Crimean supra-dialect koine, which based on the capital’s koine. Of particular value are the records of Sharia courts of the Crimean kadiys, on the one hand, and the khan’s yarliks along with letters, on the other. Both types of documents demonstrate two literary styles that were forming by different Turkic linguo-cultural traditions: that of the Golden Horde and the Crimean proper, the latter being a regional one which was influenced by the Ottoman language. The fact of lingual archaism and the mixing the phonetic, lexical and grammatical traditions of different Turkic languages in the texts of the manuscripts of official and business writing testify to the mixed character of norms in the Bakhchisaray, pre-dialect koine norms, and the norms of the literary language on the basis of the interaction of homogeneous Turkic idioms (Cumanian and Seljukian) with a small share of heterogeneous, mostly lexical, borrowings.
Meishan Zhang, Yue Zhang, Guohong Fu. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019.
We present the first open-source graphical annotation tool for combinatory categorial grammar (CCG), and the first set of detailed guidelines for syntactic annotation with CCG, for four languages: English, German, Italian, and Dutch. We also release a parallel pilot CCG treebank based on these guidelines, with 4x100 adjudicated sentences, 10K singleannotator fully corrected sentences, and 82K single-annotator partially corrected sentences.
Reference corpus for word alignment is an important resource for developing and evaluating word alignment methods. For Myanmar-English language pairs, there is no reference corpus to evaluate the word alignment tasks. Therefore, we created the guidelines for Myanmar-English word alignment annotation between two languages over contrastive learning and built the Myanmar-English reference corpus consisting of verified alignments from Myanmar ALT of the Asian Language Treebank (ALT). This reference corpus contains confident labels sure (S) and possible (P) for word alignments which are used to test for the purpose of evaluation of the word alignments tasks. We discuss the most linking ambiguities to define consistent and systematic instructions to align manual words. We evaluated the results of annotators agreement using our reference corpus in terms of alignment error rate (AER) in word alignment tasks and discuss the words relationships in terms of BLEU scores.
Aleks is a lexical database with 463 entries typical for general Slovene academic discourse. The entries include typical context examples (collocations and examples of use) taken from KAS, a corpus of Slovene academic texts, i.e. a morphosyntactically tagged synchronous and monolingual corpus, containing more than 1.5 billion words.
This article presents a new method for reducing socially desirable responding in Internet self-reports of desirable and undesirable behavior. The method is based on moving the request for honest responding, often included in the introduction to surveys, to the questioning phase of the survey. Over a quarter of Internet survey participants do not read survey instructions, and therefore, instead of asking respondents to answer honestly, they were asked whether they responded honestly. Posing the honesty message in the form of questions on honest responding draws attention to the message, increases the processing of it, and puts subsequent questions in context with the questions on honest responding. In three studies (nStudy I = 475, nStudy II = 1,015, nStudy III = 899), we tested whether presenting the questions on honest responding before questions on desirable and undesirable behavior could increase the honesty of responses, under the assumption that less attribution of desirable behavior and/or admitting to more undesirable behavior could be taken to indicate more honest responses. In all studies the participants who were presented with the questions on honest responding before questions on the target behavior produced, on average, significantly less socially desirable responses, though the effect sizes were small in all cases (Cohen’s d ranging between 0.02 and 0.28 for single items, and from 0.17 to 0.34 for sum scores). The overall findings and the possible mechanisms behind the influence of the questions concerning honest responding on subsequent questions are discussed, and suggestions are made for future research.
Most research groups studying human navigational behavior with virtual environment (VE) technology develop their own tasks and protocols. This makes it difficult to compare results between groups and to create normative data sets for any specific navigational task. Such norms, however, are prerequisites for the use of navigation assessments as diagnostic tools—for example, to support the early and differential diagnosis of atypical aging. Here we start addressing these problems by presenting and evaluating a new navigation test suite that we make freely available to other researchers (https://osf.io/mx52y/). Specifically, we designed three navigational tasks, which are adaptations of earlier published tasks used to study the effects of typical and atypical aging on navigation: a route-repetition task that can be solved using egocentric navigation strategies, and route-retracing and directional-approach tasks that both require allocentric spatial processing. Despite introducing a number of changes to the original tasks to make them look more realistic and ecologically valid, and therefore easy to explain to people unfamiliar with a VE or who have cognitive impairments, we replicated the findings from the original studies. Specifically, we found general age-related declines in navigation performance and additional specific difficulties in tasks that required allocentric processes. These findings demonstrate that our new tasks have task demands similar to those of the original tasks, and are thus suited to be used more widely.
This article introduces a package developed for R (R Core Team, 2017) for performing an integrated analysis of multiple data blocks (i.e., linked data) coming from different sources. The methods in this package combine simultaneous component analysis (SCA) with structured selection of variables. The key feature of this package is that it allows to (1) identify joint variation that is shared across all the data sources and specific variation that is associated with one or a few of the data sources and (2) flexibly estimate component matrices with predefined structures. Linked data occur in many disciplines (e.g., biomedical research, bioinformatics, chemometrics, finance, genomics, psychology, and sociology) and especially in multidisciplinary research. Hence, we expect our package to be useful in various fields.
WordNets have been used in a wide variety of applications, including in design and development of intelligent and human assisting systems.Although WordNet was initially developed as an online lexical database, (Miller, 1995 andFellbaum, 1998) later developments have inspired using WordNet database as resources in NLP applications, Language Technology developments, and as sources of structured learned materials.This paper proposes, conceptualizes, designs, and develops a voice enabled information retrieval system, facilitating WordNet knowledge presentation in a spoken format, based on a spoken query.In practice, the work converts the WordNet resource into a structured voiced based knowledge extraction system, where a spoken query is processed in a pipeline, and then extracting the relevant WordNet resources, structuring through another process pipeline, and then presented in spoken format.Thus the system facilitates a speech interface to the existing WordNet and we named the system as "Spoken WordNet".The system interacts with two interfaces, one designed and developed for Web, and the other as an App interface for smartphone.This is also a kind of restructuring the WordNet as a friendly version for visually challenged users.User can input query string in the form of spoken English sentence or word.Jaccard Similarity is calculated between the input sentence and the synset definitions.The one with highest similarity score is taken as the synset of interest among multiple available synsets.User is also prompted to choose a contextual synset, in case of ambiguities.
To qualitative researchers, social media offers a novel opportunity to harvest a massive and diverse range of content without the need for intrusive or intensive data collection procedures. However, performing a qualitative analysis across a massive social media data set is cumbersome and impractical. Instead, researchers often extract a subset of content to analyze, but a framework to facilitate this process is currently lacking. We present a four-phased framework for improving this extraction process, which blends the capacities of data science techniques to compress large data sets into smaller spaces, with the capabilities of qualitative analysis to address research questions. We demonstrate this framework by investigating the topics of Australian Twitter commentary on climate change, using quantitative (non-negative matrix inter-joint factorization; topic alignment) and qualitative (thematic analysis) techniques. Our approach is useful for researchers seeking to perform qualitative analyses of social media, or researchers wanting to supplement their quantitative work with a qualitative analysis of broader social context and meaning.
This paper presents NorNE, a manually annotated corpus of named entities which extends the annotation of the existing Norwegian Dependency Treebank. Comprising both of the official standards of written Norwegian (Bokmål and Nynorsk), the corpus contains around 600,000 tokens and annotates a rich set of entity types including persons, organizations, locations, geo-political entities, products, and events, in addition to a class corresponding to nominals derived from names. We here present details on the annotation effort, guidelines, inter-annotator agreement and an experimental analysis of the corpus using a neural sequence labeling architecture.
An important question that arises from autobiographical memory research is whether the variables that influence memory in the laboratory also drive memory for autobiographical episodes in real life. We explored this question within the context of e-mail communications and investigated the variables that influence recall for personally familiar names and temporal information in e-mails. We designed a Web-based program that analyzed each participant’s year-old sent e-mail archive and applied textual analysis algorithms to identify a set of sentences likely to be memorable. These sentences were then used as the stimuli in a cued recall task. Participants saw two sentences from their sent e-mail as a cue and attempted to recall the name of the e-mail recipient. Participants also rated the vividness of recall for the e-mail conversation and estimated the month in which they had written the e-mail. Linear mixed-effect analyses revealed that recipient name recall accuracy decreased with longer retention intervals and increased with greater frequency of contact with the recipient. Also, with longer retention intervals, participants dated e-mails as being more recent than their actual month. This telescoping error was moderately larger for e-mails with greater sentiment. These findings suggest that memory for personally familiar names and temporal information in e-mails closely follows the patterns for autobiographical memory and proper-name recall found in laboratory settings. This study introduces an innovative, Web-based experimental method for studying the cognitive processes related to autobiographical memories using ecologically valid, naturalistic communications.
The electronic lexical databases WordNets, have become essential for many computer applications, especially in linguistic research. Free French WordNet is an XML lexical database for French language based on Princeton WordNet for the English language and other multilingual resources. So far, research on Free French WordNet has focused on the construction and relevance of lexico-semantic information. However, no effort is made to facilitate the exploitation of this database under the Java language. In this context, this paper proposes our approach for the development of a new Java API based on Java Architecture for XML Binding. This Java API will make it easier for developers to exploit and use Free French WordNet to create applications for natural language processing. In order to assess the usefulness of our API, The API performance has been evaluated in the context of a Browser that we developed to extract semantic and lexical relations connecting synsets contained in this database, such as: the tree of hypernymy, the tree of hyponymy, synonyms, etc. The results showed that our API perfectly meets the needs of programmatically exploitation, exploration and consultation of this database in a Java application.
We present an ongoing project of enriching an annotation of a parallel dependency treebank, namely the Prague Czech-English Dependency Treebank, with verb-centered semantic annotation using a bilingual synonym verb class lexicon, CzEngClass. This lexicon, in turn, links the predicate occurrences in the corpus to various external lexicons, such as FrameNet, VerbNet, PropBank frame files, OntoNotes, and WordNet. We briefly describe the content of the CzEngClass synonym class lexicon and then we focus on its use for an enrichment of corpus annotation, which proceeds in two steps -automatic preprocessing and manual correction. This paper describes a first milestone of a long-term project; so far, approx. 100 CzEngClass classes, containing about 1800 different verbs each for both Czech and English, are available for such annotation. The corpus coverage at the moment is about 50%, allowing us to extract some basic statistics and discover a set of issues that appeared during the annotation process. The ultimate goal is to have a high-coverage, multilingual verbal synonym lexicon and corpora with all events annotated by such lexicon, to serve both theoretical studies in lexical semantic, translatology, corpus annotation studies etc. as well as a usable resource for training automatic semantic text processing systems for event/participant detection and linking and for general information extraction.
This article describes the first release version of a new lexicostatistical database of Northern Eurasia, which includes Europe as the most well-researched linguistic area. Unlike in other areas of the world, where databases are restricted to covering a small number of concepts as far as possible based on often sparse documentation, good lexical resources providing wide coverage of the lexicon are available even for many smaller languages in our target area. This makes it possible to attain near-completeness for a substantial number of concepts. The resulting database provides a basis for rich benchmarks that can be used to test automated methods which aim to derive new knowledge about language history in underresearched areas.
Treebanks traditionally treat punctuation marks as ordinary words, but linguists have suggested that a tree’s “true” punctuation marks are not observed (Nunberg, 1990). These latent “underlying” marks serve to delimit or separate constituents in the syntax tree. When the tree’s yield is rendered as a written sentence, a string rewriting mechanism transduces the underlying marks into “surface” marks, which are part of the observed (surface) string but should not be regarded as part of the tree. We formalize this idea in a generative model of punctuation that admits efficient dynamic programming. We train it without observing the underlying marks, by locally maximizing the incomplete data likelihood (similarly to the EM algorithm). When we use the trained model to reconstruct the tree’s underlying punctuation, the results appear plausible across 5 languages, and in particular are consistent with Nunberg’s analysis of English. We show that our generative model can be used to beat baselines on punctuation restoration. Also, our reconstruction of a sentence’s underlying punctuation lets us appropriately render the surface punctuation (via our trained underlying-to-surface mechanism) when we syntactically transform the sentence.
espanolEn el presente trabajo se delinea la aportacion de la lengua vasca a la formacion de la norma castellana durante la Edad Media y el Siglo de Oro. Para llegar a tal fin, se perfila primero el marco historico de redes sociales y linguisticas operantes en el contacto del castellano con el euskera desde los origenes de la convivencia vasco-latino-romanica hasta el siglo XVII y se tiene en cuenta, despues, la doble direccion en el contacto vasco-castellano-romanico, sobrevenido historicamente tanto en la direccion del euskera hacia el castellano como del castellano al vascuence, sin olvidar que ha afectado tambien a territorio hoy frances. EnglishThe present article studies the Basque Language contribution to the Castilian linguistic Norm in the Middle Ages and the Golden Age. For this purpose will be first designed the historical frames operating since ancient times on the linguistic contact between Basque and Romance Languages in order to explain how the direction of the Basque-Romance contact has occurred in both directions respectively, taken into account that France domain has been also concerned.
Communication is an essential part of human life. For the person with hearing and speaking disability, it is inconvenient to communicate with other people. In this paper, an end-to-end system to convert English voice to Indian Sign Language (ISL) gloss (written form of sign language) is proposed which will help deaf to communicate with others and vice versa. This system accepts English voice as an input and, converts it into the text using the speech recognition. From the recognized English text, ISL gloss is generated using the lexical database called WordNet. The focus of our work is to build a robust sign language machine translation system to convert the English text to ISL gloss using the linguistic database WordNet.
Lexical database of the lesser Sunda islands
The CHLG at a glance A parsed (= syntactically annotated) corpus of historical Low German in Penn-Treebank style. Allows for efficient, reproducible searches for a large number of morphosyntactic structures.
We recently introduced the EmojiGrid as an intuitive graphical self-report tool to measure food­evoked valence and arousal. The EmojiGrid is a Cartesian grid, labeled with facial icons (emoji) expressing different degrees of valence and arousal. The lack of verbal labels makes it a valuable, language-independent tool for cross-cultural research. Users can efficiently report their subjective ratings of both valence and arousal with a single click on the location of the grid that best represents their affective state after perceiving a stimulus. The EmojiGrid has previously been validated for the assessment of emotions evoked by food images. In this study we validated the EmojiGrid for the affective appraisal of odors. Observers (N=56, 24 males, mean age=24.3±4.6) smelled 40 randomly presented odors (27 food and 13 non­food smells), ranging from very unpleasant and arousing (e.g., feces, fish), via pleasant and calming (e.g., clove, cinnamon), to very pleasant and stimulating (e.g., peach, caramel). The odor samples consisted of felt pens, with tips that were impregnated with 4 ml of fluid odorant substance. Each pen was presented once, for about 5 seconds at 2 cm below both nostrils. The participants sniffed following a verbal command. Immediately after sniffing the pen was removed, and participants were given at least 30 s to smell fresh air. The participants reported their affective appraisal of each odor using the EmojiGrid. The resulting mean valence and arousal ratings closely agree with those from previous studies in the literature that were obtained with alternative rating tools. In addition, we find that the EmojiGrid yields the typical universal U-shaped relation between mean valence and arousal that is commonly observed for a wide range of affective sensory (visual, auditory, tactile, gustatory) stimuli. We conclude that the EmojiGrid is also a valid affective self-report tool for the assessment of odor evoked emotions.
The article is devoted to the assessment of the network community as a collective subject, as a group of interconnected and interdependent persons performing joint activities. According to the main research hypothesis, various forms of group subjectness, which determine its readiness for joint activities, are manifested in the discourse of the network community. Discourse constitutes a network community, mediates the interaction of its participants, represents ideas about the world, values, relationships, attitudes, sets patterns of behavior. A procedure is proposed for identifying discernible traces of the subjectness of a network community at various levels (lexical, semantic, content-analytical scales, etc.). The subjective structure of the network community is described based on experts' implicit representations. The revealed components of the subjectness of network communities are compared with the characteristics of the subjectness of offline social groups. It is shown that the structure of the subjectness of network communities for some components is similar to the structure of the characteristics of the subjectness of offline social groups: the discourse of the network community represents a discussion of joint activities, group norms, and values, problems of civic identity. The specificity of network communities' subjectness is revealed, which is manifested in the positive support of communication within the community, the identification and support of distinction between "us" and "them". Two models of the relationship between discursive features and the construct "subjectness" are compared: additive-cumulative and additive. The equivalence of models is established based on the discriminativeness and the level of consistency with expert evaluation by external criteria.
Theoretical background: It is a well-known fact that a name should be original and distinguish a business from the competition. It should also meet several other criteria, such as being easy to pronounce and remember, and being connected with the specifics of the company’s operations. Company names, as with any other names, are important because people react to a word the same way they react to the object this name denotes. Therefore, the length of a company name and the words it contains are important. Moreover, in the marketing nomenclature what carries the most significance is not the word itself but the so-called connotation, i.e. the direct reference to the object of the word and the entire set of features connected with it. When creating their own company names, entrepreneurs often use this aspect to better communicate with their target audience. This is enforced by both the growing competition and the expectations of consumers. In recent years, considerable innovativeness can be observed on the part of entrepreneurs who depart from traditional terms and decide on unusual names. Purpose of the article: The following article presents trends in Polish company names from the perspective of marketing efficiency on the one hand, and linguistic innovation on the other. The purpose of the article is to determine the kind of linguistic changes and their assessment from the viewpoint of communication effectiveness. Research methods: The empirical section includes the analysis of 247 names of companies that provide hairdressing services, while the theoretical section concerns the issues of creating a brand and the lexical side of it. In particular, the considerations concern the linguistic norms and marketing principles behind creating company names. Main findings: The findings indicate that the names ceased to be original but from a marketing perspective they became more effective. The names became more efficient in terms of marketing communication, for example, the words "studio" and "academy" ( studio or akademia in Polish) carry a lot of content and connote expertise, knowledge and elite. Of course, the name and surname of an owner (which were popular in the past for hairdressing companies' names) does not include such information. Of further interest is that foreign sounding words have also disappeared almost completely; in particular, the number of words from English has decreased. Therefore, by using words with a more universal meaning and domestication in the Polish language, a company evokes positive reactions and associations, which are very important in the first contact between a customer and a company.
The paper deals with studying language deviations of different types in James Joyce’s Ulysses and Finnegans Wake. Deviations in general are known to be a departure from a norm or accepted standard; in linguistics deviations are viewed as an artistic device that can be applied in different forms and at various textual levels. The author’s language deformation is analyzed as a form of deviation used for expressing the writer’s language knowledge. It is concluded that in Ulysses the destruction of the language is thoroughly thought out and multi-aimed. For instance, occasional compound units that dominate the novel imitate the style of Homer, reviving the ancient manner in contemporary language. Despite the use of conventional word-building patterns, rich semantic abundance being the basic principle of Joyce's poetics seriously complicates interpretation of the new words in the source language. The attempt is also made to systematize deviation techniques in Finnegans Wake. In particular, multilinguality is found to be the base of the lexical units created by J. Joyce. Such hybrid nonce words produce the polyphony effect and trigger the mechanism of polysemantism together with unlimited associativity of the textual material, broadening the boundaries of linguistic knowledge as a whole. Additionally, certain results of a deeper comparative analysis of the ways to translate the author’s deviations into Russian are given. The analysis of three Russian versions of Ulysses and the experimental fragmentary translation of Finnegans Wake show that there exists some regularity in the choice of translation method, particularly its dependence on the structural similarities/ differences of the source and the target languages, as well as the language levels affected by J. Joyce in the process of lingual destruction. The impossibility of complete conveyance of the semantic depth of the text and stylistic features in the target language is noted.
The article presents the results of an original research into a hypothetical dependence of translation techniques selection on the term structure in the target text. The data was obtained following the analysis of translation techniques applied to render into Ukrainian 932 English terminological word combinations related to Teaching Foreign Languages and Applied Linguistics. It was established that word combinations constitute the most numerous category in the English terminology corpus selected for the analysis. It was also found that the share of two-component terminological word combinations considerably supersedes the share of lexical units with a larger amount of words. Adjective-Noun and Noun-Noun models turned out to be the most frequent models the two-component word combinations are based on in the said sphere. In the category of the two-component word combinations, the share of adjective-headed lexical units amounts to 52%, while the Adjective-Noun model accounts for half (49%) of them. The share of the noun-headed word combinations, where the Noun-Noun model predominates, is 37%. The rest of the models have a tendency to follow the adjective-headed word combinations in their behaviour. The analysis of the correlation of translation techniques selected to render into Ukrainian the Language Pedagogy English terminological word combinations allowed to assume that the choice of the techniques is dependent on their structure. The two-component adjective-headed word combinations tend to be translated by means of calque. However, in rendering the two-component noun-headed word combinations the share of calque diminishes by half. The increase of the amount of components in a word combination is accompanied by a sharp fall in the share of calque and the simultaneous rise in the proportion of transformations, dominated by transposition, often in combination with word addition or deletion. The research results do not give any ground to assume that the selection of translation techniques in rendering English terminological word combinations related to Teaching Foreign Languages and Applied Linguistics into Ukrainian has any specific features as compared to other specialized areas, because the said results are quite similar to those observed in other spheres of human activity. Like in those spheres, calquing is used if the principles of the word combination structure in English and Ukrainian coincide, while transformations occur in case of their discrepancies. Words are added into the target text to ensure a greater degree of rendering the meaning of the term from the source text, while the reason for the word deletion is the redundancy of the word combination in the source text, i.e. the possibility to render its meaning in the target text with fewer words. Transposition is applied to meet the target language norms, and the simultaneous use of several types of transformation is explained by the desire to comply with several requirements mentioned above.
The article deals with the formation of the philosophical and psychological basis of the moral code of masters of the oriental struggle. The constructivism of the religious and philosophical thought of Buddhism and Shintoism in this process is emphasized. The term of «busido» is analysed in lexical-semantic, artistic-figurative and constructive planes. The «warrior’s path» is the key conception of the phenomenon of samurai, the core of life and life-creation, the essence and meaning of a personality’s being, the basic model and the cultural core of the masters in oriental martial arts and samurai. Busido – «warrior’s path» – the samurai code is a set of rules, recommendations and norms of the behavior of a true warrior in battle and everyday life, military philosophy, known from the ancient times. The Samurai Code of Honor proved its ability to educate true warriors, fearless defenders of the native land, brave and courageous, whom the country has been proud of for centuries. The Code has contributed to the formation of the socio-political, moral, psychological, strategic and tactical and technical potential of samurai warriors. The typology of individuals is shown: searchers of death, players with death, who deliberately neglected mortal danger. Life laid on the altar of freedom of the Motherland is the holy fate of the warrior. The existence of life in a vast existential space reveals a universal opportunity for the creation of good deeds and love to a man. The victory of this eternal and boundless world emerges on the verge of life and death. The results of research of moral qualities in sport activities of karatists in kyokushinkai style of Donetsk region are presented. The training program of karatist’s moral qualities optimization is conducted, which includes: mastering of relaxation methods, self-development of consciousness, forming of moral qualities. The program provides the creation of the special social environment, realization of methods, stimulation of psychological mechanisms of sportsmen’s moral qualities development.
The volume collects articles which discuss complexity, conventionality and creativity in the English language from perspectives as diverse as specialised discourse, language teaching and learning, language varieties, lexical creativity, stylistics, knowledge dissemination through the media and audio-visual translation. It offers a multifaceted picture of the ways in which opposing forces exerted by conventionality and creativity contribute to shaping all levels of the linguistic system. The interpretive paradigm is offered by the theory of complex systems, a rich research framework attempting to describe and explain the dynamics which emerge in the many forms of situational adaptation of natural systems. Norms and conventions are, in fact, constantly exploited and manipulated through the creative behaviour of language users. This may lead to unpredictable synchronic effects and variation and, ultimately, to diachronic innovation.
Cross-disciplinary communication is often impeded by terminological ambiguity. Hence, cross-disciplinary teams would greatly benefit from using a language technology-based tool that allows for the (at least semi-) automated resolution of ambiguous terms. Although no such tool is readily available, an interesting theoretical outline of one does exist. The main obstacle for the concrete realization of this tool is the current lack of an effective method for the automatic detection of the different meanings of ambiguous terms across different disciplinary jargons. In this paper, we set up a pilot study to experimentally assess whether the word sense induction technique of ‘context clustering’, as implemented in the software package ‘SenseClusters’, might be a solution. More specifically, given several sets of sentences coming from a cross-disciplinary corpus containing a specific ambiguous term, we verify whether this technique can classify each sentence in accordance to the meaning of the ambiguous term in that sentence. For the experiments, we first compile a corpus that represents the disciplinary jargons involved in a project on Bone Tissue Engineering. Next, we conduct two series of experiments. The first series focuses on determining appropriate SenseClusters parameter settings using manually selected test data for the ambiguous target terms ‘matrix’ and ‘model’. The second series evaluates the actual performance of SenseClusters using randomly selected test data for an extended set of target terms. We observe that SenseClusters can successfully classify sentences from a cross-disciplinary corpus according to the meaning of the ambiguous term they contain. Hence, we argue that this implementation of context clustering shows potential as a method for the automatic detection of the meanings of ambiguous terms in cross-disciplinary communication.
The article is concerned with the issue of linguistic specificity of small-sized texts, describing their text structure (also referred to as composition) as well as linguistic properties and characteristics. A cooking recipe may be defined as a specific genre of this text category. In particular, the paper aims to describe semantic structure and linguistic features of the oral cooking recipes of the Russian Germans collected during a dialectal expedition in Krasnoyarsk region, Siberia. Culinary recipes of Russian Germans may be regarded as an evidence of the preservation of their linguo-culture, traditions and ethnic identity because they reflect a number of sociocultural parameters, societal and individual values. Sociocultural parameters make it possible to regard the text of the cooking recipe as a linguocultural phenomenon. The basis of the study is a field method: digital audio recordings of the recipe texts from German dialectal informants and a comprehensive analysis thereof as compared with the recipes found in the published books. For the graphic fixation of ‘dialectal’ recipes (that is in the transcripts) which combine features of the West German and East Middle German dialects the author used the spelling close to the norms of the standard German language. As a result of the study, certain distinctive textual and linguistic features of culinary recipes were established. As far as the text and semantic structure of the oral ‘dialectal’ recipe is concerned it is normal to omit the ingredients section, information about the duration of cooking. They don’t include any clarifications or parts equivalent to subtitles and footnotes of the standard written recipes. At the phonetic, morphological, syntactic, and lexical levels culinary recipes possess typical colloquial and dialect features of the language of Russian Germans. The analysis of the recipes allowed to define the nature of changes in their linguistic components, to identify discrepancies between the recipes in the German standard language and those recorded in Krasnoyarsk region, Siberia at all linguistic levels. In addition, it became possible to discover in the dialectal material some cultural borrowings as well as to determine the linguistic impact of the other languages encountered by the informants during their life in Russia and earlier. The influence of the Russian language is exemplified among others in the mixing of elements of the Russian and German syntax and, in particular, violation of the typically German sentence framework.
The article discusses specific features of the anecdote as a speech genre, analyzes communicative-pragmatic principles of creating a comic effect in anecdotes based on the wordplay. It is noted that the concept of «anecdote», despite the fact that it is widely used in modern literary criticism and linguistics, does not have a single interpretation and a precise theoretical definition, which is explained by its genre uniqueness and complexity of a cognitive-pragmatic nature. It is emphasized that the most important part of the work of this genre is its finale, originally known to the narrator. It is the last climax phrase that contains the unexpected and unpredictable final semantic resolution that constitutes the anecdote as such. Among the features inherent in the actual anecdotal texts are: small volume, lack of authorship, reproducibility, indefinite chronotope, stereotypicality of plot schemes, relatively constant set of characters, ambivalence of the meaning of language units, intertextuality, situational functioning, etc. The dominant category of the anecdote text is minimalism, manifested in the choice of details, the number of heroes, laconic form, the volume of compositional components. It is stated that formation of the types of anecdotes took place along two lines: folk and literary. A modern anecdote, in contrast to the literary jokes of previous years, as a rule, is a speech genre, not a literary one, which determines its specificity. It is noted that anecdotes are divided into situational (subject, referential), in which the comic nature of the described situation is not related to the linguistic design, and language (linguistic), which are based on the playing out of certain linguistic phenomena. The comic effect in the latter is based on purely linguistic mechanisms and depends on the choice of the used speech means. An integral part of creating a comic effect in linguistic anecdotes is violation of certain norms, or incongruence, in the implementation of which the leading role is played by the language game. The game potential of phonetic, lexical, word-building, morphological and syntactic means, as well as precedent phenomena involved in speech works of this type are described. Particular attention is paid to punning outplaying of polysemy and various types of homonymy as one of the most popular means of creating a language joke. It is concluded that peculiarity of the game means, used for creating a humorous effect, lies in their function: they have an additional evaluative connotation, express different degrees of negative loading and take part in creating comic ambiguity in the statement.
In this paper we present a pipeline for the detection of spelling variants, i.e., different spellings that represent the same word, in non-standard texts. For example, in Middle Low German texts in and ihn (among others) are potential spellings of a single word, the personal pronoun ‘him’. Spelling variation is usually addressed by normalization, in which non-standard variants are mapped to a corresponding standard variant, e.g. the Modern German word ihn in the case of in. However, the approach to spelling variant detection presented here does not need such a reference to a standard variant and can therefore be applied to data for which a standard variant is missing. The pipeline we present first generates spelling variants for a given word using rewrite rules and surface similarity. Afterwards, the generated types are filtered. We present a new filter that works on the token level, i.e., taking the context of a word into account. Through this mechanism ambiguities on the type level can be resolved. For instance, the Middle Low German word in can not only be the personal pronoun ‘him’, but also the preposition ‘in’, and each of these has different variants. The detected spelling variants can be used in two settings for Digital Humanities research: On the one hand, they can be used to facilitate searching in non-standard texts. On the other hand, they can be used to improve the performance of natural language processing tools on the data by reducing the number of unknown words. To evaluate the utility of the pipeline in both applications, we present two evaluation settings and evaluate the pipeline on Middle Low German texts. We were able to improve the F1 score compared with previous work from \(0.39\) to \(0.52\) for the search setting and from \(0.23\) to \(0.30\) when detecting spelling variants of unknown words.
Differences in nomenclature, regarding the legal reasoning of judex facti and judex juris decisions, occur in determining the act of taking part in a criminal act. This is due to different reasoning methods. The legal consideration approach to judex facti decisions, in verifying facts as norms, is performed lexically. The way the judge's logic works is by using deductive logic and verifying the facts of the defendant's actions to normalise elements that are merely restrictive. The judex juris decisions of judges and the judex facti legal judgments understand the act of participation in corruption case by using an inductive reasoning method. Judex juris decisions examine judex facti legal considerations by determining the major premise more extensively. Judges search for the legal principles underlying norms to verify the facts of the defendant's condition. The results of the verification and the conclusion of judex juris state that the defendant's actions are proven but there are no faults. Thus, judex facti decisions are cancelled and it is decided that the defendant is free from all legal charges.
Previous research findings supporting the advantages of the go/no-go choice over the yes/no choice in lexical decision task (LDT) have suggested that the go/no-go choice might require less cognitive resources in the non-decisional processes. This study aims to test such an idea using the event-related potential method. In this study, the tasks (yes/no LDT and go/no-go LDT) and word frequency (high and low) were manipulated, and the difference between the go/no-go choice and yes/no choice were examined with BP, pN, pN1, P200, N400, and P3 components that were assumed to be closely related with the various parameters in the diffusion model. The results showed that BP, pN and pN1 amplitudes reflecting the preparation stage were not differently affected by word frequency and the task type. However, ERPs after stimulus onset showed differences. The P200 amplitudes were smaller in the go/no-go task than in the yes/no task only for low-frequency words. N400 and P3 amplitudes were only affect)
The authors of the article focus on the difficulties experienced by younger students in mastering the spelling norms of the Russian language. This is the inability to immediately distinguish in the consciousness of the signified and signifying, the inability to correctly determine the word stress and a number of others. The teacher should know the methods of formation of students ' concept of "phoneme" and the ability to recognize other phonetic units of the language. It is emphasized that the phonetic work should precede the graphic one, based on the development of the speech-motor apparatus. The authors present a description of some methods of formation of the phonetic competence, such as: exercises on the distinction between words as lexical units and as a "phonetic word", the correct syllabification, accent, modelling, awareness similarsocial functions.
The Routledge Handbook of language and culture represents a comprehensive study on the inextricable relationship between language and culture. It is structured into seven parts and 33 chapters. Part 1, Overview and historical background, by Farzad Sharifian, starts with an outline of the book and a synopsis of research on language and culture. The second chapter, John Leavitt’s Linguistic relativity: precursors and transformations discusses further the historical development of the concept of linguistic relativity, identifying different schools’ of thought views on the relation between language and culture. He also tries to demystify some misrepresentations held towards Boas, Sapir, and Whorf’ theories (pp. 24-26). Chapter 3, Ethnosyntax, by Anna Gladkova provides an overview of research on ethnosyntax, starting from the theoretical basis laid by Sapir and Whorf and investigates the differences between a narrow sense of ethnosyntax, which focuses on cultural meanings of various grammatical structures and a broader sense, which emphasises the pragmatic and cultural norms’ impact on the choice of grammatical structures. John Leavitt presents in the fourth chapter, titled Ethnosemantics, a historical account of research on meaning across cultures, introducing three traditions, i.e. ‘classical’ ethnosemantics (also referred to as ethnoscience or cognitive anthropology), Boasian cultural semantics (linguistically inspired anthropology) and Neohumboldtian comparative semantics (word-field theory, or content-oriented Linguistics). In Chapter 5, Goddard underlines the fact that ethnopragmatics investigates emic (or culture-internal) approaches to the use of different speech practices across various world languages, which accounts for the fact that there exists a connection between the cultural values or norms and the speech practices peculiar to a speech community. One of the key objectives of ethnopragmatics is to investigate ‘cultural key words’, i.e. words that encapsulate culturally construed concepts. The concept of ‘linguaculture’ (or languaculture) is tackled in Risager’s Chapter 6, Linguaculture: the language–culture nexus in transnational perspective. The author makes reference to American scholars that first introduced this notion, Paul Friedrich, who looks at language and culture as a single domain in which verbal aspects of culture are mingled with semantic meanings, and Michael Agar, for whom culture resides in language while language is loaded with culture. Risager himself brought forth a new global and transnational perspective on the concept of linguaculture, i.e. the use of language (linguistic practice) is seen as flows in people’s social networks and speech communities. These flows enhance as people migrate or learn new languages, in permanent dynamics. Lidia Tanaka’s Chapter 7, Language, gender, and culture deals with research on language, gender, and culture. According to her, the language-gender relationship has been studied by researchers from various fields, including psychology, linguistics, and anthropology, who mainly consider gender as a construct that preserves inequalities in society, with the help of language, too. Tanaka lists diachronically different approaches to language and gender, focusing on three specific ones: gender stereotyped linguistic resources, semantically, pragmatically or lexically designated language features (including register) and gender-based spoken discourse strategies (talking-time imbalances or interruptions). In Chapter 8, Language, culture, and context, Istvan Kecskes delves into the relationship between language, culture, and context from a socio-cognitive perspective. The author considers culture to be a set of shared knowledge structures that encapsulate the values, norms, and customs that the members of a society have in common. According to him, both language and context are rooted in culture and carriers of it, though reflecting culture in a different way. Language encodes past experience with different contexts, whereas context reflects present experience. The author also provides relevant examples of formulaic language that demonstrate the functioning of both types of context, within the larger interplay between language, culture, and context. Sara Miller’s Chapter 9, Language, culture, and politeness reviews traditional approaches to politeness research, with particular attention given to ‘discursive approach’ to politeness. Much along the lines of the previous chapter, Miller stresses the role of context in judgements of (im)polite language, maintaining that individuals represent active agents who challenge and negotiate cultural as well as linguistic norms in actual communicative contexts. Chapter 10, Language, culture, and interaction, by Peter Eglin focuses on language, culture and interaction from the perspective of the correspondence theory of meaning. According to him, abstracting language and culture from their current uses, as if they were not interdependent would not lead to an understanding of words’ true meaning. David Kronenfeld introduces in Chapter 11, Culture and kinship language, a review of research on culture and kinship language, starting with linguistic anthropology. He explains two formal analytic definitional systems of kinship terms: the semantic (distinctions between kin categories, i.e. father vs mother) and pragmatic (interrelations between referents of kin terms, i.e. ‘nephew’ = ‘child of a sibling’). Chapter 12, Cultural semiotics, by Peeter Torop deals with the field of ‘semiotics of culture’, which may refer either to methodological instrument, to a whole array of methods or to a sub-discipline of general semiotics. In this last respect, it investigates cultures as a form of human symbolic activity, as well as a system of cultural languages (i.e. sign systems). Language, as “the preserver of the culture’s collective experience and the reflector of its creativity” represents an essential component of cultural semiotics, being a major sign system. Nigel Armstrong, in Chapter 13, Culture and translation, tackles the interrelation between language, culture, and translation, with an emphasis on the complexities entailed by translation of culturally laden aspects. In his opinion, culture has a double-sided dimension: the anthropological sense (referring to practices and traditions which characterise a community) and a narrower sense, related to artistic endeavours. However, both sides of culture permeate language at all levels. Chapter 14, Language, culture, and identity, by Sandra Schecter tackles several approaches to research on language, culture, and identity: social anthropological (the limits at play in the social construction of differences between various groups of people), sociocultural (the interplay between an individual’s various identities, which can be both externally and internally construed, in sociocultural contexts), participatory-relational (the manner in which individuals create their social–linguistic identities). Patrick McConvell, in Chapter 15, Language and culture history: the contribution of linguistic prehistory reviews research in this field where historical linguistic evidence is exploited in the reconstruction and understanding of prehistoric cultures. He makes an account of research in linguistic prehistory, with a focus on proto- and early Indo-European cultures, on several North American language families, on Africa, Australian, and Austronesian Aboriginal languages. McConvell also underlines the importance of interdisciplinary research in this area, which greatly benefits from studies in other disciplines, such as archaeology, palaeobiology, or biological genetics. Part four starts with Ning Yu’s Chapter 16, Embodiment, culture, and language, which gives an account of theory and research on the interplay between language, culture, and body, as seen from the standpoint of Cultural Linguistics. Yu presents a survey of embodiment (in embodied cognition research) from a multidisciplinary perspective, starting with the rather universalistic Conceptual Metaphor Theory. On the other hand, Cultural Linguistics has concentrated on the role played by culture in shaping embodied language, as various cultures conceptualise body and bodily experience in different ways. Chapter 17, Culture and language processing, by Crystal Robinson and Jeanette Altarriba deals with research in the field of how culture influence language processing, in particular in the case of bilingualism and emotion, alongside language and memory. Clearly, the linguistic and cultural character of each individual’s background has to be considered as a variable in research on cognition and cognitive processing. Frank Polzenhagen and Xiaoyan Xia, in Chapter 18, Language, culture, and prototypicality bring forth a survey of prototypicality across different disciplines, including cognitive linguistics and cognitive psychology. According to them, linguistic prototypes play a critical part in social (re-)cognition, as they are socially diagnostic and function as linguistic identity markers. Moreover, individuals may develop ‘culturally blended concepts’ as a result of exposure to several systems of conceptual categorisation, especially in the case of L2 learning (language-contact or culture-contact situations). In Chapter 19, Colour language, thought, and culture, Don Dedrick investigates the issue of the colour words in different languages and how these influence cognition, a question that has been addressed by researchers from various disciplines, such as anthropology, linguistics, cognitive psychology, or neuroscience. He cannot but observe the constant debate in this respect, and he argues that it is indeed difficult to reach consensus, as colour language occasionally reveals effects of language on thought and, at other times, it is impervious to such effects. Chapter 20, Language, culture, and spatial cognition, by Penelope Brown