Papers reviewed and determined not to be word norm studies. Use the flag icon to report errors or suggest re-inclusion.
18265 papers
Berkeley FrameNet is a lexico-semantic resource for English based on the theory of frame semantics. It has been exploited in a range of natural language processing applications and has inspired the development of framenets for many languages. We present a methodological approach to the extraction and generation of a computational multilingual FrameNet-based grammar and lexicon. The approach leverages FrameNet-annotated corpora to automatically extract a set of cross-lingual semantico-syntactic valence patterns. Based on data from Berkeley FrameNet and Swedish FrameNet, the proposed approach has been implemented in Grammatical Framework (GF), a categorial grammar formalism specialized for multilingual grammars. The implementation of the grammar and lexicon is supported by the design of FrameNet, providing a frame semantic abstraction layer, an interlingual semantic application programming interface (API), over the interlingual syntactic API already provided by GF Resource Grammar Library. The evaluation of the acquired grammar and lexicon shows the feasibility of the approach. Additionally, we illustrate how the FrameNet-based grammar and lexicon are exploited in two distinct multilingual controlled natural language applications. The produced resources are available under an open source license.
<p>This paper analyzes the phenomenon of synonymy in translated texts in Russian and Tatar, with various existing and published translations of the Quran used as the main source. The primary goal of the study is to reveal the main regularities of the way synonyms function in the diachronic translations of the Quran into Russian and Tatar, as well as to follow the alterations in the vocabulary and stylistic norms of Russian and Tatar. Comparison between various translations allows shedding light on many of the peculiarities of the target language at the time the translation was completed and establishing the chronologic sequence of certain changes in the languages. The primary methods used in the study are the analysis of academic literature on the problem, consolidation of the prior research of synonyms in Russian and Tatar, studying text sources and dictionaries and comparison between the lexical units. The study shows that synonymic units found in the diachronic translations can be of varying degrees of equivalence. The most frequent in the diachronic translations of the Quran are the so-called partial synonyms, and this reflects the translators’ attempts to single out one specific lexical-semantic variant or a certain seme.</p>
The research into Bazhov’s tales has revealed frequent occurrence of the dialect and colloquial language, which is a deviation from the literary norm. Such deviations are caused by the author’s intention to preserve the local Ural language, to make the speech of the characters more colorful and authentic, and to make the text more expressive. As a result, a translator meets certain challenges – to follow the style and preserve the specific character of the original tale, to recreate it in a foreign language in such a way that the target text could have an effect on a foreign reader similar to that produced by the source text on a native reader. It is especially difficult to convey the Russian realities of the times of the Old Urals reflected in the author’s socio-cultural comments. Looking for the best solutions to these translation problems, it is necessary to rely on recommendations of reputable translation experts. The authors also give a number of recommendations for translation of such texts. The research has both theoretical and practical orientation. The research object of the article are Pavel Bazhov’s tales, the subject is the analysis of their lexical special features. The relevance of this research can be explained by the scientific interest in folklore and its language as a significant part of any culture and mentality. Bazhov’s tales have been studied by Russian philologists but not so deeply in the aspect of translation into other languages. Bazhov’s collection of tales is a good example of using a living poetic language of the Ural region, with its specific phraseology and local dialect features, to create an authentic atmosphere in literary works. This causes difficulties in translating, which can be overcome by means of pre-translation stylistic analysis of texts. Key words: genre of tales; pre-translation analysis; problem of folk tales translation; national and cultural specific character; lexical special features; translatability.
Mondzish (Mangish) lexical database, including transcriptions of my audio recordings collected in China in from 2012-2015.
The paper contains the research of noun-compounds from modern Tibetan corpus with the use of a relational lexical database. The lexical database represents a consistent classification of meanings of Tibetan lexical units with different relations between them. The paper describes the structure of the database; principles of process work with Tibetan compounds; recognized types of compounds semantic structure.
The purpose of this paper is to conduct word level and key word analysis of the dialogue in Korean textbooks using WordSmith tools. For this study, we compiled 12 Korean textbooks used in two major Korean language institutes(i.e. Yonsei and Korea). We measure the frequency of token and type of words found in the texts using the WordList function. We also use the KeyWords function for extracting the lexical saliency of texts. Key words are those whose frequency is unusually high in comparison with some norm. We extend the keyness concept to clusters which are found repeatedly near each other. Finding these key words and key clusters could help to verify the comparison of dialogues in Korean textbooks.
The all-Russian lexeme время and its derivatives in the dialects of the Russian language are considered. The author believes that the semantic volume of the word, well-known to the literary language, and its dialectal counterparts may not be identical due to different discursive conditions generated by the culture. The relevance of the study is determined by the increased attention in modern linguistics to the problems of reflection of traditional culture in the language, as well as the issues of diachronic description of the vocabulary of the Russian language and the history of particular words. Based on the analysis of lexical-semantic variants of the word время and meanings of the words with - врем - root in Russian dialects the understanding of time in traditional culture is refined. It is reported that the word время in the traditional sense names not the whole period of human life from birth to death, but only the period of biological maturity associated with the ability to procreate. It is proved that the period of maturity in the people’s culture and language is assessed as a period of prosperity, which becomes the basis for the submission of the norm in human life and a landmark in the awareness of the life space of a person. It is established that the semantics of the Russian word время has accumulated the most ancient etymological meanings of the words год and пора.
Developing a thesis is a process that demands time and dedication by the student since it is necessary to comply with conditions and norms established by institutional guides of the universities. This work describes a computational web tool that allows to evaluate the conclusion section of a thesis, focusing on three aspects: "Coverage", i.e. the connection between the general objective and the conclusion, "Opinion", value judgments about the concluded research, and "Speculation", i.e. evidence of a reflection on future work. This tool is incorporated into TURET 2.0. With the release of this updated version, TURET becomes a tool that the student can employ to analyze his/her thesis draft under acceptable parameters before submitting it to his/her adviser for further review. TURET will provide the analysis of the lexical richness and the analysis of key features of a conclusion section. We present details about the performance and the interfaces of the computational tool developed.
This study introduced the specific purpose translation teaching to Indonesian undergraduate students at Universitas Al-Azhar Medan, Indonesia. The courses were attended by the Business and Economics students who are new to translation. As parallel corpus, bilingual contract documents in Indonesian and English were chosen to help the students to grasp the conventions and norms in both languages. Dealing with difficulties in teaching specific purpose translation, the procedures and sequence analysis were conducted. The procedures consist of preliminary test, introduction to translation strategy, discussion by compare two translation text, and final test. The sequence analysis were conducted on discussion. This analysis based on semantic, lexical and syntactical aspect. The analysis shows that contract terms were characterized by nominalization, passive voice, sentence length and complexity, impersonality, binominal and multinominal expressions, unusual word order, one syllable and phrase equivalence. The students also recognizes the archaicsm, repetition and redundancy, synonymy and redundancy and absorption of foreign words. Based on the commentaries of the students, the use of parallel corpus as a tool in translation exercise has improving their ability in translating and drafting bilingual contract documents. In the end of course, 24 students completing the course and 19 (80%) of them are ready to attend the advance course.
In this paper we discuss the characteristics that are habitually assigned to collocations: binarism, belonging to the norm and arbitrariness and show that these properties do not in fact define them. 1) The binary relationship between the base and the collocate is contradicted by the possibility of combining the same collocate with a number of bases. 2) The assumption that collocations belong to the norm is refuted by the existence of semantic paradigms formed by bases which combine with the same lexical unit: the collocate. 3) The very existence of semantic paradigms is not an arbitrary phenomenon, but rather, a consequence of the meaning of the units that constitute them, the bases, one of whose semantic features forms part of the meaning of the collocate. Thus, we can conclude that collocations are radial syntagmatic structures in which one semantic feature of the collocate determines its combination with a lexical class of units, bases, which share the same feature.
This dissertation investigates linguistic and metalinguistic practices in everyday Twitter discourse in relation to aspects of speech and writing. The overarching aim is to investigate how the spoken–written interface is reconfigured in the digital writing spaces of social media. The dissertation comprises four empirical case studies and six chapters. The first study investigates communicative functions of hashtags in a speech act pragmatic framework, focalizing tagging practices that not only mark topics or organize hypertextual interaction, but rather have more specific locally meaningful functions. Two studies investigate reported speech in tweets, focusing on quotatives typically associated with informal conversational interaction (e.g., BE like). The studies identify strategies by which Twitter users animate (Tannen, 2007) speech reports. Further, one of the studies explores how such animating practices are afforded (Hutchby, 2001). Lexically, orthographically, and with images, but primarily through typography, users make voice, gesture, and stance present in their tweets, digitally re-embodying the rich nonverbal expressivity of animation in talk. Finally, a study investigates notions of talk-like tweeting from an emic perspective, showing users' negotiations of how tweets can and should correspond to speech in relation to social identity, linguistic competence, and personal authenticity. Six chapters situate and synthesize the case studies in an expanded theoretical framework. Together, the studies show how Twitter's speech–writing hybridity extends beyond a mix of linguistic features, and challenges a traditional idea of writing as a mere representation of speech. Talk-like tweeting remediates (Bolter & Grusin, 2000) presence and embodiment, forgoing the abstraction of phonetic print literacy for nonverbal expressivity and an embodied written surface. Twitter talk is shown not simply to substitute literacy norms for oral norms, but to complicate and reconfigure these norms. Talk-like tweeting makes manifest an ongoing cultural renegotiation of the meanings of speech and writing in the era of digital social media.
Text-setting, the arrangement of language to music, is a common source of evidence in the debate over the relevance of the syllable in Japanese prosody (e.g., Labrune 2012). Although Japanese text-setting is typically treated as mora-based, the present corpus analysis reveals that syllable-based text-setting is pervasive in Japanese. Two studies presented here compare native Japanese songs with those translated into Japanese. The results demonstrate use of syllabic settings throughout the corpora and across the lexical strata of Japanese. Syllabic settings are shown to arise with greater likelihood in response to pressures imposed by restrictive translation contexts, information density mismatch, and knowledge of correspondence to English loans. We argue that, given the viability of syllabic text-setting in Japanese, moraic text-setting is a stylistic norm of Japanese music that is shifting over time, rather than evidence of a lack of syllable structure in the language’s prosodic system.
A controlled natural language (CNL) is based on a natural language but includes restrictions on vocabulary, grammar, and/or semantics, in order to reduce or eliminate ambiguity and complexity.
This paper presents a new method with which to assist individuals with no background in linguistics to create monolingual dictionaries such as those used by the morphological analysers of many natural language processing applications. The involvement of non-expert users is especially critical for under-resourced languages which either lack or cannot afford the recruitment of a skilled workforce. Adding a word to a morphological dictionary usually requires identifying its stem along with the inflection paradigm that can be used in order to generate all the word forms of the new entry. Our method works under the assumption that the average speakers of a language can successfully answer the polar question “is x a valid form of the word w to be inserted?”, where x represents tentative alternative (inflected) forms of the new word w. The experiments show that with a small number of polar questions the correct stem and paradigm can be obtained from non-experts with high success rates. We study the impact of different heuristic and probabilistic approaches on the actual number of questions.
This paper presents an overview of studies on automated hand gesture analysis, which is mainly concerned with recognition and segmentation issues related to functional types and gesture phases. The issues selected for discussion have been arranged in a way that takes account of problems within the Theory of Gestures that each study seeks to address. Their principal computational factors that were involved in conducting the analysis of automated hand gesture have been examined, and an analysis of open research issues has been carried out for each application dealt with in the studies.
One particular problem in large vocabulary continuous speech recognition for low-resourced languages is finding relevant training data for the statistical language models. Large amount of data is required, because models should estimate the probability for all possible word sequences. For Finnish, Estonian and the other fenno-ugric languages a special problem with the data is the huge amount of different word forms that are common in normal speech. The same problem exists also in other language technology applications such as machine translation, information retrieval, and in some extent also in other morphologically rich languages. In this paper we present methods and evaluations in four recent language modeling topics: selecting conversational data from the Internet, adapting models for foreign words, multi-domain and adapted neural network language modeling, and decoding with subword units. Our evaluations show that the same methods work in more than one language and that they scale down to smaller data resources.
The present event-related potential (ERP) study investigated for the first time whether children with early-onset social anxiety disorder (SAD) process affective facial expressions of varying intensities differently than non-anxious controls. Participants were 15 SAD patients and 15 non-anxious controls (mean age of 9 years). They were presented with schematic faces displaying anger and happiness at four intensity levels (25%, 50%, 75%, and 100%), as well as with neutral faces. ERPs in early and later time windows (P100, N170, late positivity [LP]), as well as affective ratings (valence and arousal) for the faces, were recorded. SAD patients rated the faces as generally more arousing, regardless of the type of emotion and intensity. Moreover, they displayed enhanced right-parietal LP (350-650 ms). Both arousal ratings and LP reflect stimulus intensity. Therefore, this study provides first evidence of an intensity amplification bias in pediatric SAD during facial affect processing.
The past half-century has witnessed remarkable growth in the study of language variation, and it has now become a highly productive subfield of research in sociolinguistics. Variability is everywhere in language, from the unique details in each production of a sound or sign to the auditory or visual processing of the linguistic signal. All languages that we can observe today show variation; what is more, they vary in identical ways, namely geographically and socially. It's no secret that languages like English are full of variation. So, the aim of the article is to detect the reasons of variation and to uncover rates of usage of different free variations for a given set of lexical items. The research work is carried out by using the descriptive, comparative methods by subjecting to analysis the specific language materials. The discovery of law of variation became a starting point for the evolution of linguistics. The problem of search of variation facts and its role in the functioning of language system concerns many specialists from the outset. The scope of the investigation was to set up a system out of chaos of phenomena. Currently, the fact of conditionality of variation by system relations existing in the language is considered to be established.
The development of a normalized morpho-syntactic Arabic lexicon is not an easy task. In fact, many norms allow the structuration and representation of lexical data. The adoption of a stable standard will guarantee the interoperability and interchangeability of lexical resources. Still, research work that deals with normalization for Arabic lexical resources is not well developed yet, especially for some standards such as the TEI (Text Encoding Initiative). In this context, we aim at creating an Arabic lexicon editor with a constraint checker based on both the ISO standard LMF (Lexical Markup Framework) and the TEI guidelines. To develop this editor, we use a linguistic approach composed of several steps. The editor's prototype named ALIF can guarantee the construction of two types of output lexicon files: one in LMF and the other in TEI. The evaluation of this system is based upon a lexical database that contains all the derived and inflected forms generated from a lexicon of 10 000 canonical verbs. The results obtained were encouraging despite some flaws related to exceptional cases of difficult words.
INTRODUCTION: Acute postoperative pain following major surgery has come under increasing scrutiny as a harbinger for the development of potentially debilitating chronic postsurgical pain (CPSP), which is defined as the new onset of pain or intensification of presurgical pain persisting at least two to three months following surgery. We examined the prevalence of and risk factors associated with CPSP among women undergoing breast reconstruction. METHODS: Women ≥18 years undergoing immediate or delayed post-mastectomy breast reconstruction were recruited as part of the NCI-funded Mastectomy Reconstruction Outcomes Consortium Study, a prospective cohort study including 10 centers across the U.S. and Canada. In the current analysis, women were assessed preoperatively and at two-years postoperatively for pain experience (NPRS, MPQ-SF), severity of anxiety (GAD-7), and depression (PHQ-9), relevant medical/surgical variables, and reconstructive procedure type. Mixed-effects regression modeling was used to assess the relationship between patient-specific factors as the independent variables and two-year postoperative pain. RESULTS: Of the 1,996 patients included in the analysis, 92.7% (n=1851) underwent immediate reconstruction, with the majority (n=1263, 63.3%) choosing tissue expander-implant (TE/I) reconstruction. There was no significant difference between women reporting moderate or severe pain at two-year follow-up compared to preoperatively (11 vs. 10%, p=0.083). Regression modeling indicated that both preoperative pain (p<0.001) and depression severity (p<0.004) were related to CPSP. Autologous flap reconstruction was associated with more severe CPSP than TE/I on the MPQ-Sensory and Affective ratings. BMI, bilateral reconstruction, axillary lymph node dissection, and adjuvant radiation and chemotherapy were associated with CPSP for at least one pain measure. CONCLUSION: Women undergoing autologous reconstruction had significantly greater pain levels on the MPQ compared to TE/I patients two-years postoperatively, an important point to consider during preoperative counseling. The presence of preoperative pain was also a risk factor for CPSP. While multiple studies in recent years have focused on the widespread and under-reported prevalence of post-mastectomy CPSP, only 10% of the sample had moderate-to-severe pain at two years. Therefore, approximately 90% were either pain-free or living with a level of pain that would not be expected to interfere with daily function. Overall, our data suggest that CPSP for this cohort may be of less clinical concern than previously described, and reports of persistent pain after breast reconstruction may not necessarily reflect surgery-induced pain.
A set of Python scripts that convert function-head style encodings in dependency treebanks in a content-head style encoding (as used in the UD treebanks) and vice versa (for adpositions, copula and coordination). For more information, see (Rehbein, Steen, Do & Frank 2017).
We demonstrate the current state of INESS, the Infrastructure for the Exploration of Syntax and Semantics.INESS is making treebanks more accessible to the R&D community.Recent work includes the hosting of more treebanks, now covering more than fifty languages.Special attention is paid to NorGramBank, a large treebank for Norwegian, and to the inclusion of the Universal Dependency treebanks, all of which are interactively searchable with INESS Search.
Full text discourse parsing relies on texts comprehensively annotated with discourse relations. To this end, we address a significant gap in the inter-sentential discourse relations annotated in the Penn Discourse Treebank (PDTB), namely the class of cross-paragraph implicit relations, which account for 30% of inter-sentential relations in the corpus. We present our annotation study to explore the incidence rate of adjacent vs. non-adjacent implicit relations in cross-paragraph contexts, and the relative degree of difficulty in annotating them. Our experiments show a high incidence of non-adjacent relations that are difficult to annotate reliably, suggesting the practicality of backing off from their annotation to reduce noise for corpusbased studies. Our resulting guidelines follow the PDTB adjacency constraint for implicits while employing an underspecified representation of non-adjacent implicits, and yield 62% inter-annotator agreement on this task.
The growing emergence of cell phones and caller ID have reduced response rates among telephone polling, causing some concern among those who conduct public opinion polls in politics. That issue has led some to consider online surveys as either an alternative or at least a supplemental technique for gauging political opinions. This study sought to test this concept by conducting two identical surveys – on with live telephone interviews and one with an online survey. The results indicated that the data from the two surveys were not identical. Hillary Clinton scored higher on image ratings with the online survey, and the data for the voter optimism were also different. One possible explanation is that the online surveys are less susceptible to errors caused by a socially desired response pattern. That offers the potential for more accuracy from online surveys.
International audience
Today, German language islands in Russia and Brazil are on the way to language shift. On this way, the varieties of these communities display certain features of decomposition and simplification in terms of morphology. Regular and irregular morphology, however, are developing differently: while case reduction is the main characteristic of regular noun inflection, in personal pronouns case distinctions are maintained. Results are presented from a research project about language change in case morphology of German language islands with 125 speakers living in close contact to the majority populations in Brazil and Ruguage obsolescence as from language emergence which has been the subject of linguistic research in the past. Through its comparative perspective, it seems possible to accoussia. The core idea of the project is the assumption that we can learn as well from lannt for internally or externally induced linguistic change. Language decay is apparently not just disorder, not amorphous, but somehow structured. Certain lexical classes are more subject to reduction than others, and some residual features retain morphological “core” functions (in terms of case semantics). Language change is accelerated in times of blurring sociolinguistic differences and fading linguistic norms as an implication of losing ethnic boundaries. The recent co-officialization of minority languages in Brazil might slow down these processes. In a transcultural approach, teaching of Pomeranian as minority language (alongside the national language) could stabilize the local linguistic community, building a bridge to the High German standard language, and even to English as a lingua franca of international communication.
BACKGROUND: Apart from a progressive decline of motor functions, Parkinson's disease (PD) is also characterized by non-motor symptoms, including disturbed processing of emotions. This study aims at assessing emotional processing and its neurobiological correlates in PD with the focus on how medicated Parkinson patients may achieve normal emotional responsiveness despite basal ganglia dysfunction. METHODS: Nineteen medicated patients with mild to moderate PD (without dementia or depression) and 19 matched healthy controls passively viewed positive, negative, and neutral pictures in an event-related blood oxygen level-dependent functional magnetic resonance imaging study (BOLD-fMRI). Individual subjective ratings of valence and arousal levels for these pictures were obtained right after the scanning. RESULTS: Parkinson patients showed similar valence and arousal ratings as controls, denoting intact emotional processing at the behavioral level. Yet, Parkinson patients showed decreased bilateral putaminal activation and increased activation in the right dorsomedial prefrontal cortex (PFC), compared to controls, both most pronounced for highly arousing emotional stimuli. CONCLUSIONS: Our findings revealed for the first time a possible compensatory neural mechanism in Parkinson patients during emotional processing. The increased medial PFC activity may have modulated emotional responsiveness in patients via top-down cognitive control, therewith restoring emotional processing at the behavioral level, despite striatal dysfunction. These results may impact upon current treatment strategies of affective disorders in PD as patients may benefit from this intact or even compensatory influence of prefrontal areas when therapeutic strategies are applied that rely on cognitive control to modulate disturbed processing of emotions.
The amount of data that is available for research grows rapidly, yet technology to efficiently interpret and excavate these data lags behind. For instance, when using large treebanks for linguistic research, the speed of a query leaves much to be desired. GrETEL Indexing, or GrInding, tackles this issue. The idea behind GrInding is to make the search space as small as possible before actually starting the treebank search, by pre-processing the treebank at hand. We recursively divide the treebank into smaller parts, called subtree-banks, which are then converted into database files. All subtree-banks are organized according to their linguistic dependency pattern, and labeled as such. Additionally, general patterns are linked to more specific ones. By doing so, we create millions of databases, and given a linguistic structure we know in which databases that structure can occur, leading up to a significant efficiency boost. We present the results of a benchmark experiment, testing the effect of the GrInding procedure on the SoNaR-500 treebank.
While dependency parsers reach very high overall accuracy, some dependency relations are much harder than others. In particular, dependency parsers perform poorly in coordination construction (i.e., correctly attaching the conj relation). We extend a state-of-the-art dependency parser with conjunction-specific features, focusing on the similarity between the conjuncts head words. Training the extended parser yields an improvement in conj attachment as well as in overall dependency parsing accuracy on the Stanford dependency conversion of the Penn TreeBank.
In recent years, the research on Treebank has made great progress. However, the application of the Treebank research in international Chinese teaching is not very satisfactory. In view of international Chinese teaching, this paper constructs a diagrammatic Treebank based on the Li Jinxi's Sentence-based Grammar. With the constructing of the diagrammatic Treebank, we have made an exploration in word interpretation based on context, accurate example sentences recommendations based on word senses, words exercise based on dynamic word patterns, and specific grammar point example sentences recommendation.