Papers reviewed and determined not to be word norm studies. Use the flag icon to report errors or suggest re-inclusion.
18265 papers
The study of biological point-light displays (PLDs) has fascinated researchers for more than 40 years. However, the mechanisms underlying PLD perception remain unclear, partly due to difficulties with precisely controlling and transforming PLD sequences. Furthermore, little agreement exists regarding how transformations are performed. This article introduces a new free-access program called PLAViMoP (Point-Light Display Visualization and Modification Platform) and presents the algorithms for PLD transformations actually included in the software. PLAViMoP fulfills two objectives. First, it standardizes and makes clear many classical spatial and kinematic transformations described in the PLD literature. Furthermore, given its optimized interface, PLAViMOP makes these transformations easy and fast to achieve. Overall, PLAViMoP could directly help scientists avoid technical difficulties and make possible the use of PLDs for nonacademic applications.
Engineering activities often produce considerable documentation as a by-product of the development process. Due to their complexity, technical analysts can benefit from text processing techniques able to identify concepts of interest and analyze deficiencies of the documents in an automated fashion. In practice, text sentences from the documentation are usually transformed to a vector space model, which is suitable for traditional machine learning classifiers. However, such transformations suffer from problems of synonyms and ambiguity that cause classification mistakes. For alleviating these problems, there has been a growing interest in the semantic enrichment of text. Unfortunately, using general-purpose thesaurus and encyclopedias to enrich technical documents belonging to a given domain (e.g. requirements engineering) often introduces noise and does not improve classification. In this work, we aim at boosting text classification by exploiting information about semantic roles. We have explored this approach when building a multi-label classifier for identifying special concepts, called domain actions, in textual software requirements. After evaluating various combinations of semantic roles and text classification algorithms, we found that this kind of semantically-enriched data leads to improvements of up to 18% in both precision and recall, when compared to non-enriched data. Our enrichment strategy based on semantic roles also allowed classifiers to reach acceptable accuracy levels with small training sets. Moreover, semantic roles outperformed Wikipedia- and WordNET-based enrichments, which failed to boost requirements classification with several techniques. These results drove the development of two requirements tools, which we successfully applied in the processing of textual use cases.
The article considers variants of paronyms compatibility and some mistakes of Turkmen students in the use of paronyms – words closed in pronunciation but not identical, and sometimes different in meaning. The vocabulary of foreign students entering the university does not always correspond to the needs of their language practice, often there is no knowledge of the paronyms necessary for it. The study of various paronyms combining variants with other words, semantic connections are as necessary as knowledge of grammatical rules and spelling norms. The correct using of the word in speech suggests, firstly, knowledge of the word structure and the lexical-semantic variants of the polysemy; secondly, analysis of the word-building composition with the individual morphemes meaning identification; thirdly, the ability to choose the word needed for a given context, for which it is necessary to determine its place in the lexical-semantic system, i.e. to find its connection with other words, to identify general and differential features in a given lexical group making up a certain semantic unity. As a result of observations of Turkmen students“ oral and written speech, the main reasons for the erroneous paronyms substitution are revealed, involving roots consonance. The same root paronyms are also confused because of inaccurate understanding of the prefixes, suffixes meaning and difference they bring to the word meaning. Considering the same root-words, there are many problems due to the fact that they are heterogeneous in the semantic and word-formation aspect. Paronyms like the same-root words, similar to synonyms. They have a number of common features which leads to the similarity of these two language units, and finally to the error occurrence while using paronyms in speech. The article also contains examples of training exercises for speech skills developing in the paronyms using at Russian language classes in the Turkmen audience.
Although Modern Standard Arabic is taught in schools and used in written communication and TV/radio broadcasts, all informal communication is typically carried out in dialectal Arabic. In this work, we focus on the design of speech tools and resources required for the development of an Automatic Speech Recognition system for the Tunisian dialect. The development of such a system faces the challenges of the lack of annotated resources and tools, apart from the lack of standardization at all linguistic levels (phonological, morphological, syntactic and lexical) together with the mispronunciation dictionary needed for ASR development. In this paper, we present a historical overview of the Tunisian dialect and its linguistic characteristics. We also describe and evaluate our rule-based phonetic tool. Next, we go deeper into the details of Tunisian dialect corpus creation. This corpus is finally approved and used to build the first ASR system for Tunisian dialect with a Word Error Rate of 22.6%.
Modern data-driven spoken language systems (SLS) require manual semantic annotation for training spoken language understanding parsers. Multilingual porting of SLS demands significant manual effort and language resources, as this manual annotation has to be replicated. Crowdsourcing is an accessible and cost-effective alternative to traditional methods of collecting and annotating data. The application of crowdsourcing to simple tasks has been well investigated. However, complex tasks, like cross-language semantic annotation transfer, may generate low judgment agreement and/or poor performance. The most serious issue in cross-language porting is the absence of reference annotations in the target language; thus, crowd quality control and the evaluation of the collected annotations is difficult. In this paper we investigate targeted crowdsourcing for semantic annotation transfer that delegates to crowds a complex task such as segmenting and labeling of concepts taken from a domain ontology; and evaluation using source language annotation. To test the applicability and effectiveness of the crowdsourced annotation transfer we have considered the case of close and distant language pairs: Italian–Spanish and Italian–Greek. The corpora annotated via crowdsourcing are evaluated against source and target language expert annotations. We demonstrate that the two evaluation references (source and target) highly correlate with each other; thus, drastically reduce the need for the target language reference annotations.
Bilingual termbanks are important for many natural language processing applications, especially in translation workflows in industrial settings. In this paper, we apply a log-likelihood comparison method to extract monolingual terminology from the source and target sides of a parallel corpus. The initial candidate terminology list is prepared by taking all arbitrary n-gram word sequences from the corpus. Then, a well-known statistical measure (the Dice coefficient) is employed in order to remove any multi-word terms with weak associations from the candidate term list. Thereafter, the log-likelihood comparison method is applied to rank the phrasal candidate term list. Then, using a phrase-based statistical machine translation model, we create a bilingual terminology with the extracted monolingual term lists. We integrate an external knowledge source—the Wikipedia cross-language link databases—into the terminology extraction (TE) model to assist two processes: (a) the ranking of the extracted terminology list, and (b) the selection of appropriate target terms for a source term. First, we report the performance of our monolingual TE model compared to a number of the state-of-the-art TE models on English-to-Turkish and English-to-Hindi data sets. Then, we evaluate our novel bilingual TE model on an English-to-Turkish data set, and report the automatic evaluation results. We also manually evaluate our novel TE model on English-to-Spanish and English-to-Hindi data sets, and observe excellent performance for all domains.
To push the state of the art in text mining applications, research in natural language processing has increasingly been investigating automatic irony detection, but manually annotated irony corpora are scarce. We present the construction of a manually annotated irony corpus based on a fine-grained annotation scheme that allows for identification of different types of irony. We conduct a series of binary classification experiments for automatic irony recognition using a support vector machine (SVM) that exploits a varied feature set and compare this method to a deep learning approach that is based on an LSTM network and (pre-trained) word embeddings. Evaluation on a held-out corpus shows that the SVM model outperforms the neural network approach and benefits from combining lexical, semantic and syntactic information sources. A qualitative analysis of the classification output reveals that the classifier performance may be further enhanced by integrating implicit sentiment information and context- and user-based features.
This study examines Swedish morning TV’s framing of the phenomenon of exercising. Morgonstudion in SVT, and Nyhetsmorgon in TV4, is the morning shows that has been investigated. The aim of the study is to enlighten and enhance the understanding of how the phenomenon of exercising are framed in Swedish morning television, as well as to contribute to the theorization of media’s representation of exercise. A qualitative content analysis has been used to capture the language, to see how they present exercising and if they legitimate it. Theoretical framework applied are framing theory, representation, legitimize, healthism and public service vs commercial television. The research fields are health communication, exercise in media and morning television journalism. The result shows that Swedish morning television, through different methods and lexical choices, legitimize the phenomenon of exercise. SVT is more focused on exercise that suits everyone and that viewers can change their lifestyle by making small changes in their everyday life. For example, they have an idea of how to get the pulse up by exercises in the garden while TV4 is aiming at reaching out to those that is already exercising and specifies types of exercising during the show. They give advice on how to find out what training tools you need to complete a triathlon for example. Lexical choices reinforce that exercise is something positive and both the guests and hostess sees exercise as a norm. The elements of exercise differ between the programs, as well as the studio environment and the content. On the other hand, there are similarities such as the language that is used and they are both visited by experts.
In this paper, we show that building a treebank can be used as a way to establish a language. Annotated corpus can be used as tools when arguing that some linguistic data belongs to a separate language (rather than a dialect or variety of another established language). We provide here a case study on a treebank of Naija, a Post-creole spoken in Nigeria which presents us with significant differences from treebanks of English in terms of existing constructions and frequency of several syntactic units.
It is generally believed that concepts can be characterized by their properties (or features). When investigating concepts encoded in language, researchers often ask subjects to produce lists of properties that describe them (i.e., the Property Listing Task, PLT). These lists are accumulated to produce Conceptual Property Norms (CPNs). CPNs contain frequency distributions of properties for individual concepts. It is widely believed that these distributions represent the underlying semantic structure of those concepts. Here, instead of focusing on the underlying semantic structure, we aim at characterizing the PLT. An often disregarded aspect of the PLT is that individuals show intersubject variability (i.e., they produce only partially overlapping lists). In our study we use a mathematical analysis of this intersubject variability to guide our inquiry. To this end, we resort to a set of publicly available norms that contain information about the specific properties that were informed at the individual subject level. Our results suggest that when an individual is performing the PLT, he or she generates a list of properties that is a mixture of general and distinctive properties, such that there is a non-linear tendency to produce more general than distinctive properties. Furthermore, the low generality properties are precisely those that tend not to be repeated across lists, accounting in this manner for part of the intersubject variability. In consequence, any manipulation that may affect the mixture of general and distinctive properties in lists is bound to change intersubject variability. We discuss why these results are important for researchers using the PLT.
Abstract Language production ultimately aims to convey meaning. Yet, words differ widely in the richness and density of their semantic representations and these differences impact conceptual and lexical processes during speech planning. Here, we replicate the recent finding that semantic richness, measured as the number of associated semantic features according to semantic feature production norms, facilitates object naming, while intercorrelational semantic feature density, measured as the degree of intercorrelation of a concept’s features, has an inhibitory influence, and investigate the electrophysiological correlates of the obtained effects. Both the facilitatory effect of high semantic richness and the inhibitory influence of high feature density were reflected in an increased posterior positivity starting at about 250 ms, in line with previous reports of posterior positivities in paradigms employing contextual manipulations to induce semantic interference during language production. Furthermore, amplitudes at the same posterior electrode sites were positively correlated with object naming times between about 230 and 380 ms. The observed effects follow naturally from the assumption of conceptual facilitation and simultaneous lexical competition, and are difficult to explain by language production theories dismissing lexical competition.
We present the first freely available dependency treebank of Sanskrit. It is based on text from Panchatantra, an ancient Indian collection of fables. The annotation scheme we chose is that of Universal Dependencies, a current de-facto standard for cross-linguistically comparable morphological and syntactic annotation. In the present paper, we discuss word segmentation issues, morphological inventory and certain interesting syntactic constructions in the light of the Universal Dependencies guidelines. We also present an initial parsing experiment.
The translation of certain toponyms that had not yet been assimilated to the Romanian language at the beginning of the 19th century represented a real challenge for translators at that time. A first aspect to be considered here is the linguistic status as proper names and the possible translation options that could not be correlated to any tradition. A second aspect is the precarious stage of Romanian geographical terminology, reflected by the terminological variation for the same concept and the lack of semantic affinity, either real or related to the actual terminology. This article addresses mainly the first aspect mentioned above. The issues addressed are as follows: a concise presentation of the concept of proper names translation, the distinction between untranslatable and translatable or partially translatable proper names, the factors motivating the option of translating or not the translatable terms from a toponymic collocation. Our corpus reflects the incipient stage of the translation of translatable or partially translatable toponyms in Romanian, a stage in which the translator is free to decide upon translatability. Compared to the actual norm, the different choices from one translator to another or even those opted for by the same translator—especially the option of not translating toponyms that are nowadays translated in most languages—reveal the lack of importance of linguistic meaning (that is the lexical meaning of the etymon) of the proper name as far as its functioning was concerned, as well as the role of this non-functionality in identifying the linguistic status of a proper name.
RESUMEN EN CASTELLANO Esta tesis contribuye al estudio de analisis linguistico del ingles antiguo con bases de datos lexicas basadas en corpus. Aunque la lematizacion es considerada una de las tareas necesarias para la creacion de diccionarios, no se dispone de corpus lematizados en ingles antiguo. Ademas, en el caso de este periodo historico del ingles, que presenta numerosas variantes morfologicas y carece de estandar ortografico, es imprescindible disponer de un corpus lematizado. Por ello, el objetivo de esta tesis es lematizar una parte del lexico verbal derivado del ingles antiguo, lo que combina aspectos de morfologia, lexicografia y analisis de corpus. El alcance se restringe a las clases verbales mas complejas morfologicamente del ingles antiguo, verbos irregulares y verbos reduplicativos, que incluyen los preterito-presentes, los anomalos, los contractos y los fuertes de la clase VII. Esto requiere, en primer lugar, la seleccion y el manejo de las fuentes de datos y de verificacion de resultados, y en segundo lugar, la formulacion y secuenciado de los pasos de las tareas de lematizacion. Este trabajo tambien plantea la cuestion de la automatizacion en el proceso de la lematizacion, sobre la que escasa bibliografia se ha encontrado. La metodologia combina busquedas automaticas en el lematizador Norna y la revision manual de los resultados con las fuentes lexicograficas disponibles. El lematizador esta basado en la version 2004 del corpus de The Dictionary of Old English (DOE), que contiene aproximadamente tres mil textos y tres millones de palabras. Las fuentes lexicograficas consultadas son, por un lado, la base de datos The Grid (Nerthus Project), y por otro lado, los diccionarios de ingles antiguo, icluyendo el DOE, Bosworth and Toller, Hall-Meritt, and Sweet. Se han tenido en cuenta dos enfoques diferentes para la lematizacion en esta investigacion. Los verbos fuertes de la clase VII se han lematizado aplicando un algoritmo de busqueda basado en las formas principales del verbo (Metola Rodriguez 2015). Este algoritmo se ha creado a partir de los radicales, las flexiones y los elementos preverbales de los verbos fuertes del ingles antiguo. Por otra parte, los verbos derivados de los preterito-presentes, contractos y anomalos se han buscado a partir de sus formas simples. En conclusion, esta tesis ofrece un inventario de lemas y formas flexivas de los verbos analizados. Desde el punto de vista de la aplicabilidad, este trabajo presenta diferentes procedimientos de lematizacion automatica y manual que pueden ser aplicados a los campos de la lexicografia y la linguistica de corpus. RESUMEN EN INGLES This thesis contributes to the research in the linguistic analysis of Old English with corpus-based lexical databases. Although lemmatisation is generally accepted as one of the necessary tasks of dictionary making, no lemmatised corpus is available in Old English. In the specific area of Old English, which presents numerous morphological variations and lacks a written standard, a lemmatised corpus is necessary. Thus, the aim of this thesis is to lemmatise a part of the derived verbal lexicon of Old English, combining aspects of Morphology, Lexicography and Corpus Analysis. The scope is restricted to the most morphologically complex verbal classes of Old English, including irregular verbs and reduplicative verbs, which comprise preterite-present, anomalous, contracted and strong VII verbs. This aim requires, firstly, the selection and management of the sources of data and verification of results; and secondly, the design and sequencing of the steps of the lemmatisation tasks. This research also raises the issue of the automatisation of the process of lemmatisation of Old English verbs, on which little previous literature has been found. The methodology comprises automatic searches on the lemmatiser Norna and the manual revision of the hits with the available lexicographical sources. The lemmatiser is based on the 2004 version of The Dictionary of Old English Corpus (DOE), which contains approximately three thousand texts and three million words. The lexicographical sources checked are, in the first place, the database The Grid (Nerthus Project), and secondly, the Old English dictionaries, including the DOE, Bosworth and Toller, Hall-Meritt, and Sweet. Two different approaches to lemmatisation have been taken in this research. On the one hand, the class VII strong verbs are lemmatised by means of a search algorithm that is based on the main forms of the verbs (Metola Rodriguez 2015). The search algorithm is created on the basis on the roots, the set of inflections and the preverbal items of the strong verbs of Old English. On the other hand, the derived preterite-present, anomalous and contracted verbs are searched by means of their simplexes. In conclusion, this thesis offers an inventory of inflectional forms and lemmas of the verbs under analysis. On the applied side, this work presents different procedures of automatic and manual lemmatisation that can be applied to the fields of Lexicography and Corpus Linguistics.
This paper is an analysis of the Hill‟s Strategy Development Framework and application of the framework to a Fast Food Restaurant business which is operating in Brunei Darussalam.The study of this article concentrates on corporate objectives, marketing strategy, order qualifiers, order winners, and the operations strategy within the company.A review of the relevant literature conducted on corporate goals, competitive priorities, Order Qualifier and Order Winner. The methodology based on a desk review of secondary data and non-participant observations research approach.The findings demonstrated that Fast Food Restaurant Business in Brunei Darussalam required more considerable attention to focus on improving the Customer Service Relationship (CSR) value.It is crucial to concentrate on CSR that would enhance the business brand value, a better image rating and thereby contributing to company‟s sales and gaining a new customer.The study recommended that Fast Food Restaurant Business create a new market such as selling the frozen product.Fast Food Restaurant Business potentially can sell this in their store or export their patent right on the product internationally.The contribution of this paper is to provide provides a review of the application of the strategic framework to uncover the potential customer benefits package of the strategy.
The article is devoted to a modern problems research of television titles editing of the media addressee by the media addressor. The objectives of work are achieved by application of methods of deductive and inductive logical analysis, descriptive method, content analysis, lexical and semantic, lexical and grammatical, and stylistic analysis, comparative analysis, deep interview and poll of informants. Titles as fragments of media texts of the television program “Time Will Show” act as material of the research. In the article, attention is paid to a problem of television titles editing of the mediaaddressee in which editorial work of the media addressor is often limited to an inscription: “The spelling and a punctuation of the author are kept”. Authors in details analyze television titles of the mass media addressee of the “Time Will Show” program broadcast on Channel 1 of the Russian television; sort a number of examples from other elements of media system subject to influence of the research object and from fiction with justified violation of language norms. Authors offer the answer to a question how the editor has to work with text elements on the screen, support the position with opinions of scientists and results of the comparative analysis of the actual material, deep interview and poll of social networks users. The received results are significant for development of psycholinguistics, pragmalinguistics, cognitive linguistics, linguosemiotics, cultural linguistics, discursive linguistics, in particular, of media discourse theory, influence theory. This is because the peculiarities of verbal self-presentation of the media addressee in television titles characterizing the language personality of the media addressee as the media addressor producing the media message as reaction to a television message of the media addressor, and making influence on the mass addressee – television audience are revealed. The article will be useful to philologists, editors, journalists not only in the theoretical plan, but also in practical work.
The predictive validity of various corpus-based frequency norms in first-language lexical processing has been intensively investigated in previous research, but less attention has been paid to this issue in second-language (L2) processing. To bridge the gap, in the present study we took English as a case in point and compared the predictive power of a large set of corpus-based frequency norms for the performance of an L2 English visual lexical decision task (LDT). Our results showed that, in general, the frequency norms from SUBTLEX-US and WorldLex–Blog tended to predict L2 performance better in reaction times, whereas the frequency norms from corpora with a mixture of written and spoken genres (CELEX, WorldLex–Blog, BNC, ANC, and COCA) tended to predict L2 accuracy better. Although replicated in both low- and high-proficiency L2 English learners, these patterns were not exactly the same as those found in LDT data from native English speakers. In addition, we only observed some limited advantages of the lemma frequency and contextual diversity measures over the wordform frequency measure in predicting L2 lexical processing. The results of the present study, especially the detailed comparisons among the different corpora, provide methodological implications for future L2 lexical research.
Large-scale semantic norms have become both prevalent and influential in recent psycholinguistic research. However, little attention has been directed towards understanding the methodological best practices of such norm collection efforts. We compared the quality of semantic norms obtained through rating scales, numeric estimation, and a less commonly used judgment format called best-worst scaling. We found that best-worst scaling usually produces norms with higher predictive validities than other response formats, and does so requiring less data to be collected overall. We also found evidence that the various response formats may be producing qualitatively, rather than just quantitatively, different data. This raises the issue of potential response format bias, which has not been addressed by previous efforts to collect semantic norms, likely because of previous reliance on a single type of response format for a single type of semantic judgment. We have made available software for creating best-worst stimuli and scoring best-worst data. We also made available new norms for age of acquisition, valence, arousal, and concreteness collected using best-worst scaling. These norms include entries for 1,040 words, of which 1,034 are also contained in the ANEW norms (Bradley & Lang, Affective norms for English words (ANEW): Instruction manual and affective ratings (pp. 1-45). Technical report C-1, the center for research in psychophysiology, University of Florida, 1999).
Automatically recognized terminology is widely used for various domain-specific texts processing tasks, such as machine translation, information retrieval or ontology construction. However, there is still no agreement on which methods are best suited for particular settings and, moreover, there is no reliable comparison of already developed methods. We believe that one of the main reasons is the lack of state-of-the-art method implementations, which are usually non-trivial to recreate—mostly, in terms of software engineering efforts. In order to address these issues, we present ATR4S, an open-source software written in Scala that comprises 13 state-of-the-art methods for automatic terminology recognition (ATR) and implements the whole pipeline from text document preprocessing, to term candidates collection, term candidate scoring, and finally, term candidate ranking. It is highly scalable, modular and configurable tool with support of automatic caching. We also compare 13 state-of-the-art methods on 7 open datasets by average precision and processing time. Experimental comparison reveals that no single method demonstrates best average precision for all datasets and that other available tools for ATR do not contain the best methods.
The traditional understanding of data from Likert scales is that the quantifications involved result from measures of attitude strength. Applying a recently proposed semantic theory of survey response, we claim that survey responses tap two different sources: a mixture of attitudes plus the semantic structure of the survey. Exploring the degree to which individual responses are influenced by semantics, we hypothesized that in many cases, information about attitude strength is actually filtered out as noise in the commonly used correlation matrix. We developed a procedure to separate the semantic influence from attitude strength in individual response patterns, and compared these results to, respectively, the observed sample correlation matrices and the semantic similarity structures arising from text analysis algorithms. This was done with four datasets, comprising a total of 7,787 subjects and 27,461,502 observed item pair responses. As we argued, attitude strength seemed to account for much information about the individual respondents. However, this information did not seem to carry over into the observed sample correlation matrices, which instead converged around the semantic structures offered by the survey items. This is potentially disturbing for the traditional understanding of what survey data represent. We argue that this approach contributes to a better understanding of the cognitive processes involved in survey responses. In turn, this could help us make better use of the data that such methods provide.
In this paper, we present an approach for automatically creating a combinatory categorial grammar (CCG) treebank from a dependency treebank for the subject–object–verb language Hindi. Rather than a direct conversion from dependency trees to CCG trees, we propose a two stage approach: a language independent generic algorithm first extracts a CCG lexicon from the dependency treebank. An exhaustive CCG parser then creates a treebank of CCG derivations. We also discuss special cases of this generic algorithm to handle linguistic phenomena specific to Hindi. In doing so we extract different constructions with long-range dependencies like coordinate constructions and non-projective dependencies resulting from constructions like relative clauses, noun elaboration and verbal modifiers.
This article presents a comparison of different Word Sense Induction (wsi) clustering algorithms on two novel pseudoword data sets of semantic-similarity and co-occurrence-based word graphs, with a special focus on the detection of homonymic polysemy. We follow the original definition of a pseudoword as the combination of two monosemous terms and their contexts to simulate a polysemous word. The evaluation is performed comparing the algorithm’s output on a pseudoword’s ego word graph (i.e., a graph that represents the pseudoword’s context in the corpus) with the known subdivision given by the components corresponding to the monosemous source words forming the pseudoword. The main contribution of this article is to present a self-sufficient pseudoword-based evaluation framework for wsi graph-based clustering algorithms, thereby defining a new evaluation measure (top2) and a secondary clustering process (hyperclustering). To our knowledge, we are the first to conduct and discuss a large-scale systematic pseudoword evaluation targeting the induction of coarse-grained homonymous word senses across a large number of graph clustering algorithms.
Evaluation is crucial in the research and development of automatic summarization applications, in order to determine the appropriateness of a summary based on different criteria, such as the content it contains, and the way it is presented. To perform an adequate evaluation is of great relevance to ensure that automatic summaries can be useful for the context and/or application they are generated for. To this end, researchers must be aware of the evaluation metrics, approaches, and datasets that are available, in order to decide which of them would be the most suitable to use, or to be able to propose new ones, overcoming the possible limitations that existing methods may present. In this article, a critical and historical analysis of evaluation metrics, methods, and datasets for automatic summarization systems is presented, where the strengths and weaknesses of evaluation efforts are discussed and the major challenges to solve are identified. Therefore, a clear up-to-date overview of the evolution and progress of summarization evaluation is provided, giving the reader useful insights into the past, present and latest trends in the automatic evaluation of summaries.
This paper proposes the model for searching similar collocations in English texts in order to determine semantically connected text fragments for social network data streams analysis. The logical-linguistic model uses semantic and grammatical features of words to obtain a sequence of semantically related to each other text fragments from different actors of a social network. In order to implement the model, we leverage Universal Dependencies parser and Natural Language Toolkit with the lexical database WordNet. Based on the Blog Authorship Corpus, the experiment achieves over 0.92 precision.
We propose three approaches for disambiguating the Kannada word based on an adaptation of dictionary-based Lesk’s word sense disambiguation technique. Instead of making use of the regular dictionary as the repository of glosses, we used Indo – WordNet lexical database as the source of senses. Here we adopt a current method of measuring semantic relatedness between the concepts of the Kannada words taken from Indo – WordNet. This measure is dependent on identifying and counting the number of common words present between the glosses of a pair of concepts in accordance with Indo – WordNet.
SNOMED CT provides about 300,000 codes with fine-grained concept definitions to support interoperability of health data. Coding clinical texts with medical terminologies it is not a trivial task and is prone to disagreements between coders. We conducted a qualitative analysis to identify sources of disagreements on an annotation experiment which used a subset of SNOMED CT with some restrictions. A corpus of 20 English clinical text fragments from diverse origins and languages was annotated independently by two domain medically trained annotators following a specific annotation guideline. By following this guideline, the annotators had to assign sets of SNOMED CT codes to noun phrases, together with concept and term coverage ratings. Then, the annotations were manually examined against a reference standard to determine sources of disagreements. Five categories were identified. In our results, the most frequent cause of inter-annotator disagreement was related to human issues. In severa)
The syntax of newspaper headlines in English displays features which, on a superficial level, set it apart from the norm of Standard English. This, however, is not to say that headlinese stands outside the notion of a linguistic norm: drawing from a corpus of various headlines, and looking at the specific fields of determiners and verb tenses, this article explores how the English of newspaper headlines constitutes a norm in itself which actually builds on the very potentialities of Standard English. Such an extension of the syntactic possibilities of English is designed to fit the pragmatic purpose of headlines, i.e. heighten the relevance of the newspaper article to the reader.
In recent years, with increase in the use of internet the multimedia contents on it have rapidly increased. Users may need to go through a video in a top down manner i.e. browsing the videos, or in bottom up manner i.e. retrieving specific information from videos. They may also want to go through the summary or through the highlights of the videos. This has necessitated the need to handle multimedia resources effectively. This paper proposes an automatic method for aligning scripts of lecture videos with captions. Alignment is needed to extract time information from captions and insert it in the scripts, to create index of the videos. No alignment work has been previously done in lecture videos domain. Alignment methods proposed for other type of videos are not applicable for lecture videos because, different similarity techniques behave differently on different types of datasets. The proposed method uses transcripts of lecture videos, SRT file of captions available along with lecture videos and captions generated from auto-caption generation feature of YouTube. The captions and scripts are then aligned using a dynamic programming technique. No such work has been previously done for lecture videos. Most important aspect of alignment is similarity measure. In the proposed work we have used three similarity measures cosine, jaccard, and dice. A comparative analysis of these measures is given in the paper. We also use a large lexical database of English words known as WordNet for word-to-word similarity. The experimental result shows comparison of various similarity techniques and YouTube captions.
On the material of the corpus of English-language texts of articles informing about military actions, armed conflicts and defence issues, the analysis of metaphors and their translation into Russian is carried out. The main types of metaphor in military discourse are identified: metaphorical phraseological units, worn out metaphors and metaphors-cliché, terminated metaphor, recent metaphor (neologisms), non-deployed speech metaphor. It is indicated that the most frequent translation techniques used for translation of metaphors in military discourse are demetaphorization, calques, remetaphorization, translation by equivalent metaphorical expression in the target language with the use of translation transformations (addition, omission, antonymy translation). The article focuses on the technique of demetaphorization, which text realizations form 50 % of the total number of cases examined. It is shown that demetaphorization may be accompanied by a combination of translation transformations (generalization, adequate substitute, descriptive translation). It is established that the choice of translation technique depends on the type of metaphor and its structure (component composition). The analysis revealed that the result of the use of demetaforization is the replacement of more expressive, pragmatically loaded lexical units of the original text with neutral equivalents, that is motivated by the norms of the language of translation: Russian military texts are characterized by lexical and stylistic uniformity and much less saturation with stylistically coloured elements in comparison with English texts. However, despite this, the use of demetaforization technique makes it possible to convey the necessary element of the original content, the preservation of which in translation is a minimum condition for providing the recipient of the expression of the translation language with a pragmatic impact, for which the sender of the original language is seeking.
The Prague Dependency Treebank 3.5 is the 2018 edition of the core Prague Dependency Treebank (PDT). It contains all PDT annotation made at the Institute of Formal and Applied Linguistics under various projects between 1996 and 2018 on the original texts, i.e., all annotation from PDT 1.0, PDT 2.0, PDT 2.5, PDT 3.0, PDiT 1.0 and PDiT 2.0, plus corrections, new structure of basic documentation and new list of authors covering all previous editions. The Prague Dependency Treebank 3.5 (PDT 3.5) contains the same texts as the previous versions since 2.0; there are 49,431 annotated sentences (832,823 words) on all layers, from tectogrammatical annotation to syntax to morphology. There are additional annotated sentences for syntax and morphology; the totals for the lower layers of annotation are: 87,913 sentences with 1,502,976 words at the analytical layer (surface dependency syntax) and 115,844 sentences with 1,956,693 words at the morphological layer of annotation (these totals include the annotation with the higher layers annotated as well). Closely linked to the tectogrammatical layer is the annotation of sentence information structure, multiword expressions, coreference, bridging relations and discourse relations.
propos des nologismes lis la mode et de leur circulation en franais et en tchque Rsum La mondialisation, qui acclre les contacts entre les langues, facilite normment la circulation des expressions nologiques dans des langues non apparentes. Un des domaines touchs par des apparitions particulirement nombreuses des expressions nologiques est celui de la mode et du style. De nouveaux mots, tels que hipster, preppy, girly, se propagent rapidement dans la culture anglo-amricaine et envahissent aussi les langues qui sont en contact avec cette culture. Le franais et le tchque ne font pas exception. tant donn que ce champ lexical n'a pas t exploit de manire comparative, nous avons dcid de dcrire certains mots lis au style dans ces deux langues. Sur un corpus fond sur nos propres connaissances du thme et sur un lexique trouv dans la presse, nous essayons de dcrire la prsence des lexmes choisis dans les deux langues et de comparer leur existence dans les diffrents types de documents, leur diffusion ainsi que leur nature et leur place dans les deux langues tudies, le franais et le tchque.
Dans le cadre général d’une sémiotique des cultures, cette recherche a utilisé les propositions épistémologiques et méthodo-logiques de la sémantique interprétative pour renouveler l’analyse de textes irlandais médiévaux. Le but était d’apporter une contribution à une problématique générale intéressant les sciences du langage, mais aussi les sciences historiques: comment fonder la pertinence scientifique d’une interprétation de textes et signes anciens appartenant à une culture différente? Pour cela il a été choisi de viser les faits sémantiques qui interviennent dans les processus de transfert de sens que la tradition rhétorique nomme comparaison, métaphore ou symbole. L’approche méthodologique a nécessité l’édition d’un corpus interlinéaire offrant un accès direct aux données de l’Electronic Dictionary of the Irish Language. L’étude du lexique a confronté les possibilités définitoires aux afférences contextuelles observées par un relevé systématique des isotopies ciblées.Sur le plan sémantique, l’analyse des processus différentiels qui structurent les molécules sémiques a permis d’observer la circulation des sèmes marquant les analogies intentionnelles. Sur le plan diachronique, la description du système de valeur, pris dans sa globalité, fonde la pertinence de l’interprétation en ce qu’il intègre les normes sociales du contexte historique du signe. Sur le plan des études celtiques, l’analyse des correspondances entre les domaines de l’orientation spatiale, des cycles temporels et des fonctions sociales donne un nouvel accès à la complexité du système de pensée de cette culture. Les formes sémantiques décrites fournissent de nouveaux modèles pour des comparaisons. Sur cette base, les expressions de l’association arbre-savoir ont été décrites pour apporter une solution aux problèmes de l’étymologie de la lexie druid- et proposer le dépassement des approches lexicales monographiques par l’approche intertextuelle.
he object of this paper is the variant of quasi-standard language, i.e. the variant of the perceived standard language formed by the young generation of the Aukštaitian area. The aims of the study are to examine whether the assessments made by respondents representing the Aukštaitian area suggest the presence of such a quasi-standard and, if they do, to provide the characterisation of this quasi-standard variant on the basis of the collected data. The data of the study consists of eight audio texts-stimuli which represent six Lithuanian regiolects (A and D represent Southern Aukštaitian; B and E represent Žemaitian, C represents Northwestern Aukštaitian, F represents the western part of Eastern Aukštaitian, G represents Southwestern Aukštaitian, while H represents the eastern part of East Aukštaitian) and the responses to two questions in the questionnaire designed according to the principles of perceptual dialectology which ask the respondents to rate the similarity of the audio text-stimulus to the standard language. The study demonstrated that the respondents perceived texts-stimuli B and E (representative of the Žemaitian dialect) as the least similar to the standard language. Texts-stimuli H (representative of the eastern part of East Aukštaitian regiolect), A and D (representative of East Aukštaitian regiolect), on the contrary, were seen as the closest to the standard language. Since the ratings of these texts-stimuli in the respondents’ assessment were substantially higher in comparison to the rest of the texts-stimuli, the results suggest the existence of a quasi-standard. The analysis of respondents’ motives of giving high scores to audio texts-stimuli A, D, and H demonstrates that the morphological and lexical characterisation of the quasi-standard of the young generation representing the Aukštaitian area is only fragmentary. The most prominent are phonetic features, namely: more open and more closed pronunciation of vowels i and u, shortening of unstressed long vowels o, u, and i, lengthening of the stressed short vowels u and i, correct accentuation and non-reduced endings. Based on the analysis carried out, it is possible to assume that the quasi-standard variety formed by the young generation representing the Aukštaitian area consists of some tertiary phonetic features of the eastern parts of East Aukštaitian and South Aukštaitian regiolectal zones, norms of standard language pronunciation, shortened verb forms typically characteristic of dialects and standard language and mixed lexis (containing that of standard language / dialects / borrowings).
Processing of nouns and action verbs can be differentially compromised following lesions to posterior and anterior/motor brain regions, respectively. However, little is known about how these deficits progress in the course of neurodegeneration. To address this issue, we assessed productive lexical skills in a patient with posterior cortical atrophy at two different stages of his pathology. On both occasions, he underwent a structural brain imaging protocol and completed semantic fluency tasks requiring retrieval of animals (nouns) and actions (verbs). Imaging results were compared with those of controls via voxel-based morphometry, whereas fluency performance was compared to age-matched norms through Crawford’s t-tests. In the first assessment, the patient exhibited atrophy of more posterior regions supporting multimodal semantics (medial temporal and lingual gyri), together with a selective deficit in noun fluency. Then, by the second assessment, the patient’s atrophy had progressed mainly towards fronto-motor regions (rolandic operculum, inferior and superior frontal gyri) and subcortical motor hubs (cerebellum, thalamus), and his fluency impairments had extended to action verbs. These results offer unprecedented evidence of the specificity of the pathways related to noun and action-verb impairments in the course of neurodegeneration, highlighting the latter’s critical dependence on damage to regions supporting motor functions, as opposed to multimodal semantic processes.
This work investigates legal concepts and their expression in Portuguese, concentrating on the “Order of Attorneys of Brazil” Bar exam. Using a corpus formed by a collection of multiple-choice questions, three norms related to the Ethics part of the OAB exam, language resources (Princeton WordNet and OpenWordNet-PT) and tools (AntConc and Freeling), we began to investigate the concepts and words missing from our repertory of concepts and words in Portuguese, the knowledge base OpenWordNet-PT. We add these concepts and words to OpenWordNet-PT and hence obtain a representation of these texts that is mostly “contained” in the lexical knowledge base.
This study explores the use of kin terms in a corpus of Vietnamese–English bilingual spontaneous conversation. While the corpus features a range of single Vietnamese lexical items in otherwise English discourse, kin terms, as in the example below, are overwhelmingly the most frequent (accounting for 84%, 164/196, of single Vietnamese words in an English context). Borrowing or Code-switching? Traces of community norms in Vietnamese-English speechAll authorsLi Nguyen http://orcid.org/0000-0001-8632-7909https://doi.org/10.1080/07268602.2018.1510727Published online:09 October 2018Table Download CSVDisplay Table The study puts forward an empirical attempt at determining whether such items should be considered code-switches or borrowings, and the role that pragmatic norms play in shaping this linguistic behaviour. Discourse distribution of the kin terms in terms of person reference and syntactic role are used as cross-language ‘conflict sites’, to determine the level of integration of such items as a test of their status as code-switches or borrowings. This reveals that the distribution of Vietnamese kin terms in an otherwise English context mirrors that of Vietnamese kin terms in monolingual Vietnamese, and is distinct from that of English kin terms. This measure of integration suggests that these may be single-word code-switches. Nonetheless, the high frequency of use and their diffusion across the community are suggestive of borrowings. Follow-up interviews with the participants reveal specific community norms that underlie the use of these terms, namely as a linguistic resource to retain, promote and conform to community cultural practice. While the paper acknowledges the difficulty in determining the exact status of these forms based on existing criteria, it demonstrates how judicious application of empirical methodology enables us to pinpoint such strategies in studying language in contact.
The article analyses the attitudes of the students of Lithuanian University of Health Sciences (LSMU) regarding the use of loanwords in medical terminology. Many of the previously unacceptable and avoidable barbarisms used in the book of language tips “Lexicology: usage of loanwords” (published in 2013), became standard versions of the general language. Among such loanwords are the medical terms, which led to carrying out this research. Survey respondents were second-year students of the Faculty of Medicine of LSMU. The study involved 100 students. A list of 40 medical terms was specifically made for this research and was submitted to the students. The aim of this article is to research whether students are able to determine the ratio of a loanword to the norm of the standard language. The article raises the hypothesis that the students will be able to differentiate normative and non-standard loanwords. The analysis showed that 85.8% of respondents indicated loanwords being totally unacceptable and avoidable. Twice as many respondents reported that the loanwords, considered as side version of the norm, did not meet the requirements of the standard language norm. More than half (56.5%) of respondents indicated that loanwords, considered as an equivalent variant of the norm, did not meet the requirements of the standard language norm. The research hypothesis proved to be true.
The article is devoted to the study of associative field of the word using the method of experiment. The paper presents a brief review of the definitions of concepts such as: free associative experiment, association, content analysis. The article attempts to analyze the results of the free associative experiment, associative field study of steppe. The factual material for illustration of the main provisions was the direct lexical association of the respondents. The experimental data allow us to observe mental stereotypes of society, to reveal its cultural memory, modern values verbalized in associations. The following methods were used in the research: questionnaire survey, descriptive, generalization, systematization, observation, content analysis.
Sentiment analysis involves classifying text into positive, negative and neutral classes according to the emotions expressed in the text. Extensive study has been carried out in performing sentiment analysis using the traditional ‘bag of words’ approach which involves feature selection, where the input is given to classifiers such as Naive Bayes and SVMs. A relatively new approach to sentiment analysis involves using a deep learning model. In this approach, a recently discovered technique called word embedding is used, following which the input is fed into a deep neural network architecture. As sentiment analysis using deep learning is a relatively unexplored domain, we plan to perform in-depth analysis into this field and implement a state of the art model which will achieve optimal accuracy. The proposed methodology will use a hybrid architecture, which consists of CNNs (Convolutional Neural Networks) and RNNs (Recurrent Neural Networks), to implement the deep learning model on the SAR14 and Stanford Sentiment Treebank data sets.
Best-worst scaling is a judgment format in which participants are presented with a set of items and have to choose the superior and inferior items in the set. Best-worst scaling generates a large quantity of information per judgment because each judgment allows for inferences about the rank value of all unjudged items. This property of best-worst scaling makes it a promising judgment format for research in psychology and natural language processing concerned with estimating the semantic properties of tens of thousands of words. A variety of different scoring algorithms have been devised in the previous literature on best-worst scaling. However, due to problems of computational efficiency, these scoring algorithms cannot be applied efficiently to cases in which thousands of items need to be scored. New algorithms are presented here for converting responses from best-worst scaling into item scores for thousands of items (many-item scoring problems). These scoring algorithms are validated through simulation and empirical experiments, and considerations related to noise, the underlying distribution of true values, and trial design are identified that can affect the relative quality of the derived item scores. The newly introduced scoring algorithms consistently outperformed scoring algorithms used in the previous literature on scoring many-item best-worst data.
Recent investigations have established the value of using rebus puzzles in studying the insight and analytic processes that underpin problem solving. The current study sought to validate a pool of 84 rebus puzzles in terms of their solution rates, solution times, error rates, solution confidence, self-reported solution strategies, and solution phrase familiarity. All of the puzzles relate to commonplace English sayings and phrases in the United Kingdom. Eighty-four rebus puzzles were selected from a larger stimulus set of 168 such puzzles and were categorized into six types in relation to the similarity of their structures. The 84 selected problems were thence divided into two sets of 42 items (Set A and Set B), with rebus structure evenly balanced between each set. Participants (N = 170; 85 for Set A and 85 for Set B) were given 30 s to solve each item, subsequently indicating their confidence in their solution and self-reporting the process used to solve the problem (analysis or insight), followed by the provision of ratings of the familiarity of the solution phrases. The resulting normative data yield solution rates, error rates, solution times, confidence ratings, self-reported strategies and familiarity ratings for 84 rebus puzzles, providing valuable information for the selection and matching of problems in future research.
Human free association (FA) norms are believed to reflect thestrength of links between words in the lexicon of an averagespeaker. Large-scale FA norms are commonly used as a datasource both in psycholinguistics and in computational mod-eling. However, few studies aim to analyze FA norms them-selves, and it is not known what are the most important factorsthat guide speakers’ lexical choices in the FA task. Here, wefirst provide a statistical analysis of a large-scale data set ofEnglish FA norms. Second, we argue that such analysis caninform existing computational models of semantic memory,and present a case study with the topic model to support thisclaim. Based on our analysis, we provide the topic model withdictionary-based knowledge about word synonymy/antonymy,and demonstrate that the resulting model predicts human FAresponses better than the topic model without this information.
<p>Gender identity, one of the most important social categories in people’s lives, is socially constructed and language is claimed to have a significant role in constructing the gender identity. This paper studies the construction of Sundanese women through five Sundanese nouns referring to women found in the corpus of <em>Manglè </em>magazine, published between 1958–2013. The research employs a mixed-method design in which quantitative analysis is combined with qualitative analysis to investigate how the nouns referring to women are used to construct Sundanese women from the periods of Guided Democracy (1958–1965) to Reformation (2004–2013). The quantitative analysis is used to examine the frequency of word occurrence diachronically. The frequency of word accurrence is subsequently interpreted qualitatively by considering social and cultural contexts, such as the norms of speech levels in Sundanese, Sundanese belief about marriage, and gender issues. The result of analysis shows that women are constructed in various identities by every noun referring to them. The lexical choices used to contruct women are greatly influenced by the social and cultural contexts. </p>
At the interface between scene perception and speech production, we investigated how rapidly action scenes can activate semantic and lexical information. Experiment 1 examined how complex action-scene primes, presented for 150 ms, 100 ms, or 50 ms and subsequently masked, influenced the speed with which immediately following action-picture targets are named. Prime and target actions were either identical, showed the same action with different actors and environments, or were unrelated. Relative to unrelated primes, identical and same-action primes facilitated naming the target action, even when presented for 50 ms. In Experiment 2, neutral primes assessed the direction of effects. Identical and same-action scenes induced facilitation but unrelated actions induced interference. In Experiment 3, written verbs were used as targets for naming, preceded by action primes. When target verbs denoted the prime action, clear facilitation was obtained. In contrast, interference was observed when)
The article highlights an issue of the spiritual relationship of a person and his language, and hence the national-linguistic behavior of the individual and community. The issues are topical in the current linguistic studies. The translator plays an important role in the perception of the linguistic picture of the world created by one people by representatives of another nation, especially when talking about cognition of the foreign reality via the literary text, where the word reflects the national-linguistic behavior of the authors of the original and translated text. The article examines the peculiarities of P. Kulish's national-linguistic behavior as a pioneer of the Ukrainian translation school, whose creative work has repeatedly been the subject of research by scientists in various fields of knowledge. In accordance with the purpose and objectives of the study, the author argues the definition of the concept of national-linguistic behavior. The typological and nationally-marked features of addresses in the Ukrainian speech discourse and specific features of address constructions translation are compared and generalized according to the theoretical provisions of scholars. The emphasis is laid on the word-forming and lexical capabilities of transferring addresses in Ukrainian translation. The most frequent techniques are transcribing personal and some general foreign appellatives in accordance with the spelling norms, existing in the time of P. Kulish, rendering honoratives/honorifics by Ukrainian equivalents, or sporadically by transliteration, using diminutives and emotive addresses with nationally-marked lexemes of that time, or emotive attributes in the structure of commonly used addressing constructions, typical of the Ukrainians.
This paper analyzes distributional properties that facilitate the categorization of words into lexical categories. First, word-context co-occurrence counts were collected using corpora of transcribed English child-directed speech. Then, an unsupervised k-nearest neighbor algorithm was used to categorize words into lexical categories. The categorization outcome was regressed over three main distributional predictors computed for each word, including frequency, contextual diversity, and average conditional probability given all the co-occurring contexts. Results show that both contextual diversity and frequency have a positive effect while the average conditional probability has a negative effect. This indicates that words are easier to categorize in the face of uncertainty: categorization works best for words which are frequent, diverse, and hard to predict given the co-occurring contexts. This shows how, in order for the learner to see an opportunity to form a category, there needs to)
The recent growth in low-cost eye-tracking systems makes it feasible to incorporate real-time measurement and analysis of eye position data into activities such as learning to read. It also enables field studies of reading behavior in the classroom and other learning environments. We present a study of the data quality provided by two remote eye trackers, one being a low-sampling-rate, low-cost system. Then we present two algorithms for mapping fixations derived from the data to the words being read. One is for immediate (or real-time) mapping of fixations to words and the other for deferred (or post hoc) mapping. Following this, an evaluation study is reported. Both studies were carried out in the classroom of a Finnish elementary school with students who were second graders. This study shows very high success rates in automatically mapping fixations to the lines of text being read when the mapping is deferred. The success rates for immediate mapping are comparable with those obtained in earlier studies, although here the data is collected some 10 min after initial calibration of low-sample (30 Hz) remote eye trackers, rather than a laboratory setting using high-sampling-rate trackers. The results provide a solid basis for developing systems for use in classrooms and other learning environments that can provide immediate automatic support with reading, and share data between a group of learners and the teacher of that group. This makes possible new approaches to the learning of reading and comprehension skills.