Papers reviewed and determined not to be word norm studies. Use the flag icon to report errors or suggest re-inclusion.
18265 papers
It is now widely acknowledged that, between the fifteenth and the seventeenth centuries, most of the European national grammatical traditions were derived from the long-established Graeco-Latin descriptive and normative framework. However, a thorough investigation into the particulars of such a ‘transfer’, at a time when the first grammars of the English vernacular were progressively translated from the Latin, or written directly in English, is still to be carried out. How were classical linguistic norms practically transferred to English? How was usage then looked upon? Why was it decided that the vernacular should be taught? In this paper, I will examine in some detail how the first English grammarians integrated the vernacular of England into the paradigm of Latin grammar. We shall also study the conditions under which the specificities of English were revealed. But our main concern will be to try and determine whether any challenging position to the prevalent model that had sprung from the Graeco-Roman tradition can be traced back. If this long-established tradition appeared at that time as the only one apt to guarantee the efficiency of grammatical description, it seems that it was consistently challenged even by the most prominent figures.
L’informatisation des alignements textuels est confrontée à la complexité de l’organisation textuelle et discursive. L’architecture modulaire Trame/Cadre issue des recherches menées en textométrie facilite la navigation dans l’espace textuel multilingue. Le flux textuel est représenté par un système de coordonnées sur le texte (la Trame ). Le calcul d’une Trame permet une identification précise des objets ( contenants et contenus ) nécessaires aux repérages contextuels (le Cadre ). La construction d’un Cadre permet de stocker non seulement les découpages du texte mais aussi les annotations produites par différentes procédures informatiques (y compris les alignements) et, éventuellement, de les faire passer d’une procédure de traitement à l’autre. Ces états successifs de traitement induisent la notion de ressource textuelle incrémentale qui conserve la trace de séquences de traitement apportées à la ressource textuelle initiale, avec apport de méthodes quantitatives. Cette approche est implémentée au sein du logiciel Le Trameur qui permet d’explorer les corpus multilingues richement annotés ( treebanks ).
This chapter focuses on ‘female’ and ‘male’ speech genres and cults, and their representations in linguistic practices. Social Genders in traditional societies tend to be associated with different domains, and different speech styles. This is where inequality between the male and the female Social Genders comes to light. Special languages and language registers can come to be used in male-only rituals. A whole set of terms may be forbidden to women. Women may have to exercise special caution and use deferential forms to address their relatives through marriage. Women are often seen as keepers and promoters of prestigious linguistic norms, and of traditional language. Or they can be viewed as a dangerous ‘other’ who lead the society in the wrong direction.
In this work, we have created a semantic similarity calculation system between text documents to contribute to their semantic clustering. Indeed, semantic clustering of documents is a promising field of research, since it guarantees a quick and targeted access to information. The aim of document clustering is to put together similar documents. We used the algebraic model VSM (Vector Space Model) [2] to represent text documents and the WordNet [1] lexical database, in that it groups words together based on their meanings. In this paper, we will present an overview of the static and semantic methods for calculating the similarity measure and the appropriateness of these methods. As our research is focusing on the treatment of text documents on e-learning systems. We worked on a corpus of a set of text documents from the computer science textbook for high school students in Morocco. To evaluate our system, an experiment has been conducted among students who produced text documents. Experimental evaluations using WordNet prove that the system presented in this work improves the accuracy of semantic similarity between the text documents.
Dans le cadre des recherches sur l’interface prosodie, syntaxe et genres de discours, l’objectif de cet article est d’exposer une modélisation de la structure prosodique du français parlé, guidée par les données d’usage, et de présenter une première exploration statistique du modèle Rhapsodie (un treebank multilocuteurs et multigenres de 3 heures, annoté en prosodie), qui offre un retour précis sur le rôle du marquage prosodique dans la caractérisation linguistique des genres discursifs.
The contents and structure of semantic networks have been the focus of much recent research, with major advances in the development of distributional models.In parallel, connectionist modeling has extended our knowledge of the processes engaged in semantic activation.However, these two lines of investigation have rarely brought together.Here, starting from a standard textual model of semantics, we allow activation to spread throughout its associated semantic network, as dictated by the patterns of semantic similarity between words.We find that the activation profile of the network, measured at various time points, can successfully account for response times in the lexical decision task, as well as for subjective concreteness and imageability ratings.
Grammar-based surface realizers require inputs compatible with their reversible, constraint-based grammars, including a proper representation of unbounded dependencies and coordination. In this paper, we report on progress towards creating realizer inputs along the lines of those used in the first surface realization shared task that satisfy this requirement. To do so, we augment the Universal Dependencies that result from running the Stanford Dependency Converter on the Penn Treebank with the unbounded and coordination dependencies in the CCGbank, since only the latter takes the Penn Treebank's trace information into account. An evaluation against gold standard dependencies shows that the enhanced dependencies have greatly enhanced recall with moderate precision. We conclude with a discussion of the implications of the work for a second realization shared task.
The primary goal in this thesis is to identify better syntactic constraint or bias, that is language independent but also efficiently exploitable during sentence processing. We focus on a particular syntactic construction called center-embedding, which is well studied in psycholinguistics and noted to cause particular difficulty for comprehension. Since people use language as a tool for communication, one expects such complex constructions to be avoided for communication efficiency. From a computational perspective, center-embedding is closely relevant to a left-corner parsing algorithm, which can capture the degree of center-embedding of a parse tree being constructed. This connection suggests left-corner methods can be a tool to exploit the universal syntactic constraint that people avoid generating center-embedded structures. We explore such utilities of center-embedding as well as left-corner methods extensively through several theoretical and empirical examinations. Our primary task is unsupervised grammar induction. In this task, the input to the algorithm is a collection of sentences, from which the model tries to extract the salient patterns on them as a grammar. This is a particularly hard problem although we expect the universal constraint may help in improving the performance since it can effectively restrict the possible search space for the model. We build the model by extending the left-corner parsing algorithm for efficiently tabulating the search space except those involving center-embedding up to a specific degree. We examine the effectiveness of our approach on many treebanks, and demonstrate that often our constraint leads to better parsing performance. We thus conclude that left-corner methods are particularly useful for syntax-oriented systems, as it can exploit efficiently the inherent universal constraints in languages.
This paper address the problem of extracting learning objects or keywords from bunch of Documents,with the goal of use this objects for various purpose like Resume filtering,Email filtering,content classification etc. Keyword extraction, concept finding are in learning objects is very important subject in today’s eLearning environment. Keywords are subset of words that contains the useful information about the content of the document. Keyword extraction is a process that is used to get the important keywords from documents. In this proposed System I Calculate the TF-IDF of each word, then Decision tree algorithm is used for feature selection process using wordnet dictionary. WordNet is a lexical database of English which is used to find similarity from the candidate words. The words having highest similarity are taken as keywords.
Individuals differ in their ability to feel their own and others' internal states, with those that have more autistic and less empathic traits clustering at the clinical end of the spectrum. However, when we consider semantic competence, this group could compensate with a higher capacity to imagine the meaning of words referring to emotions. This is indeed what we found when we asked people with different levels of autistic and empathic traits to rate the degree of imageability of various kinds of words. But this was not the whole story. Individuals with marked autistic traits demonstrated outstanding ability to imagine theoretical concepts, i.e., concepts that are commonly grasped linguistically through their definitions. This distinctive characteristic was so pronounced that, using tree-based predictive models, it was possible to accurately predict participants' inclination to manifest autistic traits, as well as their adherence to autistic profiles - including whether they fell above or below the diagnostic threshold - from their imageability ratings. We speculate that this quasi-perceptual ability to imagine theoretical concepts represents a specific cognitive pattern that, while hindering social interaction, may favor problem solving in abstract, non-socially related tasks. This would allow people with marked autistic traits to make use of perceptual, possibly visuo-spatial, information for "higher" cognitive processing.
In this paper, we investigate four important issues together for explicit discourse relation labelling in Chinese texts: (1) discourse connective extraction, (2) linking ambiguity resolution, (3) relation type disambiguation, and (4) argument boundary identification. In a pipelined Chinese discourse parser, we identify potential connective candidates by string matching, eliminate non-discourse usages from them with a binary classifier, resolve linking ambiguities among connective components by ranking, disambiguate relation types by a multiway classifier, and determine the argument boundaries by conditional random fields. The experiments on Chinese Discourse Treebank show that the F1 scores of 0.7506, 0.7693, 0.7458, and 0.3134 are achieved for discourse usage disambiguation, linking disambiguation, relation type disambiguation, and argument boundary identification, respectively, in a pipelined Chinese discourse parser.
Continuous word representations appeared to be a useful feature in many natural language processing tasks.Using fixed-dimension pre-trained word embeddings allows avoiding sparse bag-of-words representation and to train models with fewer parameters.In this paper, we use fixed pre-trained word embeddings as additional features for a neural scoring function in the MST parser.With the multi-layer architecture of the scoring function we can avoid handcrafting feature conjunctions.The continuous word representations on the input also allow us to reduce the number of lexical features, make the parser more robust to out-of-vocabulary words, and reduce the total number of parameters of the model.Although its accuracy stays below the state of the art, the model size is substantially smaller than with the standard features set.Moreover, it performs well for languages where only a smaller treebank is available and the results promise to be useful in cross-lingual parsing.
Hoarding is a complex and impairing psychiatric disorder and a public health problem. Traditionally it is assessed through observation and interview, but recently a new method has been proposed where living quarters of an individual are visually compared with a set of template images ranked according to the “Clutter Image Rating” (CIR) scale from 1 to 9. However, such an assessment is time-consuming, subjective, and weak in repeatability. We propose an automatic method for classifying hoarding images according to the CIR scale. Since clutter in living quarters (e.g., piles of boxes, newspapers, clothing) corresponds to “busy” areas with lots of edges in captured images, we use the histogram-of-gradients (HOG) descriptor to characterize images and estimate the CIR value using two methods: regression and classification. In 4-fold cross-validation on 620 images that we harvested from the internet, both methods result in mean-absolute CIR error of about 1.2. Given the simplicity of our method, this is an encouraging result as it approximates ratings by trained professionals who admit assigning CIR values within ± 1 CIR point.
Recently, these has been a surge on studying how to obtain partially annotated data for model supervision. However, there still lacks a systematic study on how to train statistical models with partial annotation (PA). Taking dependency parsing as our case study, this paper describes and compares two straightforward approaches for three mainstream dependency parsers. The first approach is previously proposed to directly train a log-linear graph-based parser (LLGPar) with PA based on a forest-based objective. This work for the first time proposes the second approach to directly training a linear graph-based parse (LGPar) and a linear transition-based parser (LTPar) with PA based on the idea of constrained decoding. We conduct extensive experiments on Penn Treebank under three different settings for simulating PA, i.e., random dependencies, most uncertain dependencies, and dependencies with divergent outputs from the three parsers. The results show that LLGPar is most effective in learning from PA and LTPar lags behind the graph-based counterparts by large margin. Moreover, LGPar and LTPar can achieve best performance by using LLGPar to complete PA into full annotation (FA).
A known way to improve the accuracy of dependency parsers is to combine several different parsing algorithms, in such a way that the weaknesses of each of the models can be compensated by the strengths of others. For example, voting-based combination schemes are based on variants of the idea of analyzing each sentence with various parsers, and constructing a combined output where the head of each node is determined by “majority vote” among the different parsers. Typically, such approaches combine very different parsing models to take advan- tage of the variability in the parsing errors they make. In this paper, we show that consistent improvements in accuracy can be obtained in a much simpler way by combining a single parser with itself. In particular, we start with a greedy implementation of the Nivre pseudo-projective arc-eager algorithm, a well-known left-to-right transition-based parser, and we combine it with a “mirrored” version of the algorithm that analyzes sentences from right to left. To determine which of the two obtained outputs we trust for the head of each node, we use simple criteria based on the length and position of dependency arcs. Experiments on several datasets from the CoNLL-X shared task and the WSJ section of the English Penn Treebank show that the novel combination system obtains better performance than the baseline arc-eager parser in all cases. To test the generality of the approach, we also perform experiments with a different transition system (arc-standard) and a different search strategy (beam search), obtaining similar improvements in all these settings.
Perirhinal cortex (PrC) has been implicated as a brain region in the medial temporal lobes (MTL) that critically contributes to familiarity-based recognition memory, a process that allows for recognition to occur independently of contextual recollection. Informed by neurophysiological research in non-human primates, fMRI, as well as behavioural work in humans, the current thesis research tests the novel hypothesis that PrC cortex functioning also underlies the ability to assess cumulative lifetime familiarity with object concepts that are characterized by a lifetime of experiences. In Chapter 2, a patient (NB) with a left anterior temporal lobe (ATL) lesion that included PrC as well as an amnesic patient (HC) with a bilateral lesion to the hippocampus were tested on their ability to make lifetime familiarity judgements for object concepts (i.e., concrete nouns). Patient NB made abnormal familiarity ratings for objects concepts relative to matched controls, while patient HC produced ratings that did not differ from control participants. In Chapter 3, I tested healthy young adults on a frequency judgement task and lifetime familiarity task while they underwent fMRI. A region in the left PrC tracked both the perceived frequency of recent laboratory exposure as well as perceived lifetime familiarity. Finally, in Chapter 4, I tested whether indeed lifetime familiarity judgements are based on conceptual processing by making use of an associative priming paradigm. Associatively-related primes increased the perceived familiarity of object concepts while also reducing the latency of these judgements. Overall, the results from all three empirical chapters provides evidence that warrants an extension of PrC functioning to include the cumulative assessment of lifetime familiarity with object concepts.
The OPT submission to the Shared Task \nof the 2016 Conference on Natural Language \nLearning (CoNLL) implements a \n‘classic’ pipeline architecture, combining \nbinary classification of (candidate) explicit \nconnectives, heuristic rules for non-explicit \ndiscourse relations, ranking and ‘editing’ \nof syntactic constituents for argument identification, \nand an ensemble of classifiers to \nassign discourse senses. With an end-toend \nperformance of 27.77 F1 on the English \n‘blind’ test data, our system advances \nthe previous state of the art (Wang & Lan, \n2015) by close to four F1 points, with particularly \ngood results for the argument identification \nsub-tasks. OPT system results appear \nmore competitive on the new, ‘blind’ \ntest data than on the ‘test’ and ‘development’ \nsections of the Penn Discourse Treebank \n(PDTB; Prasad et al., 2008), which \nmay indicate reduced over-fitting to specific \nproperties of the venerableWall Street \nJournal (WSJ) text underlying the PDTB
This thesis describes unexpected constructions based on THIS and THAT by French and Spanishlearners of English. Chapter 1 raises the issue of the study of THIS and THAT as markers in the twomicrosystems of proforms and deictics. Chapter 2 covers different types of analyses of referencewith THIS and THAT in native English and refers to different theoretical frameworks (Cornish,Cotte, Halliday & Hasan, Kleiber, Fraser & Joly, Lapaire & Rotgé). It crossreferencesrepresentations (anaphora/deixis; endophoricity/exophoricity) with an analysis of functionalrealisations. Chapter 3 broaches the issue of interlanguage analysis, and it shows that a dynamicsystemic approach grounded in the functional distinction of the forms is necessary. Chapter 4 givesdetails about existing annotation tagsets for English corpora (Penn Treebank, Claws7, ICEGB). Itshows the need for a finergrained annotation relying on functional tags and for semanticinformation on the positions (subject v. oblique). Chapter 5 describes the multilayer annotationstructure which is implemented for the analysis of different corpora. It also covers the methods usedto automatically annotate functional categories (as well as their evaluation), and it justifies thechoices made to support corpus interoperability. Chapter 6 offers a regression analysis whichprovides evidence on the tendencies of the operationalised variables (the L1, the written or spokenmode of the corpora and the type of reference). Chapter 7 examines the role of the previously codedlinguistic properties of the analysis. With the use of classifiers, it describes a system for automaticerror analysis. Chapter 8 concludes on the methodologies used in the thesis and their implicationsin linguistic analysis.
In many natural language processing based intelligent systems, parsing is the first task to perform. However, in the next stages, many systems often have the capacity of processing a limited number of parsed structures. The problem is to determine what parsed sentences can be recognized by a system. The decision of syntactic structures which can be processed by a system is consider as the task of "classification" of a parsed sentence into one of given classes of recognizable parses. In this paper we deal with this issue by proposing a method for mapping Vietnamese chunked sentences to a set of pre-defined shallow structures. Also, we tag lexicons and chunk phrases of the original sentences using our Functional Part-of-Speech (FPOS) tagset with Apache OpenNLP tools (Tokenizer, POS Tagger, Chunker). Based on the foundation of Functional Grammar, we define new lexical tags and combine with Penn-Treebank tagset to build our FPOS tagset. Due to our set of shallow structures is finite, instead of using a parser, we propose a rule-based algorithm for the mapping process. We establish conversion rules according to the reality experiences when using Vietnamese in common communication. The experiment shows that we converse successfully for the major of testing sentences and the algorithm can be applied for different languages.
This dataset introduces a companion reproducibility Java console program, called HESML_vs_SML_test.jar, of the work introduced by Lastra-Díaz and García-Serrano [1]. This latter work introduces the Half-Edge Semantic Measures Library (HESML), and carries-out an experimental survey between HESML V1R2, the Semantic Measures Library (SML) 0.9 [2] and the WNetSS [4] semantic measures libraries. The HESML_vs_SML_test.jar program runs the set of performance and scalability benchmarks detailed in [1] and generates the figures and tables of results reported in the aforementioned work, which are also enclosed as complementary files of this dataset (see files below). Licensing note: The 'HESML_vs_SML_test.jar' program is based on the HESML V1R2 [3], SML 0.9 [2] and WNetSS [4] semantic measures libraries, and it includes these libraries in its distribution, as well as WordNet 3.0 [6] and the SimLex665 [5] dataset. Thus, if you use this dataset, you should also cite the works related to these resources. References: [1] Lastra-Díaz, J. J., and García-Serrano, A. (2016). HESML: a scalable ontology-based semantic similarity measures library with a set of reproducible experiments and a replication dataset. To appear in Information Systems Journal. [2] Harispe, S., Ranwez, S., Janaqi, S., and Montmain, J. (2014). The Semantic Measures Library: Assessing Semantic Similarity from Knowledge Representation Analysis. In E. Métais, M. Roche, & M. Teisseire (Eds.), Proc. of the 19th International Conference on Applications of Natural Language to Information Systems (NLDB 2014) (Vol. 8455, pp. 254–257). Montpelier, France: Springer. http://dx.doi.org/10.1007/978-3-319-07983-7_37 [3] Lastra-Díaz, J. J., & García-Serrano, A. (2016). HESML V1R2 Java software library of ontology-based semantic similarity measures and information content models. Mendeley Data, v2. https://doi.org/10.17632/t87s78dg78.2 [4] Ben Aouicha, M., Taieb, M. A. H., and Ben Hamadou, A. (2016). SISR: System for integrating semantic relatedness and similarity measures. Soft Computing, 1–25. http://dx.doi.org/10.1007/s00500-016-2438-x [5] Hill, F., Reichart, R., & Korhonen, A. (2015). SimLex-999: Evaluating Semantic Models with (Genuine) Similarity Estimation. Computational Linguistics, 41(4), 665–695. http://dx.doi.org/10.1162/COLI_a_00237 [6] Miller, G. A. (1995). WordNet: A Lexical Database for English. Communications of the ACM, 38(11), 39–41. http://dx.doi.org/10.1145/219717.219748
The purpose of this paper is to discuss the social legitimacy of the non-dominant variety of French that is used in Belgium (henceforth ‘Belgian French’). As will be detailed, Francophone Belgians’ attitudes have shifted from early 19th c. – late 20th c. purism and subsequent linguistic subjection to France to more recent acceptation of endogenous traits and increasing distance from the Hexagonal model. Nevertheless, these attitudes remain characterized by a “double distance” from both Hexagonal and Belgian French. The idea that French is viewed by Francophone Belgians as a polycentric/polynomic language will thus be questioned: do they really consider that there is a legitimate Belgian variety of French? What is the relevance of the national criterion in the way they define linguistic norms? What other criteria lie behind the definition and legitimization of their linguistic norms?
Nonostante una secolare tradizione lessicografica, la lingua latina manca ancora di risorse lessicali di tipo computazionale aggiornate allo stato dell’arte. Cio e strettamente connesso alla limitata disponibilita di corpora testuali latini annotati linguisticamente, sulla cui base empirica possano essere costruite nuove risorse lessicali. Tuttavia, una serie di progetti mirati allo sviluppo di avanzate risorse linguistiche per il latino (tra cui alcune treebank) e stata avviata nel corso dell’ultimo decennio. In questo articolo, presentiamo Latin Vallex, un lessico di valenza per il latino realizzato in stretta connessione con l’annotazione semantico-pragmatica di due treebank latine comprensive di testi di epoche e generi diversi. Cio consente di connettere biunivocamente le strutture valenziali registrate nel lessico e le loro occorrenze nei dati testuali delle treebank.
Tokenizer, POS Tagger, Lemmatizer and Parser models for all 50 languages of Universal Depenencies 2.0 Treebanks, created solely using UD 2.0 data (http://hdl.handle.net/11234/1-1983). The model documentation including performance can be found at http://ufal.mff.cuni.cz/udpipe/users-manual#universal_dependencies_20_models. To use these models, you need UDPipe binary version at least 1.2, which you can download from http://ufal.mff.cuni.cz/udpipe. In addition to models itself, all additional data and value of hyperparameters used for training are available in the second archive, allowing reproducible training.
We describe results of investigation of a specific type of discontinuous constructions, namely non-projective constructions concerning verbs and their arguments. This topic is especially important for languages with a relatively free word order, such as Czech, which is the language we have primarily worked with. For comparison, we have included some results for English. The corpora used for both languages are the Prague Czech-English Dependency Treebank and the Prague Dependency Treebank, which are both annotated at a dependency syntax level as well as a deep (semantic) level, including verbs and their valency (arguments). We are using traditionally defined non-projectivity on trees with full linear ordering, but the two levels of annotation are innovatively combined to determine if a particular (deep) verb -argument structure is non-projective. As a result, we have identified several types of discontinuities, which we classify either by the verb class or structurally in terms of the verb, its arguments and their dependents. In addition, we have quantitatively compared selected phenomena found in Czech translated texts (in the PCEDT) to the native Czech as found in the original Prague Dependency Treebank.
Nonostante una secolare tradizione lessicografica, la lingua latina manca ancora di risorse lessicali di tipo computazionale aggiornate allo stato dell’arte. Ciò è strettamente connesso alla limitata disponibilità di corpora testuali latini annotati linguisticamente, sulla cui base empirica possano essere costruite nuove risorse lessicali. Tuttavia, una serie di progetti mirati allo sviluppo di avanzate risorse linguistiche per il latino (tra cui alcune treebank) è stata avviata nel corso dell’ultimo decennio. In questo articolo, presentiamo Latin Vallex, un lessico di valenza per il latino realizzato in stretta connessione con l’annotazione semantico-pragmatica di due treebank latine comprensive di testi di epoche e generi diversi. Ciò consente di connettere biunivocamente le strutture valenziali registrate nel lessico e le loro occorrenze nei dati testuali delle treebank.
Universal Dependencies (UD) are gaining much attention of late for systematic evaluation of cross-lingual techniques for crosslingual dependency parsing. In this paper we present our work in line with UD. Our contribution to this is manifold. We extend UD to Indian languages through conversion of Pān inian Dependencies to UD for the Hindi Dependency Treebank (HDTB). We discuss the differences in annotation in both the schemes, present parsing experiments for both the formalisms and empirically evaluate their weaknesses and strengths for Hindi. We produce an automatically converted Hindi Treebank conforming to the international standard UD scheme, making it useful as a resource for multilingual language technology.
Dataset for TACL submission "The Galactic Dependencies Treebanks: Getting More Data by Synthesizing New Languages".<br> The scripts and model parameters for replicating this dataset are available at https://github.com/gdtreebank/gdtreebank.
The article deals with the phenomenon of diglossia in the context of development of Greek language in the Hellenistic period, observes in diachrony the correlation between the linguistic norm of Atticism and Koine Greek – from the beginning of theory of Atticist mimesis to the times of second sophistry; determines the specificity of Koine Greek of the New Testament and its relation to, on the one hand, Atticism canon and literary Koine Greek, and on the other hand, – to the language of the early Patristic literature.
espanolCon el desarrollo de la informatica, en la investigacion del lenguaje se introdujo la teoria y metodologia de redes complejas, que transforma el sistema de la lengua en las redes complejas compuestas de nodos y enlaces para hacer un analisis cuantitativo de la estructura de la lengua. El desarrollo de la gramatica de dependencias proporciona un apoyo teorico a la construccion del corpus anotado (treebank), por lo que el analisis estadistico con las redes complejas se hace posible. Este articulo presenta la teoria y metodologia de las redes complejas y construye las redes sintacticas de dependencia a base del corpus anotado (treebank) de las expresiones orales del examen EEE-4 (Examen del Espanol como Especialidad - Nivel 4). Mediante el analisis de las caracteristicas generales de las redes, incluyendo el numero de nodos, los enlaces, el grado medio, la longitud media de los caminos, la distribucion de grados y la centralizacion, tiene como objetivo descubrir la diferencia y similitud potencial entre las expresiones orales de distintos niveles. Ademas, con el analisis de conglomerados, esta investigacion pretende demostrar la capacidad discriminatoria de las variables de las redes complejas y proporcionar una referencia potencial para el trabajo de calificacion. EnglishWith the development of information technology, the theory and methodology of complex network has been introduced to the language research, which transforms the system of language in a complex networks composed of nodes and edges for the quantitative analysis about the language structure. The development of dependency grammar provides theoretical support for the construction of a treebank corpus, making possible a statistic analysis of complex networks. This paper introduces the theory and methodology of the complex network and builds dependency syntactic networks based on the treebank of speeches from the EEE-4 oral test. According to the analysis of the overall characteristics of the networks, including the number of edges, the number of the nodes, the average degree, the average path length, the network centrality and the degree distribution, it aims to find in the networks potential difference and similarity between various grades of speaking performance. Through clustering analysis, this research intends to prove the network parameters’ discriminating feature and provide potential reference for scoring speaking performance.
In recent years there has been a lot of interest in cross-lingual parsing for developing treebanks for languages with small or no annotated treebanks. In this paper, we explore the development of a cross-lingual transfer parser from Hindi to Bengali using a Hindi parser and a Hindi-Bengali parallel corpus. A parser is trained and applied to the Hindi sentences of the parallel corpus and the parse trees are projected to construct probable parse trees of the corresponding Bengali sentences. Only about 14% of these trees are complete (transferred trees contain all the target sentence words) and they are used to construct a Bengali parser. We relax the criteria of completeness to consider well-formed trees (43% of the trees) leading to an improvement. We note that the words often do not have a one-to-one mapping in the two languages but considering sentences at the chunk-level results in better correspondence between the two languages. Based on this we present a method to use chunking as a preprocessing step and do the transfer on the chunk trees. We find that about 72% of the projected parse trees of Bengali are now well-formed. The resultant parser achieves significant improvement in both Unlabeled Attachment Score (UAS) as well as Labeled Attachment Score (LAS) over the baseline word-level transferred parser.
This volume takes as its central organizing principle the foundational understanding about community knowledge that challenges “narrow conceptions of language, literacy, personal stories, bounded or contained learning contexts (e.g., home, community, schools), hegemonic cultural and linguistic norms, quantitative and static views of ‘resources,’ and limited attention to the agency, identities, and strategic actions of diverse students and their families as they traverse contexts” (daSilva Iddings, this volume). The touchstone to this approach is “Funds of Knowledge” as it has been conceptualized for nearly twenty-five years. This chapter will briefly summarize the approach as it has evolved and will lay out the programmatic implementation as it unfolded within CREATE.
Due to the constant increasing of electronic textual information, modern society needs for the automatic processing of natural language (NL). The main purpose of NL automatic text processing systems is to analyze and create texts and represent their content. The purpose of the paper is the development of linguistic and software bases of an automatic system for processing English publicistic texts. This article discusses the examples of different approaches to the creation of linguistic databases for processing systems. The author gives a detailed description of basic building blocks for a new linguistic processor: lexicalsemantic, syntactical and semantic-syntactical. The main advantage of the processor is using special semantic codes in the alphabetical dictionary. The semantic codes have been developed in accordance with a lexical-semantic classification. It helps to precisely define semantic functions of the keywords that are situated in parsing groups and allows the automatic system to avoid typical mistakes. The author also represents the realization of a developed linguistic database in the form of a training computer program.
The study examined the relationship between idiom familiarity, knowledge of idiom meaning and idiom transparency judgments in L2. A group of 23 intermediate Japanese learners of English were asked to provide familiarity ratings, transparency judgments, and definitions for 30 English idioms, 27 of which had semantically equivalent but compositionally different idiomatic counterparts in Japanese and 3 phrases for which semantic equivalents in L1 also shared the same structural properties. Transparency ratings were repeated after the instructional treatment. A comparison of pre-treatment and post-treatment transparency scores showed that knowledge of conventional idiom meanings had a strong effect on the learners’ perceptions of idiom transparency. Transparency judgments, however, were not found to be a reliable predictor of the learners’ ability to infer figurative meanings of the idiomatic phrases. Idiom familiarity was not found to have a significant effect on idiom comprehension or on transparency judgments either. A limited positive effect of language transfer on L2 idiom comprehension and transparency ratings was observed.
Facial expressions frequently involve multiple individual facial actions. How do facial actions combine to create emotionally meaningful expressions? Infants produce positive and negative facial expressions at a range of intensities. It may be that a given facial action can index the intensity of both positive (smiles) and negative (cry-face) expressions. Objective, automated measurements of facial action intensity were paired with continuous ratings of emotional valence to investigate this possibility. Degree of eye constriction (the Duchenne marker) and mouth opening were each uniquely associated with smile intensity and, independently, with cry-face intensity. In addition, degree of eye constriction and mouth opening were each unique predictors of emotion valence ratings. Eye constriction and mouth opening index the intensity of both positive and negative infant facial expressions, suggesting parsimony in the early communication of emotion.
We tackle the challenge of learning part-of-speech classified translations as part of an inversion transduction grammar, by learning translations for English words with known part-of-speech tags, both from existing translation lexica and from parallel corpora. When translating from a low resource language into English, we can expect to have rich resources for English, such as treebanks, and small amounts of bilingual resources, such as translation lexica and parallel corpora. We solve the problem of integrating these heterogeneous resources into a single model using stochastic Inversion Transduction Grammars, which we augment with wildcards to handle unknown translations.
In this paper, the importance of “culture” is focused on regarding its relation to second/foreign language acquisition/learning. By reinterpreting the Iceberg Model of Culture, the author thinks that second/foreign language learners are exposed to the dominant culture with its social and linguistic norms and therefore they experience deculturalization, which also brings about the issue of the Self and the Other. It is suggested in the paper that a shift from communicative competence (CC) to intercultural communicative competence (ICC) in multicultural second/foreign language classes per se could enhance language learning by involving these learners’ native culture elements in the language learning/teaching process. While the paper illustrates some pedagogical implications that such a shift could entail, it concludes that introducing these practices in multicultural second/foreign language classes have the potential to enable both practitioners and learners to deal with power relations and the deculturalizing forces, which might be prevalent in such classes.
Facial expressions are one of the most important types of non-verbal communication. Although interpretation of facial expressions is usually robust, studies have shown that both age-related and disease-related factors can influence recognition accuracy. In particular, older people show deficits in recognition of negative expressions. Similarly, patients suffering from Parkinson's disease (PD) also show impairments in recognition of fear, anger, and disgust expressions. These studies so far have only focused on the basic, or "universal" expressions. Here, we were interested in investigating and comparing the effects of age and disease on facial expression processing for a wider range of both emotional and communicational expressions. For our ongoing study we recruited a total of 79 participants: 20 PD patients, 15 age-matched, older healthy controls (HC), and 44 younger healthy controls (HCS). During the experiment, participants were instructed to watch videos of 27 facial expressions performed by 6 different actors and to rate each expression based on 12 evaluative dimensions (arousal, valence, naturalness, politeness, persuasiveness, dynamic, familiarity, empathy, honesty, attractiveness, intelligence, and outgoingness) using a 7-point Likert scale. Ratings were analyzed using within-group and across-group correlations, factor analysis, and item analyses. Overall, we found that ratings of expressions were more different due to age, than due to disease-prevalence: r(PD/HC)=.756 versus r(PD/HCS)=.627, r(HC/HCS)=.640. Three out of six factors in the factor analysis were common for all groups (arousal-dynamic, familiarity-empathy, and naturalness-sincerity), showing common evaluation patterns. Confirming earlier findings of a "positivity effect", valence ratings of negative expressions were higher for both older groups (although valence ratings highly correlated within-group: all r>.919). Similarly, negative expression were perceived as more natural but less persuasive by both older groups. Overall, our results show that age-related factors play a much larger role than PD-related factors in processing of both emotional and communicational facial expressions. Meeting abstract presented at VSS 2016
Con el desarrollo de la informática, en la investigación del lenguaje se introdujo la teoría y metodología de redes complejas, que transforma el sistema de la lengua en las redes complejas compuestas de nodos y enlaces para hacer un análisis cuantitativo de la estructura de la lengua. El desarrollo de la gramática de dependencias proporciona un apoyo teórico a la construcción del corpus anotado (treebank), por lo que el análisis estadístico con las redes complejas se hace posible. Este artículo presenta la teoría y metodología de las redes complejas y construye las redes sintácticas de dependencia a base del corpus anotado (treebank) de las expresiones orales del examen EEE-4 (Examen del Español como Especialidad - Nivel 4). Mediante el análisis de las características generales de las redes, incluyendo el número de nodos, los enlaces, el grado medio, la longitud media de los caminos, la distribución de grados y la centralización, tiene como objetivo descubrir la diferencia y similitud potencial entre las expresiones orales de distintos niveles. Además, con el análisis de conglomerados, esta investigación pretende demostrar la capacidad discriminatoria de las variables de las redes complejas y proporcionar una referencia potencial para el trabajo de calificación.
OBJECTIVE: Examine the production of abstract and concrete nouns in patients with neurodegenerative disease using the Cookie Theft picture description. BACKGROUND: In a previous study, we observed a double dissociation in abstract and concrete word knowledge between the semantic variant of primary progressive aphasia (svPPA) and the behavioral variant of frontotemporal dementia (bvFTD). Compared to age-matched healthy controls, bvFTD patients were significantly more impaired for abstract nouns than for concrete nouns, and their poor abstract knowledge related to atrophy in the inferior frontal gyrus. In contrast, svPPA patients were significantly more impaired for concrete nouns compared to abstract nouns, associated with atrophy to the left temporal lobe. In this study, we test if the same neural regions that are critical for effective abstract and concrete word comprehension also play a role in the production of abstract and concrete words. DESIGN/METHODS: We assess the production of abstract and concrete nouns in 42 bvFTD and 21 svPPA patients using an oral description of the Cookie Theft picture. Patients met published diagnostic criteria, and concreteness or abstractness of each word was calculated using Brysbaert concreteness ratings (Brysbaert, Warriner & Kuperman, 2014). RESULTS: We observe the same double dissociation pattern during production as we previously saw with comprehension: bvFTD patients produce a smaller proportion of abstract nouns than svPPA patients. Moreover, the average concreteness rating for all nouns produced by svPPA patients is lower than for bvFTD. Regression analyses demonstrated that decreased abstract noun production in bvFTD relates to atrophy in the left inferior frontal gyrus. In addition, decreased concrete noun production in svPPA relates to atrophy in the left inferior temporal lobe. CONCLUSIONS: These results corroborate the finding that abstract and concrete nouns are represented in partially dissociable anatomic regions.
Recent work on neural network models shows success in dependency parsing. In this paper, we present a sequence learning dependency parsing (SLDP) model using long short-term memory for shift-reduce parser. A feed-forward neural network is used to build greedy model from rich local features. With the features extracted by the local model, we further train a long short-term memory (LSTM) model optimized for global parsing sequences. Our model has the capability of learning not only atomic feature combinations automatically but also the long distance dependent information for dependency parsing. Experiments on English Penn Treebank show that our SLDP model significantly outperforms the baseline, achieving 90.7% unlabeled attachment score and 89.0% labeled attachment score.
Статтю присвячено аналізу комунікативного простору ЗМІ на предмет порушення мовних норм. Увагу зосереджено на мовному рівні сучасної інформаційної продукції. З’ясовано типологію помилок у текстах ЗМІ та причини їх виникнення. Наголошено на умінні учасників комунікативного процесу правильно користуватися засобами рідної мови, нормами, будувати висловлювання з урахуванням умов спілкування. Відповідальність за належне мовне оформлення повідомлення поділяє разом з його автором (журналістом) і редактор ЗМК, який цю інформацію оприлюднює. The article is devoted the analysis of communicative space of MASS-MEDIA for the purpose violation of linguistic norms. Attention concentrated at linguistic level of modern informative products. Tipologiyu of errors is found out in texts of MASS-MEDIA and reason of their origin. It is marked ability of participants of communicative process correctly to use facilities of the mother tongue, norms, to build an utterance taking into account the terms of intercourse. For a due linguistic registration of report divides responsibility together with his author (by a journalist) and editor ZMK, which promulgates this information.
Detecting automatically the cause relations of a text may be useful in question answering tasks and event information extraction. The aim of this paper is to study how to detect coherence relations of the cause subgroup (Cause, Result and Purpose). To achieve this aim we have used the Rhetorical Structure Theory (RST) and some automatic linguistic information from different tools developed by IXA Group. We have used a corpus of 60 scientific abstracts, the Basque RST Treebank (Iruskieta et al., 2013), of different domains: science, medicine and terminology. A linguist has annotated all the signals of that corpus and described the most important problems in such task. To report the reliability of this annotator, two linguists have annotated the signals of the cause subgroup and all the annotations were compared and evaluated. After that, a superannotator has harmonized all the signals of those cause relations. Finally, we show the most important signals for such relations.
情绪调节指的是个体对情绪的发生、体验与表达进行调控的能力和过程。良好的情绪调节能力有利于个体保持愉快的心境、改善不利的心境。现有的情绪调节研究大多都局限于负性情绪调节,而正性情绪调节的研究一直很少,当前研究综合行为和电生理方法,对被试在进行正性情绪认知重评调节时的唤醒度和颧肌肌电(zygomatic electromyography,简称zEMG)进行分析。结果发现两个实验中当进行正性情绪上调调节时,唤醒度和肌电指标相比维持调节时都有显著的增强,而下调与维持时肌电差异不显著。两个实验结果一致的证明正性情绪上调效应显著,而下调效应不显著。这表明相对于抑制正性情绪,人们可能更习惯和倾向于增强自身的正性情绪。 Emotion regulation refers to the ability and the process that individual adjusts and controls the occurrence, experience and expression of emotion. Good emotion regulation ability is beneficial for individuals to keep pleasant mood and improve the bad one. Although there are many studies investigating emotion regulation, they are mostly about negative emotion regulation, and few are known about the study of positive emotion regulation. In the current study, we use behavioral and electrophysiology measure of arousal rating and zygomatic electromyography (zEMG) to index the variation of positive emotion regulation. We collect arousal rating in each trial after participants regulate their emotion in Experiment 1 and the activity of zEMG in Experiment 2 in which the par-ticipants regulate their emotions induced by International Affective Picture System pictures. The results indicate that up-regulated effect of positive emotion regulation is significant, but down- regulated effect is not. This suggests that people may be desirable and habitual to increase their positive emotion rather than inhibit it.
Feedforward Neural Network (FNN)-based language models estimate the probability of the next word based on the history of the last N words, whereas Recurrent Neural Networks (RNN) perform the same task based only on the last word and some context information that cycles in the network. This paper presents a novel approach, which bridges the gap between these two categories of networks. In particular, we propose an architecture which takes advantage of the explicit, sequential enumeration of the word history in FNN structure while enhancing each word representation at the projection layer through recurrent context information that evolves in the network. The context integration is performed using an additional word-dependent weight matrix that is also learned during the training. Extensive experiments conducted on the Penn Treebank (PTB) and the Large Text Compression Benchmark (LTCB) corpus showed a significant reduction of the perplexity when compared to state-of-the-art feedforward as well as recurrent neural network architectures.
Semantic similarities are a cross-field research in Natural Language Processing and Ontologies with some possible fallout in Artificial Intelligence. Formerly, similarities were computed following a syntactical treatment to support case-based reasoning. Textual similarities are now guided by semantic machineries, offering various ways to compute relatedness measures. In this paper, we present both a logical and a visual framework aiming to reason with them. For that reason, we introduced FLH±, a fragment of description logic underpinning the well-known lexical database Wordnet. We illustrated this framework with the path length relatedness, one of the historical similarity measures occurring in a taxonomy. The core of our framework orchestrates the computation of similarity scores supported by REVERB, STANFORD CORENLP and WORDNET:SIMILARITY APIs and interfaces global similarities in graphical way by positioning them on segments. We also depicted some experimental results to confront our computational framework with some empirical data.
Detecting automatically the cause relations of a text may be useful in question answering tasks and event information extraction. The aim of this paper is to study how to detect coherence relations of the cause subgroup (CAUSE, RESULT and PURPOSE). TO achieve this aim we have used the Rhetorical Structure Theory (RST) and some automatic linguistic information from different tools developed by IXA Group. We have used a corpus of 60 scientific abstracts, the Basque RST Treebank (Iruskieta et al., 2013), of different domains: science, medicine and terminology. A linguist has annotated all the signals of that corpus and described the most important problems in such task. To report the reliability of this annotator, two linguists have annotated the signals of the cause subgroup and all the annotations were compared and evaluated. After that, a superannotator has harmonized all the signals of those cause relations. Finally, we show the most important signals for such relations.
The goal of language modeling techniques is to capture the statistical and structural properties of natural languages from training corpora. This task typically involves the learning of short range dependencies, which generally model the syntactic properties of a language and/or long range dependencies, which are semantic in nature. We propose in this paper a new multi-span architecture, which separately models the short and long context information while it dynamically merges them to perform the language modeling task. This is done through a novel recurrent Long-Short Range Context (LSRC) network, which explicitly models the local (short) and global (long) context using two separate hidden states that evolve in time. This new architecture is an adaptation of the Long-Short Term Memory network (LSTM) to take into account the linguistic properties. Extensive experiments conducted on the Penn Treebank (PTB) and the Large Text Compression Benchmark (LTCB) corpus showed a significant reduction of the perplexity when compared to state-of-the-art language modeling techniques.