Papers reviewed and determined not to be word norm studies. Use the flag icon to report errors or suggest re-inclusion.
16504 papers
Slang is one of the problems encountered in developing Lexical Database. There is no complete source officially for slangs in any language. This research attempts to create a list of Indonesian slangs using the proposed framework and methodology. The contents for slangs are generated by Twitter; therefore it can encounter the newest slangs available in the society.
The study is based on an experiment to measure the affective states of computer users via their use of mouse and keyboard. The experiment was replicated from a previous study by Khan et al., [5] resulting in significant correlations between the computer users pattern of interactions and their valence, arousal ratings. This study utilized the same data set from [5] and re-confirmed its validity by training Artificial Neural Networks (ANN). The data was divided into two portions for each individual. A portion to train ANN on his/her patterns of interaction and other portion to test the ANN. The study resulted in an average recognition rate of 64.72 % for valence and 61.02 % for arousal ratings. The highest recognition rates for individual participants' valence and arousal were 100% and 87% respectively. These figures suggest that ANN is a bright prospect for the measurement of affective states of individual computer users via their interaction with keyboard and mouse.
The paper analyzes 33 grammatical structures of Tsinghua University 973 Treebank from the syntactic point of view. Firstly, we explore the distribution of these structures in Tsinghua University 973 Treebank and analyze the syntactic constituents of these structures. Then, we gather statistics about these structures based on the external functional relation and the internal structural relation of its subcomponents. Finally, according to the same internal structural relations, we generate a matrix to show the ambiguity among these structures. The statistical data offers the syntactic knowledge for auto-identifying these structures in future use.
The paper presents a small empirical study into emotion and affect recognition based on auditory and visual features, which was performed in the context of the Audio-Visual Emotion Challenge (AVEC) 2012. The goal of this competition is to predict continuous-valued affect ratings based on the provided auditory and visual features, e.g., local binary pattern (LBP) features extracted from aligned face images, and spectral audio features.
Statistical machine translation has been remarkably successful for the world’s well-resourced languages, and much effort is focussed on creating and exploiting rich resources such as treebanks and wordnets. Machine translation can also support the urgent task of documenting the world’s endangered languages. The primary object of statistical translation models, bilingual aligned text, closely coincides with interlinear text, the primary artefact collected in documentary linguistics. It ought to be possible to exploit this similarity in order to improve the quantity and quality of documentation for a language. Yet there are many technical and logistical problems to be addressed, starting with the problem that – for most of the languages in question – no texts or lexicons exist. In this position paper, we examine these challenges, and report on a data collection effort involving 15 endangered languages spoken in the highlands of
Quantitative linguistics (QL) is a discipline of linguistics, that, using real texts, studies languages with quantitative mathematical approaches, aiming to precisely describe and explain, with a system of mathematical laws, the operation and development of language systems. Later in this review, we will address the relationship between QL and computational linguistics. Quantitative Syntax Analysis is a recent work on QL by Reinhard Köhler that not only provides a comprehensive introduction to the work of QL on the syntactic level, but also sketches the theoretical grounds, the research paradigm, and the ultimate goals of quantitative linguistics in general.In the first chapter, Köhler points to the vital role of syntax in language: Syntax enables language users to code structures instead of ideas as wholes. A text embodies a complex cognitive formation and meets several basic requirements in human communication, which implies that language is not autonomous, but a dynamic communicative system used by human beings. Hence, the ultimate understanding and explanation of syntax (and the whole language system) depends on usage-based investigation of the cognitive basis and the functional requirements of language, which is somewhat neglected in many mainstream syntactic studies. Even in those cases where explanatory power is acknowledged as the ultimate goal of linguistic investigation, the necessary knowledge is still required as to what a scientific linguistic theory is and how such a theory may be built. So far, it is rare for quantitative means to be used in syntactic study. One reason is that many syntacticians are too addicted to the enshrined traditional paradigms that have been proven to be somewhat inadequate when it comes to processing real texts. And that is why in computational linguistics, “devout executors of the belief in strictly formal methods as opposed to statistical ones do not have any chance to succeed” (page 4).The second chapter, entitled “The Quantitative Analysis of Language and Text,” begins with an explanation of the difference between quantitative linguistics and the formal branches of linguistics that have once been widely used in computational linguistics: QL is concerned with the quantitative properties important for understanding the development and the operation of linguistic systems, whereas the formal branches of linguistics use only qualitative mathematical means and formal logics to model structural properties of language, overlooking, in most cases, the aspects of systems that exceed structure, viz., functions, dynamics, and processes. Köhler points out that the successes of modern natural sciences (the exact, testable statements, the precise predictions, and the copious applications) all derive from their instruments and their advanced models. This implies that these instruments and models, for which the quantitative parts of mathematics (probability theory and statistics, function theory, differential equations) are indispensable ingredients, are worth integrating into linguistics, which is the aim of QL.Chapter 3, entitled “Empirical Analysis and Mathematical Modeling,” reviews the important works of quantitative syntactic analysis. In the first section, Köhler gives a long list of the important syntactic units and properties defined within the frameworks of both phrase structure syntax and dependency syntax; this reflects the fact that researchers in both fields have been engaged in some fruitful quantitative studies. In Section 3.2, he defines quantitation of syntactic concepts as counting the objects under study, because syntactic analysis investigates only discrete objects. Section 3.4 is a detailed review of the important works on various syntactic phenomena within the frameworks of both phrase structure syntax and dependency syntax, including sentence length, probabilistic grammars and probabilistic parsing, Markov chains, Frumkina’s law on the syntactic level, distribution of dependency distance, and distribution of dependency types, and so on. These quantitative models, which have been empirically corroborated with real texts, or sometimes treebanks and dictionaries (of various languages), can be linguistically, cognitively, or functionally interpreted—a rare achievement in the past statistical investigations of language. Apart from the models concerning probabilistic grammars and Markov chains, which have already been widely used in computational linguistics, there are some other works that may also have practical applications in various fields. For example, the mathematical model of sentence length, which describes the probability of neighboring length classes as a function of the probability of the first of the two given classes, may contribute to practical applications such as text classification and the measurement of text comprehensibility, and so forth. The frequency studies of word and syntactic constructions have obtained many results useful for language teaching, the construction of parsing algorithms, and the estimation of effort of (automatic) rule learning, and more. The syntactic studies on Frumkina’s law, which is concerned with the number of text blocks with x occurrences of a given syntactic element or category, may benefit certain types of computational text processing if specific constructions or categories can be differentiated and found automatically by their particular distributions. One advantage of QL is that all its findings are mathematically formulated and linguistically interpreted, which at least makes it possible to be used in constructing models necessary for computational linguistics.In Chapter 4, Köhler introduces his efforts to build a real “linguistic theory,” for he believes that “there is not yet any elaborated linguistic theory in the sense of the philosophy of science” (page 21). Building such a theory begins with “plausible hypotheses,” which may become laws when sufficiently attested and may then be further integrated into a coherent system. This is the process of setting up a scientific theory, as succinctly summarized in the title of this chapter: “Hypotheses, Laws, and Theories.”The first section of this chapter shows the first step toward a scientific linguistic theory, the process in which “plausible hypotheses” are deduced, interpreted, and empirically attested before finally becoming laws. In Section 2, Köhler introduces the foundation of his synergetic linguistics, which views language as a dynamic, self-organizing, and self-regulating system where the so-called enslaving principle and order parameters are the crucial elements. On this basis, the author builds a synergetic syntactic model in Section 4.2.7 with certain modeling principles. Within the framework of phrase-structure syntax, eight properties of syntactic constructions and four inventories are chosen to build this model, which are linked together by laws resulting from the verified hypotheses and subject to the regulation of some order parameters.Quantitative linguistics, which is an unfamiliar field of study for many linguists, depends heavily on real texts and mathematical tools. Therefore we believe it is worthwhile to briefly clarify the differences and the relations between QL and corpus linguistics, on one hand, and between QL and computational linguistics, on the other hand. In comparison with QL, corpus linguistics is in fact more of a research methodology rather than an independent linguistic discipline, reflecting a shift of focus from competence to performance, from introspection to empirical study. This makes both the common ground and the difference between corpus linguistics and QL, which aims to quantitatively and mathematically explore, on the basis of real texts and treebanks, the fundamental laws governing the structure and evolution of language, and integrate them into a systematic theory capable of explanation and prediction.Computational linguistics (CL) is an interdisciplinary field investigating the structure of natural languages from a formal, mathematical, and computational point of view. Compared with QL, CL seems to be more interested in research that can have direct applications in such fields as language understanding, language generation, machine translation, and so forth, than in the explanation of language structure, operation, and evolution. Traditionally, the mathematical models used in CL are derived from the linguistic theories via the process of formalization. Due to the limitations of the qualitative mathematical models, however, the statistical models, most of which so far fail linguistic interpretations, have recently become dominant in the field of computational linguistics. That is perhaps why Shuly Wintner (2009, page 641), in a “Last Words” article in this journal, called for “the return of linguistics to computational linguistics,” implying that the advances in the field of CL may ultimately lie in the advances in the understanding of language itself.Wintner’s appeal does reflect the present situation of computational linguistics: a discipline in which many works are heavily oriented towards engineering and weakly grounded in linguistics. The formal linguistic models seem now outshone by the purely statistical paradigms, as the result of their inadequacy in processing real-world languages. Of course, this means no renouncement of the value of the traditional formal models. But it is obvious that considerable updating and enrichment are necessary for these formal models if they are to play significant roles in the future. In this regard, QL, which has a solid linguistic foundation, may help by providing quantitative cues (which are linguistically interpretable) to improve the performance in NLP, as has been illustrated in the case of probabilistic grammar that ingeniously integrates quantitative, statistical devices into qualitative linguistic models. QL has provided some useful models for CL. And it is reasonable to believe that it will continue to do so in the future, though some of its achievements seem now not ready to be directly used in CL.But perhaps QL can do more than that. Though QL has not yet drawn much attention from computational linguistics, it is potentially a proper answer to Wintner’s call for the return of linguistics to computational linguistics. CL aims to replicate in computers the patterns of human language behavior, whereas QL endeavors to mathematically and quantitatively reveal the laws and the principles that govern human language behavior—this is a relation between theory and practice. The success of a scientific discipline is usually based on precise models that are well-grounded in profound understanding of the object of study. Computational linguistics is no exception. Two features of QL are hence noteworthy. One is that it aims to explain, within a certain linguistic framework, the operation and the evolution of language by uncovering the systematic cognitive and functional regulations that underlie human languages. The other is that it tries to mathematically model, with systematic and precise quantitative laws, these regulations and the resulting mechanism of language. In view of these two features, we believe that the success of QL will somehow and somewhat boost the studies in the field of CL and that the communication between QL and CL is and will be not only possible but also mutually beneficial. This is why we hold that it is worthwhile to recommend this book to researchers in the field of computational linguistics, a book presenting a panorama of QL in general and discoveries on the syntactic level in particular.
An investigation of the relationship between background knowledge and reading comprehension performance on standardized reading tests (the California STAR Test) was conducted with sixth, seventh, and eighth-grade ethnic minority children from low-income backgrounds (N = 68). Predictor variables examined included perceived background knowledge (overall and topic-specific), GPA, basic literacy skills, reading self-concept, race and ethnicity, language background, gender, and grade level. Research questions addressed participants' familiarity with topics discussed in STAR test reading passages and about the predictive nature of participant rankings and ratings of passages, as measured by the Topic Familiarity Ranking Measure and the Topic Familiarity Rating Scale. Results indicated that background knowledge of passage topics had a significant positive association (p <.05) with reading comprehension performance for 30% of the CST passages, for seventh and eighth-grade participants. Hierarchical regression analyses conducted on three of the passages showed that between 7% and 16% of the variance in reading comprehension performance was accounted for by background knowledge, as measured by the Topic Familiarity Rating Scale.
the effect of pathological aging on explicit memory is very well documented, but relatively few studies have addressed this issue in the musical domain. To examine learning and consolidation of melodies, we designed a melodic recognition task involving immediate and delayed recognition of 16 target melodies (8 familiar and 8 unfamiliar). Seventeen patients with mild to moderate Alzheimer's disease (AD) and 17 age-matched controls were tested. During the initial presentation of the targets, the participant had to decide whether or not the melody was familiar. Recognition was tested after one and three presentations of the target melodies using a yes/no recognition paradigm. Delayed recognition was tested after 24 hours to evaluate consolidation. In keeping with the findings of Bartlett, Halpern, and Dowling (1995), age-matched controls showed better recognition of familiar than unfamiliar melodies. Controls also showed improved performance with multiple presentations for both familiar and unfamiliar melodies, without forgetting after 24-hour delay. In contrast, patients with AD showed impaired learning and recognition of both unfamiliar and familiar melodies with no benefit of familiarity on recognition. Nevertheless, the familiarity decision-based ratings of patients was in keeping with controls. These findings suggest that musical recognition memory is impaired in AD, but the musical lexicon (as assessed by familiarity ratings) is preserved. These findings highlight the need to use both familiar and unfamiliar music in experimental tasks to study the different processes underlying recognition memory.
Dependency parsing has attracted considerable interest from researchers and developers in natural language processing. However, to obtain a high‐accuracy dependency parser, supervised techniques require a large volume of hand‐annotated data, which are extremely expensive. This paper presents a simple and effective approach for improving dependency parsing with subtrees derived from unannotated data, which are easy to obtain. First, we use a baseline parser to parse large‐scale unannotated data. Then, we extract subtrees from dependency parse trees in the auto‐parsed data. Next, the extracted subtrees are classified into several sets according to their frequency. Finally, we design new features based on the subtree sets for parsing algorithms. To demonstrate the effectiveness of our proposed approach, we conduct experiments on the English Penn Treebank and Chinese Penn Treebank. The results show that our approach significantly outperforms baseline systems. It also achieves the best accuracy for the Chinese data and an accuracy competitive with the best known systems for the English data.
The paper concentrates on which language means may be included into the annotation of discourse relations in the Prague Dependency Treebank (PDT) and tries to examine the so called alternative lexicalizations of discourse markers (AltLex’s) in Czech. The analysis proceeds from the annotated data of PDT and tries to draw a comparison between the Czech AltLex’s from PDT and English AltLex’s from PDTB (the Penn Discourse Treebank). The paper presents a lexico-syntactic and semantic characterization of the Czech AltLex’s and comments on the current stage of their annotation in PDT. In the current version, PDT contains 306 expressions (within the total 43,955 of sentences) that were labeled by annotators as being an AltLex. However, as the analysis demonstrates, this number is not final. We suppose that it will increase after the further elaboration, as AltLex’s are not restricted to a limited set of syntactic classes and some of them exhibit a great degree of variation. Key words: alternative lexicalization of discourse markers (AltLex); discourse connectives; discourse relations 1.
The common use of a single de facto standard annotation scheme for dependency treebank creation leaves the question open to what extent the performance of an application trained on a treebank depends on this annotation scheme and whether a linguistically richer scheme would imply a decrease of the performance of the application. We investigate the effect of the variation of the number of grammatical relations in a tagset on the performance of dependency parsers. In order to obtain several levels of granularity of the annotation, we design a hierarchical annotation scheme exclusively based on syntactic criteria. The richest annotation contains 60 relations. The more coarse-grained annotations are derived from the richest. As a result, all annotations and thus also the performance of a parser trained on different annotations remain comparable. We carried out experiments with four state-of-the-art dependency parsers. The results support the claim that annotating with more fine-grained syntactic relations does not necessarily imply a significant loss of accuracy. We also show the limits of this approach by giving details on the fine-grained relations that do have a negative impact on the performance of the parsers.
We show that orthographic cues can be helpful for unsupervised parsing. In the Penn Treebank, transitions between upper- and lower-case tokens tend to align with the boundaries of base (English) noun phrases. Such signals can be used as partial bracketing constraints to train a grammar inducer: in our experiments, directed dependency accuracy increased by 2.2% (average over 14 languages having case information). Combining capitalization with punctuation-induced constraints in inference further improved parsing performance, attaining state-of-the-art levels for many languages.
Title: Valency of verbs in the Prague Dependency Treebank Author: PhDr. Zdeňka Urešová Department: Institute of Formal and Applied Linguistics MFF UK Supervisor: Prof. PhDr. Eva Hajičová, DrSc. Abstract: This dissertation describes PDT-Vallex, a valency lexicon of Czech verbs, and its relation to the annotation of the Prague Dependency Treebank (PDT). The PDT-Vallex lexicon was created during the an- notation of the PDT and it is a valuable source of verbal valency information available both for linguistic research and for computer- ized natural language processing. In this thesis, we describe not only the structure and design of the lexicon (which is closely related to the notion of valency as developed in the Functional Generative De- scription of language) but also the relation between the PDT-Vallex and the PDT. The explicit and full-coverage linking of the lexicon to the treebank prompted us to pay special attention to diatheses; we propose formal transformation rules for diatheses to handle their surface realization even when the canonical forms of verb arguments as captured in the lexicon do not correspond to the forms of these arguments actually appearing in the corpus.
This paper presents a theoretical discussion about the use of Frame Semantics as corpora annotation paradigm. The objective of this paper is to evaluate the applicability of Frame Semantics theory and FrameNet paradigm for the semantic annotation of legal texts. The work presented in this paper is an initial step in the construction of a treebank for the Brazilian legal language.
The program part of the lexical database was developed as the Mozilla Firefox application (2005-2006); the processing software is designed as a modern lexicographic workstation. Since 2007, we have gradually complemented the database with the lexicographicallly relevant data (research project Creation of a Lexical Database of the Czech Language of the Beginning of the 21st Century, 2005-2011). We have focused predominantly on detailed treatment of the example part, which demonstrates the breadth of the collocability of the lemmas.
Ontology matching is a main step for integrating overlapping domains of knowledge and establishing interoperation among semantic web application. As information sources grow rapidly, manual ontology matching becomes more tedious and time-consuming and consequently leads to errors and frustration. In this paper we developed the new lexical and semantic similarity measure by using the lexical database ConceptNet. The proposed strategy used new lexical and semantic matching for finding the correspondence entities. In the semantic approach we use the electronic lexical database, ConceptNet for identifying the similar entities and create similarity matrices according to that. We evaluate the proposed measure using standard methods of precision and recall, tested on a well- known benchmark and also compared to other algorithms presented in the paper. The experimental results show the proposed algorithm is effective and outperforms other algorithms.
When porting parsers to a new domain, many of the errors are related to wrong attachment of out-of-vocabulary words. Since there is no available annotated data to learn the attachment preferences of the target domain words, we attack this problem using a model of selectional preferences based on domainspecific word classes. Our method uses Latent Dirichlet Allocations (LDA) to learn a domain-specific Selectional Preference model in the target domain using un-annotated data. The model provides features that model the affinities among pairs of words in the domain. To incorporate these new features in the parsing model, we adopt the co-training approach and retrain the parser with the selectional preferences features. We apply this method for adapting Easy First, a fast nondirectional parser trained on WSJ, to the biomedical domain (Genia Treebank). The Selectional Preference features reduce error by 4.5 % over the co-training baseline. 1
An initiative to model a Bangla to English (B2E) translation using Natural Language Processing (NLP) was proposed in our previous research. Here we implemented the model with a lot of modifications. A very successful translator is observed in Anubadok Online which is based on Penn Treebank annotation system and it can only translate English sentences to Bangla. Penn Tree Bank is the collection of English corpus, so Bangle linguistic processing is not observed there. Bangla is an Irregular Language. In our previous research, we proposed a case structure analysis for verb. There are a lot of influences of case in Bangla language. The relationship between verb and case elements is an important issue for Bangla language. But in our current implementation we used rule based approach. For Bangla-English translation first we performed morphological analysis for Bangla then we used rule based analysis where we considered a limited feature of case analysis. After that, using a dictionary we translate bangle words into English. To make an English sentence, we considered English SVO grammatical rules. Our current system is successfully implemented for the translation of Assertive-Affirmative, Negative and Interrogative sentences.
A novel method for hybrid graph-based dependency parsing of natural language text is proposed. It is based on k-best maximum spanning tree dependency parsing and evaluation of the spanning trees by using a verb valency lexicon for a given language as a reranking knowledge base. The approach is compared with existing state-of-the-art transition-based and graph-based approaches to dependency parsing. As the proposed generic method was developed specifically for improving the accuracy of Croatian dependency parsing, Croatian Dependency Treebank and CROVALLEX verb valency lexicon are used in the experiment. The suggested approach scored approximately 77.21 % LAS, outperforming the tested state-of-the-art approaches by at least 2.68 % LAS. TITLE AND ABSTRACT IN CROATIAN Ovisnosno parsanje pomoću k najboljih razapinjućih stabala i ponovnoga vrjednovanja valencijskim rječnikom glagola Predlaže se novi pristup hibridnom ovisnosnom parsanju tekstova prirodnoga jezika temeljenom na teoriji grafova. Pristup je zasnovan na ovisnosnom parsanju pomoću k najboljih razapinjućih
Features of Bingxin's translations are studied based on corpus techniques.The findings are listed below:(1) There is a general tendency of simplification in lexical use in Bingxin's translations,but some individual works don't show this tendency;the lexical density of translated works is not affected by the lexical density and register of the original works.(2) The use of high-frequency words tends to be consistent in the original Chinese texts and translated texts,however,the use of functional words tends to rise in translated texts.(3) Compared with the original Chinese texts,the average length of sentences tends to expand in translated texts with more pauses within a sentence.(4) The use of personal pronouns,conjunctions and prepositions in translated texts is right between the original Chinese texts and English texts but there is a tendency to closely follow Chinese textual norms.(5) The use of adjectives,adverbs and verbs in translated texts tends to overlap with that in the original Chinese texts.(6) Features of translated language are affected by both source texts and translator's style.
The social and cultural ‘turn’ in language education of recent years has helped move language teaching and curriculum design away from many of the more rigid dogmas of earlier generations, but the issue of the roles of the learners’ first language (L1) in language pedagogy and classroom interaction is far from settled. Some follow a strict ‘exclusive target language’ pedagogy, while others ‘resort to’ the use of the L1 for a variety of purposes (see ACTFL 2008). Underlying these competing views is the perspective of the L1 as an impediment to second language learning. Following sociocultural theory and ecological perspectives of language and learning and based on the findings of research on classroom code-switching and code choice, this paper lays out an approach to the language classroom as a multilingual social space in which learners and teacher study, negotiate, and co-construct code choice norms toward the dynamic, creative, and pedagogically effective use of both the target language and the learners’ L1(s). Learner use of the L1 for the purpose of grammatical or lexical learning is also considered, and some examples for instruction are offered.
We have developed EyeMap, a freely available software system for visualizing and analyzing eye movement data specifically in the area of reading research. As compared with similar systems, including commercial ones, EyeMap has more advanced features for text stimulus presentation, interest area extraction, eye movement data visualization, and experimental variable calculation. It is unique in supporting binocular data analysis for unicode, proportional, and nonproportional fonts and spaced and unspaced scripts. Consequently, it is well suited for research on a wide range of writing systems. To date, it has been used with English, German, Thai, Korean, and Chinese. EyeMap is platform independent and can also work on mobile devices. An important contribution of the EyeMap project is a device-independent XML data format for describing data from a wide range of reading experiments. An online version of EyeMap allows researchers to analyze and visualize reading data through a standard Web browser. This facility could, for example, serve as a front-end for online eye movement data corpora.
The present study explored different approaches for automatically scoring student essays that were written on the basis of multiple texts. Specifically, these approaches were developed to classify whether or not important elements of the texts were present in the essays. The first was a simple pattern-matching approach called “multi-word” that allowed for flexible matching of words and phrases in the sentences. The second technique was latent semantic analysis (LSA), which was used to compare student sentences to original source sentences using its high-dimensional vector-based representation. Finally, the third was a machine-learning technique, support vector machines, which learned a classification scheme from the corpus. The results of the study suggested that the LSA-based system was superior for detecting the presence of explicit content from the texts, but the multi-word pattern-matching approach was better for detecting inferences outside or across texts. These results suggest that the best approach for analyzing essays of this nature should draw upon multiple natural language processing approaches.
In many areas of the behavioral sciences, different groups of objects are measured on the same set of binary variables, resulting in coupled binary object × variable data blocks. Take, as an example, success/failure scores for different samples of testees, with each sample belonging to a different country, regarding a set of test items. When dealing with such data, a key challenge consists of uncovering the differences and similarities between the structural mechanisms that underlie the different blocks. To tackle this challenge for the case of a single data block, one may rely on HICLAS, in which the variables are reduced to a limited set of binary bundles that represent the underlying structural mechanisms, and the objects are given scores for these bundles. In the case of multiple binary data blocks, one may perform HICLAS on each data block separately. However, such an analysis strategy obscures the similarities and, in the case of many data blocks, also the differences between the blocks. To resolve this problem, we proposed the new Clusterwise HICLAS generic modeling strategy. In this strategy, the different data blocks are assumed to form a set of mutually exclusive clusters. For each cluster, different bundles are derived. As such, blocks belonging to the same cluster have the same bundles, whereas blocks of different clusters are modeled with different bundles. Furthermore, we evaluated the performance of Clusterwise HICLAS by means of an extensive simulation study and by applying the strategy to coupled binary data regarding emotion differentiation and regulation.
Context enables readers to quickly recognize a related word but disturbs recognition of unrelated words. The relatedness of a final word to a sentence context has been estimated as the probability (cloze probability) that a participant will complete a sentence with a word. In four studies, I show that it is possible to estimate local context–word relatedness based on common language usage. Conditional probabilities were calculated for sentences with published cloze probabilities. Four-word contexts produced conditional probabilities significantly correlated with cloze probabilities, but usage statistics were unavailable for some sentence contexts. The present studies demonstrate that a composite context measure based on conditional probabilities for one- to four-word contexts and the presence of a final period represents all of the sentences and maintains significant correlations (.25, .52, .53) with cloze probabilities. Finally, the article provides evidence for the effectiveness of this measure by showing that local context varies in ways that are similar to the N400 effect and that are consistent with a role for local context in reading. The Supplemental materials include local context measures for three cloze probability data sets.
En psychologie tout comme en traitement automatique des langues, les normes qui portent sur des proprietes semantiques des mots, comme le degre d’abstraction, l’imagerie ou la polarite, sont importantes. Ces normes ont systematiquement ete obtenues en demandant a des juges d’evaluer les mots sur des echelles, allant par exemple de tres concret a tres abstrait. Ce mode de recolte etant lent et couteux, des methodes de construction automatique ont vu le jour. Elles peuvent etre divisees en deux types: celles qui se basent sur des ressources linguistiques et celles qui se basent sur des corpus. Notre objectif est de comparer, pour une meme methode d’accroissement de normes lexicales basee sur les similarites entre les mots, l’utilisation d’un corpus et d’une ressource lexicale (WordNet) pour estimer ces similarites. Nous montrons que les similarites calculees a partir d’informations sur les cooccurrences des mots dans les textes sont plus efficaces, et ce pour 4 des 5 normes etendues. Nous montrons egalement que le choix du corpus influence peu les resultats, du moins pour des corpus generaux.
Cette thèse porte sur une étude de la variation et du changement lexicaux des mots référant aux notions de « véhicule automobile » et de « travail rémunéré » dans le français de l’Outaouais, une variété de français laurentien caractérisée par le bilinguisme équilibré et stable et le contact intense avec l’anglais. La thèse est réalisée dans le cadre de la sociolinguistique variationniste labovienne combinée avec des méthodes quantitatives et des techniques analytiques multivariationnelles des règles variables. Cette étude se base sur les données empiriques recueillies dans les communautés francophones de la région de la capitale canadienne parmi les locuteurs nés entre 1846 et 1994 (RFQ, Ottawa-Hull, FdO).\nLe chapitre 2 suit l’évolution sémantique des termes lexicaux, étudie un système d’interaction des facteurs historiques et ceux socialement motivés, et examine l’hypothèse du développement interne du vocabulaire du français canadien. Les chapitres 3 et 4 examinent la corrélation des facteurs liés au bilinguisme et au contact avec l’anglais avec la fréquence d’emploi des variables lexicales; et la marque sociale des variantes lexicales.\nCette thèse: i) met en valeur la méthodologie variationniste quantitative dans l’étude de la variation lexicale; ii) approfondit plusieurs réflexions théoriques et des patrons classiques sur la théorie variationniste; iii) caractérise le lien dynamique entre le parler des locuteurs et les normes de la communauté à laquelle ils se rattachent; iv) contribue à la meilleure compréhension de la dynamique lexicale en fonction du statut du français en situation de contact de langues.
Written texts, particularly published ones, are widely perceived to have legitimacy beyond that of the spoken word in literate societies. One reason perhaps is that written text generally has more staying power as a concrete and tangible documentation of thought, intention, information, agreements and so on than does the spoken word. Furthermore, in many cases written text assumes a larger, more unifi ed, identifi able audience that is refl ective of those who share some subset of cultural norms and values. With legitimacy and cultural norms as a backdrop, the reason for code-switching in written discourse becomes an interesting subject of inquiry. Why switch between languages in a medium where one has ample time and resources to produce a monolingual text per the expected norm? Researchers have identifi ed this phenomenon in texts ranging from blogs to historical documents and have proposed various accounts, some of which are presented in this volume. On the surface, the switches found in written text may look and read like typical oral code-switches, where two or more languages are used, at times inter-sententially, at times intra-sententially and occasionally intra-lexically, with bound and free morphemes of two (or more) languages collaborating to create a discourse.
This paper foregrounds one argument in Rawls's work that is crucial to his case for one, determinate, form of political economy: a property-owning democracy.1 Section one traces the evolution of this idea from the seminal work of Cambridge economist James Meade; section two demonstrates how a commitment to a property-owning democracy flows from Rawls's own principles; section three focuses on Rawls's striking critique of orthodox welfare state capitalism. This all sets the stage for an argument, presented in section four, from the complexity of economic interactions to the strategy of making markets fair in the only feasible way that they can be made fair, namely, by “patterning” their effects. Section five concludes by asking whether any scheme of this general type is a realistic form of utopianism for a society such as ours.Many early readers of A Theory of Justice took Rawls to be advocating a form of “Keynesian capitalist liberalism.”2 However, if we define a capitalist society as one where people who do not own capital work for wages paid to them by capitalists (those who exclusively hold property and other forms of capital), then Rawls's conception of a property-owning democracy would involve the rejection of capitalism. James Meade, the proximate influence on Rawls's ideas, was indeed a Keynesian. However, given the working definition of a capitalist society that I have noted, it seems that liberal Keynesianism can reasonably be characterized as anti-capitalist in both Meade's and Rawls's variants.Meade's conception of a property-owning democracy emerged when he sought new avenues for egalitarianism in Britain given that the achievements of the Attlee government of 1945 were receding into the past.3 His aim was to combine Keynesian demand management with the public ownership of natural monopolies and the institutions of a property-owning democracy. Meade further proposed educational reform, a publicly funded unit trust, and state investment funds to supply an unconditional basic income. In a break with the policies of the then Labour government, Meade believed that welfare state redistribution was a threat to overall economic efficiency. Furthermore, relying on trade unions to redress the balance between labour and capital generated constant inflationary pressure in a way that explained the perceived “failure” of post-war Keynesian demand management: [G]radually, as in our imperfectly competitive society separate groups learned to press their monopolistic bargaining powers to obtain each for itself the best possible share of the available income, the system broke down.4 Meade's new strategy for redressing the balance between labour and capital was to rely on market prices to protect individual liberties and economic efficiency, but to increase the bargaining strength of labour by giving workers capital: If private property were much more equally divided we should achieve the mixed citizen—both worker and property owner at the same time—to live in the ‘mixed economy’ of public and private enterprise. The ownership of private property could then fulfill its useful function of providing a basis for private enterprise and for individual security and independence without carrying with it the curse of social inequality as it now does.5 Rawls's adaptation of Meade's ideas contains a cleaner break with welfare state capitalism, particularly in his late, summative statement of his views in Justice as Fairness where welfare state capitalism is unequivocally described as unjust.6 In the first edition of A Theory of Justice Rawls stated that: The aim of the branches of government is to establish a democratic regime in which land and capital are widely though not presumably equally held. Society is not so divided that one small sector controls the preponderance of productive resources.7 The branches of government dealing with the economy are the Allocative branch that deals with externalities, competition, and anti-trust. The Stabilisation branch is the most Keynesian branch of government, concerned with demand management and full employment. The Transfer Branch ensures the payment of a decent social minimum compatible with economic efficiency overall, via a negative income tax. Finally, the Distributive Branch raises money for transfer and for the regulation of the top end of distributions via a flat rate expenditure tax and the imposition of inheritance tax.The upshot, then, is this: in Rawls's ideal property-owning democracy, markets operate in a context structured pervasively by fairness. The state intervenes not only to supply public goods and to counter negative externalities, but also to impose that which Rawls called “adjusted procedural justice.” Some effects are the unintended outcomes of intended behaviour; the “invisible hand” part of Rawls's view is that the market, of its nature, decentralises economic power and protects freedom of occupational choice. It does the former by protecting free association and free equality of opportunity and does the latter by giving rise to differential earnings.8The main difference between Rawls's and Meade's version of property-owning democracy, then, concerns the strategic role of progressive taxation. In the first edition of A Theory of Justice, Rawls states that steeply progressive taxes may very well be justified “given the injustice of existing institutions.”9 But in ideal theory the role of progressive taxation is minimal. Equally striking is the residual “invisible hand” role, in both Meade's and Rawls's ideal, played by markets. This is an important point to which I will return below. In order to explain why progressive taxation plays such a marginal role in the ideally just society we need a better grasp on how a property-owning democracy is justified by Rawls's principles interpreted as working together as an interlocking group.Which of Rawls's principles make the case for a property-owning democracy? All of them, but in different ways. One of the most interesting aspects of Justice as Fairness is that in his comparison of a property-owning democracy and welfare state capitalism, to the detriment of the latter, Rawls interprets the build up of private concentrations of wealth permitted in welfare state capitalism as a potential threat to basic liberty. In his earlier work, Rawls had indeed conceded that on any view that has permissible inequality, the equal basic liberties would be of different worth to different people. But that thought was not troubling if people had a broadly comparable fair value in their political liberties: their ability to hold office and to participate, broadly, in the political determination of office. The political liberties, here, are the gatekeeper for the liberties as a whole.10 Martin O'Neill has objected that this argument is overdone: there are a variety of insulation strategies that a liberal democracy can pursue that can prevent accumulations of private wealth influencing the political process. So if Rawls objects to welfare state capitalism not because of bad effects that it brings about directly, but on the grounds of a general exposure to a political risk that it does not prevent, that part of his argument is implausible.11This is not the place to discuss my disagreement with O'Neill in any detail, but in fact I think that Rawls was not only right to emphasise this argument, but also that he needed to do more within the ambit of his own theory to address it.12 The measures he actually suggests to protect the fair value of the political liberties are disappointingly thin. My answer, unsurprisingly, is that Rawls is here demonstrating that his first principle of equal basic liberty, if it is to be implemented in conjunction with the fair value proviso for the political liberties, demands implementation in a property-owning democracy. The latter is the only way to prevent the concentration of private wealth that will lead to illegitimate interference with the political process.13What of the first part of the second principle, governing equality of opportunity? Here the case for a property-owning democracy is even clearer when one notes that “property” in this phrase is standing in for capital as a whole, and human capital counts as a form of capital. Fair equality of opportunity requires an adequately funded and free public education system that brings everyone's marketable talents up to their full potential, given that education is a public good not likely to be promoted in an unconstrained capitalist market. This measure, if fully implemented, would have a transformatory effect on the labour market by substantially increasing the supply of qualified labour, thus reducing the unearned rents currently accruing to the limited supply of labour for particular occupations.14I think it is important to bear in mind that this restructuring of the labour market by the full implementation of measures genuinely designed to protect the fair value of the political liberties, liberty as a whole, and the fair equality of opportunity forms the context for the introduction of the difference principle.15 The distinctive way in which Rawls makes the labour market fair, namely, by structuring the context in which it operates in order to pattern its effects has been very insightfully highlighted by Paul Smith: The idea that the equalization of property ownership would transform the labour market, by equalizing bargaining power and eliminating the economic coercion to accept drudge jobs at low pay and thus forcing employers to make all jobs attractive, all things considered, is crucial to Rawls's idea that, in a competitive labour market located in a just basic structure, income inequalities would tend just to compensate the costs of different jobs, that is, tend to equality, all things considered.16 Smith believes that this explains some of the distinctive features of Rawls's egalitarian strategy: Economic equalization is more likely and reliably to be effected, as Rawls thinks, by institutions and policies that equalize bargaining power than by an egalitarian ethos restraining the exercise of unequal bargaining power (and egalitarian institutions and their distributional results are what, if anything, could produce an egalitarian ethos).17 I will say more about this Rawlsian strategy and its underlying rationale below. But the basic idea is to structure the labour market so that what looks like the introduction of special incentives under the difference principle works under a of that make such incentives tend to be This is crucial to any to the that such incentives within that counter to and lead to Rawls's theory as a to my point about the of Rawls's principles; the case for a property-owning democracy is I would that it is by the first principle, even if it were not and it only on the with the principle of fair equality of two principles would be by an unconstrained difference principle that not operate in a context structured by the of capital implemented by a property-owning democracy. is in of for the argument that only a equality that is then by his own critique of Rawlsian incentives can be in that The introduction of the difference principle would to the first principle and the value for the basic is in order to critique of we to all three principles as as a as and making an case for a property-owning then, for the case for a property-owning democracy from within Rawls's own But was he Rawls in about the of a property-owning democracy to the welfare state capitalism with which we are most critique of welfare state capitalism as capitalism the fair value of the political liberties, and it has some for equality of the policies to achieve that are not It very inequalities in the ownership of property and natural so that the of the economy and much of political in as the welfare may be and a decent social minimum the basic a principle of to economic and social inequalities is not I think this contains an interesting of and a critique of existing social The the of the that the idea of a property-owning democracy which made it the of the political from to the idea of a property-owning democracy is a between the of private in the form of and the here is it has been to different and to the and of the ideal of property-owning view was that the of private property the of the of political some have justified private property via its to but in the case of of capital the is, with security of and the of mind that that security of in of the ideal to Meade and to One of that is particularly of is not the to the in but also the security that it them in from the that is a standing risk for on income. In our from it is the who are to at The here, not just welfare but also the James who has with how it is to be is a case where the difference between and ideal is One of of welfare state capitalism is that it welfare and the of a welfare critique seems in a currently society where who are and are in that of their that people are without it is to impose justified by the fact that welfare to and social However, to Rawls's views it that the of comparison is a just society that has implemented his principles in the form of a property-owning democracy. the of the payment of the social minimum to make a a of at what role is there for welfare state in such an Rawls other than social security in its of providing believed that there was good to welfare state as such and that the context in which they operate is because it concentrations of wealth in private A property-owning democracy, on the other in the of such concentrations of wealth just as political aim to state capitalism an context I would a are in which a society to make a democratic to a society that is just by Rawls's But I think Rawls is right to about a society an between the very and a of the well who their to be by the of and even well welfare state Rawls's ideal of a society of free and equal is not compatible with the of a of and who are not of any scheme for likely to be even if they are of a decent social This is Rawls's that welfare state capitalism the of a principle of The under welfare state capitalism live in a society that is both and to just via any feasible democratic would like to one from Rawls's to on an of argument is a very interesting strategy in Rawls that does not on a property-owning democracy in but on a general way of that is to lead to a property-owning democracy, very much like I have in mind are Rawlsian that from the and complexity of a economy that the implementation of principles of can work only by structuring the context in which market so as to impose a pattern on their Justice Fairness Rawls his own with a of the latter general up fair and for fair individual and that all outcomes the are made Rawls that this view the to individual that may fair, but which are actually by concentrations of wealth that are to equality of the fair value of the political liberties, and so is in order to we to procedural This an of which Rawls as is one Justice as focuses first on the basic structure and on the to for all equally we rely on an of between principles to and principles that to particular between and this of is and are then free to their within the of the basic structure, in the that in the social system the to are in is a point about complexity this I it Rawls is we do not a realistic on the one the context in which a market operates in order to make its effects fair, and the of each in a market one at a to the same is a an has that Smith was right to the value by as generated by the underlying of economic in society that makes This any individual in a of Smith the of in which a in economic society as all Economic a for is best explained as in the of that both to and in progressive to Smith also the differential talents and in this system as of education and that from ideas are that it is to individual productive The income that from talents is an from a of social that of the economy is by working The for an that are the of in a that is a in a of This of argument seems to of the particular form that Rawls's egalitarian strategy it is not feasible to individual market because the very idea of an individual market is a is why Rawls believed that the only way to make a market fair is to make its effects achieve that by structuring the context in which that market operates to pattern its is how a property-owning democracy think a of this of Rawls's overall strategy critique of Rawlsian special incentives to which I have this strategy to make Rawls a latter in his of to that a society that the to all would do so as the unintended effect of by that was and I think that is a if as an of the point is that if this restructuring of the market is the only feasible way to be an egalitarian then there is feasible way to a conception of as fairness. in a political economy structured in this way and their in their labour this commitment to and do not is in the phrase in the from Rawls where he notes that are free to pursue their permissible within a just also on to point of between Rawls and namely, their of can have that are not intended by any in the market. I the phrase “invisible hand” with given of of this phrase that is of But like believes that markets involve a of externalities, and some some In particular Rawls a between his first principle the basic liberties, equality of and the of a think this point is important as it has a on the between a property-owning democracy and a market Rawls his own in ideas about political economy in an economy of One of the in Rawls's between property-owning democracy and liberal is whether the latter is to a property-owning democracy as it democratic of Rawls's for a property-owning democracy on this point as a liberal Rawls was to idea that there was an function to in the and the of the democratic ideal into a place for good we most of our of argument seems to Rawls is right to the of if was an economy not to be made up of such given that they are such to The is to be in the that Rawls makes between the regime of liberty by his first principle and market It seems to that and are also that, if worker have for society as a whole, there is a rationale for giving them tax to their further and the of a they were equally to emphasise that given competitive and a regime of liberty, we can that some people will trade democratic of their for other that Rawls is a political liberal with for the of Rawls then he is to the idea that political is itself a part of the good do not have to lead in which political is one of their they have to their role as a of when it to of public this it does not then, if a works in a that does not democratic there is in the context of society as a the the political liberal people to work in that their for but it is not that a mixed economy of both worker and need have this Furthermore, there is unintended in a property-owning democracy by Paul namely, that employers are to have to more for given the supply of labour and the of in a property-owning democracy will not be on to the market by drudge jobs, and as part of making jobs will have to increase for making at work make if they do that Rawls makes between a regime of liberty and the market via explains as and we can a property-owning democracy to a of of both I not here to competitive as Rawls argument that a society as a could a and not concerned with of capital a regime of liberty we can people to trade other for the value of democratic of his so does in my any of the of political do not political all of the even democratic of where they If they to trade this value this is not to say that it is not a but that the regime of liberty by Rawls's first principle them to do need be overall of the of if it is in and more only is this of worker ownership not a point between a property-owning democracy and market it also seems to a point in of property-owning democracy. This is because the capital given to each in the latter people to more in where they to work with their on income from labour more likely to to work in a that them democratic would like to by one important to the idea of property-owning democracy and a as a for The is that all that can be by an property-owning democracy is giving people a in their own more we are to the of the It is to that the in to ownership to with in other a at The then, is this: everyone's exposure to the of the first point to make is that this view the role played in a property-owning democracy by human capital. It the way in which the of the principle of equality of opportunity requires the of of So a property-owning democracy is not just about ownership and share even if it there is the of whether the to a property-owning democracy just all to risk to the of of the in that very Rawls's is in the we are working here in ideal it will be free from the generated on the of regulation from the very concentrations of wealth that a property democracy to The role played by a of regulation in the is a The to this is that part of Meade's was a publicly unit that share ownership the of the market thus a of from its However, my main is that it is not who is to do any better in capital and that most of are in this notes the of of the public funds in the and Meade's a striking to how the capital of some of are and In I think on is the most realistic way to to make a property-owning democracy a would like to with a very much at the that would the first a property-owning democracy. The some in this with the introduction of the and the but both of were in My is much more and comparable to and for a payment of to My is but in one way more and in way more A in the is a way of people to by their with a from a of government, and private when the its are to a first a small for a capital that is I think a good place to is with the in most that do not to on the market The first is to make a state for of hold an capital for by unit by This capital is, up My is to the as part of a to the potential of a property-owning democracy for at the state would a capital on their state to and This capital like a would the form of only for limited for education enterprise the of a first property as a in a The is that and to the money at this up on a will that by an In my are up part of own underlying in the that will be adequately by the security and it brings the of to The can their capital in to an investment in The will have all the of a capital at a when it can make a difference to whether they a own will also them a of from In with the full implementation of Rawls's this will both make a property-owning democracy a and also address the of how a like this is to be
Scots: Studies in its Literature and Language John M. Kirk and Iseabail Macleod (eds). Rodopi, 2013 ISBN 9789042037397, 65 [euro], 309pp. Scholarly Festschrifts dedicated to a specific scholar usually mark the coronation of the achievements of a lifelong career, and express the esteem and consideration in which the honorand is held by their peers; in the present case, the unquestionable importance and the extremely high quality of J. Derrick McClure's committed involvement in most aspects of Scots research makes it only too appropriate that this collection should gather some of the best names in Scottish Studies to celebrate him with diversified articles focussing on his field of expertise. Jeremy Smith's 'Textual Afterlives: Barbour's Bruce and Hary's Wallace' skilfully presents the application of historical pragmatics to five editions of John Barbour's The Bruce and Blind Hary's The Wallace, namely John Ramsay's manuscript (1489), Robert Leprevik's print (1571), Andro Hart's edition of 1620, Robert Freebaim's of 1758 and John Pinkerton's of 1790 for the first, and John Ramsay's manuscript (1488), Robert Leprevik's edition of 1570, the Glasgow editions of 1685 and 1713, Robert Freebaim's of 1758, and Robert Morison's of 1790 for the second. Details such as layout, punctuation, capitalisation, fonts and the individual treatment of distinctively Scottish lexemes are carefully sifted in order to infer the effect which they presumably exerted on their contemporary Scottish readership; special attention is devoted to the medieval and early modern understanding of a text as a conglomerate of concepts rather than grammatical units, as well as to any indicator of the shift from an oral to a visual approach which the introduction of the printing press is known to have entailed. Variations in editorial choices during the centuries suggest a growing antiquarian interest in correctness, as well as the first signs of a romantic 'mythological historicity' which drew heavily on the epic's purported authenticity inspired by the authority of the manuscript originals; prefaces and textual interpolations, on the other hand, are noted as reflecting the evolution of society's political attitude towards its southern neighbour. It is this constant redefinition of Scottish society's identity which, according to the article's premise, is reflected in the textual minutiae, and which warrants the exploration of each edition as a culturally-embedded product of its period, rather than a mere reproduction of the original. Robert McColl Millar's To bring my language near to the language of men? Dialect and Dialect Use in the Eighteenth and Early Nineteenth Centuries: Some Observations' explores how the social, economic and political changes which marked the second half of the eighteenth century influenced the increasingly self-conscious recourse to dialect for literary purposes against the opposing tendency of widespread diffidence towards anything diverging from the accepted norm. The case study concentrates on two emigrants' letters to their families, one from a Scottish indentured servant in Maryland in the early eighteenth century, the second from an English political prisoner in New South Wales in the early 1800s: the relevant dates are posited as the two approximate temporal extremes between which the standard language is presumed to have imposed itself. The theoretical premise is then tested against the entries found in the early nineteenth-century Original Statistical Account of Scotland, which are subdivided into the two categories of overt attitudes towards language as opposed to covered ones. Both concepts have been previously developed by McColl Millar, and in this article identify the self-conscious, often ambivalent comments passed by the informants on the linguistic landscape of their district on the one hand, and the incursions of Scots lexical items, idioms and proverbs into an otherwise wholly English text on the other. …
The paper studies Hausa film language through the analysis of three communication strategies, namely proverbs, imperatives and forms of address. It shows that Hausa film creates a new discourse by reflecting modern and traditional Hausa society. The films preserve some accepted cultural norms of behavior and norms of communication in order to please the more conservative public. On the other hand, combination of traditional and modern Hausa lifestyle evokes changes in the discourse. The paper shows that proverbs are commonly used as communication strategy for indirectness, rather than a specialized language. It also discovers that imperatives are used as communication strategy in close relations between interlocutors (no matter what their social status is) to express the direct message. As for forms of address, traditional and borrowed terms reflect the changing style of life. The examples extracted from the Hausa films are to show how the regular grammatical and lexical means change their discourse function in new social context.
In this paper we present and justify methodological principles and syntactic criteria to design an annotation scheme for a Persian Treebank. The advantages of the proposed scheme for annotation of the Persian Treebank will be discussed. At the same time, we present the way that different types of linguistic knowledge (morphological, syntactic and semantic) are encoded in the structures of the schema. We will show how this scheme can account for many of the syntactic constructions that appear to be unique to the Persian language.
Lexical databases following the wordnet paradigm capture information about words, word senses, and their relationships. A large number of existing tools and datasets are based on the original WordNet, so extending the landscape of resources aligned with WordNet leads to great potential for interoperability and to substantial synergies. Wordnets are being compiled for a considerable number of languages, however most have yet to reach a comparable level of coverage. We propose a method for automatically producing such resources for new languages based on WordNet, and analyse the implications of this approach both from a linguistic perspective as well as by considering natural language processing tasks. Our approach takes advantage of the original WordNet in conjunction with translation dictionaries. A small set of training associations is used to learn a statistical model for predicting associations between terms and senses. The associations are represented using a variety of scores that take into account structural properties as well as semantic relatedness and corpus frequency information. Although the resulting wordnets are imperfect in terms of their quality and coverage of language-specific phenomena, we show that they constitute a cheap and suitable alternative for many applications, both for monolingual tasks as well as for cross-lingual interoperability. Apart from analysing the resources directly, we conducted tests on semantic relatedness assessment and cross-lingual text classification with very promising results.
This paper will focus on recent and near-term future developments at FrameNet (FN) and the interoperability issues they raise. We begin by discussing the current state of the Berkeley FN database including major changes in the data format for the latest data release. We then briefly review two recent local projects, "Rapid Vanguarding”, which has created a new interface for the frame and lexical unit definition process based on the Word Sketch Engine of Kilgarriff et al. (2004), and “Beyond the Core”, which has developed tools for annotating constructions, and created a sample “construction” of especially “interesting” constructions which are neither simply lexical nor easy for the standard parsers to parse. We also cover two current collaborations, FN’s part in the development of the manually annotated subcorpus of the American National Corpus, and a pilot study on aligning WordNet and FrameNet, to exploit the complementary strengths of these quite different resources. We discuss FN-related research on Spanish, Japanese, German (SALSA), Chinese and other languages, and the language-independence of frames, along with interesting FN-related work by others, and a sketch of a large group of image-schematic frames which are now being added to FN. We close with some ideas about how FrameNet can be opened up, to allow broader participation in the development process without losing precision and coherence, including a small-scale study on acquiring data for FN using Amazon’s Mechanical Turk crowd-sourcing system.
This paper investigates how to best couple hand-annotated data with information extracted from an external lexical resource to improve part-of-speech tagging performance. Focusing mostly on French tagging, we introduce a maximum entropy Markov model-based tagging system that is enriched with information extracted from a morphological resource. This system gives a 97.75 % accuracy on the French Treebank, an error reduction of 25 % (38 % on unknown words) over the same tagger without lexical information. We perform a series of experiments that help understanding how this lexical information helps improving tagging accuracy. We also conduct experiments on datasets and lexicons of varying sizes in order to assess the best trade-off between annotating data versus developing a lexicon. We find that the use of a lexicon improves the quality of the tagger at any stage of development of either resource, and that for fixed performance levels the availability of the full lexicon consistently reduces the need for supervised data by at least one half.
This comprehensive book makes many original contributions to the field of genres on the web. The identification and characterization of genres is of obvious interest to “pure” linguistics, but as this book makes clear, there are some important practical applications. Chief amongst these will be the advent of genre-aware search engines, where users will be able to specify not only their topics of interest, but the desired genre of the returned web pages, as in the WEGA search engine described in this book by Stein et al. Crowston et al. give the example of someone wishing to buy a digital camera. A traditional search engine would return pages on the topic of the specified brand of digital cameras, most of which will just be the web sites of sellers. But what the buyer really wants is information about this type of camera in certain genres only, such as product reviews and opinion-bearing blogs, which provide the opinions of people who have already bought that camera. The...
The relationship between ontologies and natural language lexicons is a hotly debated one. An ontology is a formalized system of concepts (potentially of a specific domain) and the relations these concepts entertain. A lexicon, on the other hand, is the language component that contains the conventionalized knowledge of natural language speakers about lexical items (mostly words, but also morphemes and idioms). Ontologies ‘operate’ on the conceptual level, lexicons on the linguistic level. Ontologies systematize and relate concepts, lexicons systematize and relate words and other lexical items. However, as semantic relations between lexical items reflect meaning relatedness and meaning is essentially conceptual, both notions appear to be very close to one another (and are often wrongly used interchangeably). The interplay of and mapping between ontologies and lexical resources is therefore a vital and challenging field of research, one which has gained additional momentum and importance...
Many research questions require a within-class object recognition task matched for general cognitive requirements with a face recognition task. If the object task also has high internal reliability, it can improve accuracy and power in group analyses (e.g., mean inversion effects for faces vs. objects), individual-difference studies (e.g., correlations between certain perceptual abilities and face/object recognition), and case studies in neuropsychology (e.g., whether a prosopagnosic shows a face-specific or object-general deficit). Here, we present such a task. Our Cambridge Car Memory Test (CCMT) was matched in format to the established Cambridge Face Memory Test, requiring recognition of exemplars across view and lighting change. We tested 153 young adults (93 female). Results showed high reliability (Cronbach's alpha = .84) and a range of scores suitable both for normal-range individual-difference studies and, potentially, for diagnosis of impairment. The mean for males was much higher than the mean for females. We demonstrate independence between face memory and car memory (dissociation based on sex, plus a modest correlation between the two), including where participants have high relative expertise with cars. We also show that expertise with real car makes and models of the era used in the test significantly predicts CCMT performance. Surprisingly, however, regression analyses imply that there is an effect of sex per se on the CCMT that is not attributable to a stereotypical male advantage in car expertise.
The Assessment Battery for Communication (ABaCo) was introduced to evaluate pragmatic abilities in patients with cerebral lesions. The battery is organized into five evaluation scales focusing on separate components of pragmatic competence. In the present study, we present normative data for individuals 15–75 years of age (N = 300). The sample was stratified by age, sex, and years of education, according to Italian National Institute of Statistics indications in order to be representative of the general national population. Since performance on the ABaCo decreases with age and lower years of education, the norms were stratified for both age and education. The ABaCo is a valuable tool in clinical practice; the normative data provided here will enable clinicians to determine different kinds and specific levels of communicative impairments more precisely.
Because of wide disparities in college students’ math knowledge—that is, their math achievement—studies of cognitive processing in math tasks also need to assess their individual level of math achievement. For many research settings, however, using existing math achievement tests is either too costly or too time consuming. To solve this dilemma, we present three brief tests of math achievement here, two drawn from the Wide Range Achievement Test and one composed of noncopyrighted items. All three correlated substantially with the full achievement test and with math anxiety, our original focus, and all show acceptable to excellent reliability. When lengthy testing is not feasible, one of these brief tests can be substituted.
The measurement of executive function has a long history in clinical and experimental neuropsychology. The goal of the present report was to determine the profile of behavior across the lifespan on four computerized measures of executive function contained in the recently developed Psychology Experiment Building Language (PEBL) test battery http://pebl.sourceforge.net/ and evaluate whether this pattern is comparable to data previously obtained with the non-PEBL versions of these tests. Participants (N = 1,223; ages, 5–89 years) completed the PEBL Trail Making Test (pTMT), the Wisconsin Card Sort Test (pWCST; Berg, Journal of General Psychology, 39, 15–22, 1948; Grant & Berg, Journal of Experimental Psychology, 38, 404–411, 1948), the Tower of London (pToL), or a time estimation task (Time-Wall). Age-related effects were found over all four tests, especially as age increased from young childhood through adulthood. For several tests and measures (including pToL and pTMT), age-related slowing was found as age increased in adulthood. Together, these findings indicate that the PEBL tests provide valid and versatile new research tools for measuring executive functions.
Psychologists, psycholinguists, and other researchers using language stimuli have been struggling for more than 30 years with the problem of how to analyze experimental data that contain two crossed random effects (items and participants). The classical analysis of variance does not apply; alternatives have been proposed but have failed to catch on, and a statistically unsatisfactory procedure of using two approximations (known as F 1 and F 2) has become the standard. A simple and elegant solution using mixed model analysis has been available for 15 years, and recent improvements in statistical software have made mixed models analysis widely available. The aim of this article is to increase the use of mixed models by giving a concise practical introduction and by giving clear directions for undertaking the analysis in the most popular statistical packages. The article also introduces the djmixed add-on package for SPSS, which makes entering the models and reporting their results as straightforward as possible.