Papers reviewed and determined not to be word norm studies. Use the flag icon to report errors or suggest re-inclusion.
16504 papers
To assess the effects of discrepancy between two independent variables, investigators sometimes compute difference scores and correlate such scores with a criterion variable. However, the correlation of the difference with the criterion is accounted for by the correlations of the difference constituents with the criterion and the constituents’ variances. It follows that when investigators are testing a prediction that is not captured by the difference constituents’ main effects, using the difference correlation analysis may be misleading. Under these circumstances, the effects of a discrepancy between two independent variables can be assessed by a test of their interaction. The problems inherent in using difference scores and the advantage of testing the interaction are illustrated in relation to research programs on two separate topics in social psychology.
Parallel to, and to some degree inreaction to French poststructuralisttheorization (as championed by Derrida,Foucault, and Lacan, among others) is a Frenchneo-structuralism built directly on theachievements of structuralism using electronicmeans. This paper examines some exemplaryapproaches to text analysis in thisneo-structuralist vein: SATOR's topoidictionary, the WinBrill POS tagger andFrançois Rastier's interpretativesemantics. I consider how a computer-assisted``Wissenschaft'' accumulation of expertisecomplements the neo-structuralist approach.Ultimately, electronic critical studies will bedefined by their strategic position at theintersection of the two chief technologiesshaping our society: the new informationprocessing technology of computers and therepresentational techniques that haveaccumulated for centuries in texts.Understanding how these two informationmanagement paradigms complement each other is akey issue for the humanities, for computerscience, and vital to industry, even beyond thenarrow realm of the language industries. Thedirection of critical studies, a small planetlong orbiting in only rarefied academiccircles, will be radically altered by the sheersize of the economic stakes implied by a newkind of text, the industrial text, thetechnological heart of an information society.
It is important to give useful clues for selecting desiredcontent from a number of retrieval results obtained (usually) from avague search request. Compared with monolingual retrieval, such asupport framework is inevitable and much more significant for filteringgiven translingual retrieval results. This paper describes an attempt toprovide appropriate translation of major keywords in each document in across-language information retrieval (CLIR) result, as a browsingsupport for users. Our idea of determining appropriate translation ofmajor keywords is based on word co-occurrence distribution in thetranslation target language, considering the actual situation of WWWcontent where it is difficult to obtain aligned parallel (multilingual)corpora. The proposed method provides higher quality of keywordtranslation to yield a more effective support in identifying the targetdocuments in the retrieval result. We report the advantage of thisbrowsing support technique through evaluation experiments includingcomparison with conditions of referring to a translated documentsummary, and discuss related issues to be examined towards moreeffective cross-language information extraction.
Results of the noun–verb pair comprehension and production tests from the Test Battery for Auslan Morphology and Syntax (A. Schembri et al., 2000) are presented, reanalyzed, and compared to data from 2 other cases dealing with noun–verb pairs: the Auslan lexical database and a comparison of Auslan and American Sign Language (ASL) signs. The data confirm the existence of formationally related noun–verb pairs in Auslan in which the verb displays a single movement and the noun displays a repeated movement. The data also suggest that the best exemplars of noun–verb pairs of this type in Auslan form a distinct set of iconic (mimetic) signs archetypically based on inherently reversible actions (such as opening and shutting). This strong iconic link perhaps explains why the derivational process appears to be of limited productivity, though it does appear to have 'spread' to a number of signs that appear to have no such iconicity. There appears to be considerable variability in the use of the derivational markings, particularly in connected discourse, even for signs of the 'open and shut' variety. Overall, the derivational process is apparently still closely linked to an iconic base, is incipient in the grammar of Auslan, and is best described as only partially grammaticalized. (PsycINFO Database Record (c) 2016 APA, all rights reserved)
This paper describes how traditional andnon-traditional methods were used to identifyseventeen previously unknown articles that webelieve to be by Stephen Crane, published inthe New-York Tribune between 1889 and1892. The articles, printed without byline inwhat was at the time New York City's mostprestigious newspaper, report on activities ina string of summer resort towns on New Jersey'snorthern shore. Scholars had previouslyidentified fourteen shore reports as Crane's;these possible attributions more than doublethat corpus. The seventeen articles confirmhow remarkably early Stephen Crane set hisdistinctive writing style and artistic agenda. In addition, the sheer quantity of the articlesfrom the summer of 1892 reveals how vigorouslythe twenty-year-old Crane sought to establishhimself in the role of professional writer. Finally, our discovery of an article about theNew Jersey National Guard's summer encampmentreveals another way in which Crane immersedhimself in nineteenth-century military cultureand help to explain how a young man who hadnever seen a battle could write so convincinglyof war in his soon-to-come masterpiece,The Red Badge of Courage. We argue that thejoint interdisciplinary approach employed inthis paper should be the way in whichattributional research is conducted.
Imagine discourse between the arts in which the conventions of what we might call ordinary cognition do not apply, on site of intense lobbying neither tethered by history or cultural integrity, nor, frequently, concerned with social cohesion or communicative norms. It will be discourse in which the categories of an imperial culture are abrogated (however temporarily) by an indigenous one, yet it will undoubtedly also be site of intense colonization. On it, likewise, there will be an appropriation of language on an unprecedented scale. Past experience will play little part. Memory short and episodic, rather than semantic. It primal discourse. Primal in that it the site of first contact. Primal also in that it most often considered be the meeting of primitive culture and an advanced. Primal, likewise, in behavioral sense: in it, the satisfaction of physiological needs tantamount. Indeed, body and mind here are in state of kinetic unrest. This scene of prolonged immat urity, yet ontological and epistemological questions held in private language are encouraged be made public. Here the verbal arts have no canon. Literature has no prevailing cultural standard of merit. Questions of the popular and the high cultural are not naturalized and the fictional and nonfictional carry the same degree of verisimilitude as works of propaganda, rhetoric, and didacticism. In modem times, the West has become the site of this tenacious yet frequently unacknowledged imperial discourse, the discourse between multifarious forms of artistic representation win the attention of children. It in such discourse that the picture book located. 1 Arguing the need for critical language for the discussion of children's picture books, Peter Hunt suggests that to pictures into the same mould as words seems be potentially unproductive, except in terms of establishing conventions, when, of course, it is, by definition, necessary (181). It impossible, however, conventionalize pictorial representation the same degree as linguistic representation. Linguistic systems are mastered painstakingly, piece by piece, referent by referent, word by word. Pictorial systems, by contrast, are mastered all at once; they involve what Flint Schier has called natural generativity and are therefore much less conventional than linguistic systems. Each system, nevertheless, relies on general agreement and on willingness engage in communicative activity: the pictorial system on deep recognitional capacities that link object and its picture, the linguistic system on lexical and syntactical regularities and rules. In the media-saturated culture of the contemporary West, the commonalities and differences of our separate but shared experiences are frequently offered up in televisual or hypertextual format in which what Hunt describes as force set of discursive practices that address and interpellate both adults and children as potential viewers or listeners. The linguistic and the pictorial are frequently experienced as synergistic or polylogic systems bound up in this mass media, media whose intention, according Jean Baudrillard, transcribe the complexity of contemporary life into an ongoing procession of meaningless simulacra, hyperreal, a real without origin or reality (2). Baudrillard's disenchanted vision of postmodernity, articulated most profoundly in the late 1970s and 1980s, produced an interesting ontological metaphor. Disneyland, he claimed, is there conceal the fact that it the 'real' country, all of 'real' America, which Disneyland (just as prisons are there conceal the fact that it the social, in its entirety, in its banal omnipresence, which carceral). Disneyland presented as imaginary in order make us believe that the rest real, when in fact all of Los Angeles and the America surrounding it are no longer real, but of the order of the hyperreal and of simulation (25). …
We present a rule--based shallow--parser compiler, which allows to generate a robust shallow-parser for any language, even in the absence of training data, by resorting to a very limited number of rules which aim at identifying constituent boundaries. We contrast our approach to other approaches used for shallow--parsing (i.e. finite-state and probabilistic methods). We present an evaluation of our tool for English (Penn Treebank) and for French (newspaper corpus "LeMonde") for several tasks (NP-chunking & "deeper" parsing).
We consider several perceptual issues in the context of machine recognition ofmusic patterns. It is argued that a successful implementation of a musicrecognition system must incorporate perceptual information and error criteria.We discuss several measures of rhythm complexity which are used fordetermining relative weights of pitch and rhythm errors. Then, a new methodfor determining a localized tonal context is proposed. This method is based onempirically derived key distances. The generated key assignments are then usedto construct the perceptual pitch error criterion which is based on noterelatedness ratings obtained from experiments with human listeners.
The stability over time of serum IgG antibody levels to human papillomavirus type 16 (HPV-16) was determined by comparing the HPV-16 capsid antibody levels in serial serum samples of an age-stratified random subsample of 1656 primiparous mothers resident in Helsinki who were followed until their second pregnancy, on average 29.5 months later. The correlation between the first and second pregnancy HPV-16 serum antibody levels of the same woman was high, even when >4 years had elapsed between pregnancies (r =.822). Between negativity, indeterminate results, or quartiles of positivity, the predictive values for being classified in the same category on both occasions ranged between 42% and 91%. Correlation coefficients, predictive values, and kappa coefficients between serial samples all were comparable with those of repeat analyses of the same sample, indicating that HPV capsid antibody levels are generally stable during several years of follow-up.
Leech (1969), considering English poetry, treats poems as linguistically deviant forms of language in which the form and content are “foregrounded” against a background of nondeviant language. The deviant language that is used may be noticeably irregular or noticeably regular. The content of the poem may also be deviant as the poet creates meanings that are not expected to be taken literally. I have applied these ideas to the British Sign Language (BSL) poem “Trio” by the late Deaf poet Dorothy Miles. Analysis of this poem shows the same features in BSL poetry that Leech found in English poetry. The poetry of Dorothy (Dot) Miles is widely considered in Britain to be some of the best BSL poetry in the public domain. Her work is powerful, thoroughly crafted, and richly significant at many levels, and it easily justifies—and rewards—careful linguistic analysis. One well-known feature of Dot’s poetry is that it frequently “works” both in English and in BSL. This feature, however, will not be the focus of this chapter. Instead, I will consider the features of BSL that create the richness of the three-part poem, “Trio,” which is made up of three stanzas: Morning, Afternoon, and Evening.1 The starting point for this analysis is the fact that poetry is a deviation from ordinary language. Poetry is not only allowed to deviate from normal patterns of ordinary language, but also it is expected to do so. Careful choice of linguistic forms allows the poet to produce language that carries significance far greater than that of ordinary language. Leech (1969) has defined foregrounding as the deviation from linguistic norms for the sake of art. The foregrounding in “Trio” occurs in two different types of deviation that involve noticeably irregular and noticeably regular use of language. In the first type, poetic deviation creates a
Low interrater reliability coefficients are a common problem for behavior rating scales. One hypothesis to account for this is that raters have different frames of reference from which to judge behaviors. In the present study, the interrater reliability of the Devereux Behavior Rating Scale-School Form was examined, and the hypothesis that teacher frame of reference influences ratings was explored. Special and general education teachers rated the behavior of 51 children with emotional disturbance (ED), and general education teachers inde pendently rated the behavior of 51 matched control children. Interrater reliability coeffi cients were higher for the general education sample than for the sample of children with ED. Limited support was found for the hypoth esis that frame of reference may affect ratings. Findings suggest that many factors influence ratings and that teachers may benefit from rater training.
In statistical parsing, the probabilistic models are used to evaluate the possibility of each candidate parse tree, where the parse tree with the largest probability is deemed to be the final result of the parsing. Therefore, the core of statistical parsing is a probabilistic evaluation model. The main difference among the various probabilistic evaluation models lies in which types of features in the context are used to assign the probabilities to the parse trees. Various probabilistic evaluation models have been proposed in the field of statistical parsing, where different models use different feature types. How to evaluate a feature type's predictive power for the parsing tree? The paper proposes an information theory based feature type analysis model. Using the method, we can quantitatively analyze the power of different feature types for syntactic structure prediction from the viewpoint of information theory. The basic idea is that we use entropy and conditional entropy to measure whether a feature type grasps some of the information for syntactic structure prediction. If the average uncertainty of the syntactic structures declines apparently, the feature type is deemed to have grasped some intrinsic linguistic information in the context that has close relation to the syntactic structure. Using Penn Treebank as training and testing set, our experiment quantitatively analyze the different feature types' predictive power for syntactic structure predictive power for syntactic structure prediction in a systematic way and draws a series of conclusions which reflect the predictive power of different feature types and feature type combination for syntactic parsing.
Abstract This study examines judgments of male self-control regarding sexual aggression in dating situations. A survey was conducted using vignettes that described situations where a hypothetical man attempts to have intercourse with his female date, she resists, and he considers various ways of overcoming her resistance. Survey respondents judged these vignettes. The goals of the study were to estimate the prevalence of beliefs about low male self-control, and to examine how contextual information affects these assessments. Contrary to expectations derived from the literature, the large majority of respondents attributed high levels of self-control to men across a wide range of scenarios. As expected, the man's alcohol intoxication reduced perceived control. Against expectations, consensual foreplay and prior sex did not affect ratings. Implications for social learning and rational choice theories of sexual aggression are discussed, with particular emphasis on Ajzen's (1988) theory of planned behavior.
Computational learning of natural language is often attempted without using the knowledge available from other research areas such as psychology and linguistics. This can lead to systems that solve problems that are neither theoretically or practically useful. In this paper we present a system CLL which aims to learn natural language syntax in a way that is both computationally effective and psychologically plausible. This theoretically plausible system can also perform the practically useful task of unsupervised learning of syntax. CLL has then been applied to a corpus of declarative sentences from the Penn Treebank (Marcus et al., 1993; Marcus et al., 1994) on which it has been shown to perform comparatively well with respect to much less psychologically plausible systems, which are significantly more supervised and are applied to somewhat simpler problems.
This paper compares a number of generative probability models for a wide-coverage Combinatory Categorial Grammar (CCG) parser. These models are trained and tested on a corpus obtained by translating the Penn Treebank trees into CCG normal-form derivations. According to an evaluation of unlabeled word-word dependencies, our best model achieves a performance of 89.9%, comparable to the figures given by Collins (1999) for a linguistically less expressive grammar. In contrast to Gildea (2001), we find a significant improvement from modeling word-word dependencies.
We present a generative distributional model for the unsupervised induction of natural language syntax which explicitly models constituent yields and contexts. Parameter search with EM produces higher quality analyses than previously exhibited by unsupervised systems, giving the best published un-supervised parsing results on the ATIS corpus. Experiments on Penn treebank sentences of comparable length show an even higher F1 of 71% on non-trivial brackets. We compare distributionally induced and actual part-of-speech tags as input data, and examine extensions to the basic model. We discuss errors made by the system, compare the system to previous models, and discuss upper bounds, lower bounds, and stability for this task.
Manual, large scale (computational) grammar development is time consuming, expensive and requires lots of linguistic expertise. More recently, a number of alternatives based on treebank resources (such as Penn-II, Susanne, AP treebank) have been explored. The idea is to automatically ``induce'' or rather read off (P)CFG grammars from the parse annotated treebank resources and to use the treebank grammars thus obtained in (probabilistic) parsing or as a starting point for further grammar development. The approach is cheap, fast, automatic, large scale, ``data driven'' and based on real language resources.\n\nTreebank grammars typically involve large sets of lexical tags and non-lexical categories as syntactic information tends to be encoded in monadic category symbols. They feature flat rules (trees) that can ``underspecify'' attachment possibilities. Treebank grammars do not in general follow Xbar architectural design principles (this is not to say that treebank grammars do not have design principles). As a consequence, treebank grammars tend to have very large CFG rule bases (e.g. Penn-II > 17,000 CFG rules for about 1 million words of text) with often only minimally differing rules. Even though treebank grammars are large, they are still incomplete, exhibiting unabated rule accession rates. From a grammar engineering point of view, the size of the rule base poses problems for maintainability, extendability and, if a treebank grammar is to be used as a CF-base in a LFG grammar, for functional (feature-structure) annotations. From the point of view of theoretical linguistics, flat treebank trees and treebank grammars extracted from such trees do not express linguistic generalisations. From the perspective of empirical and corpus linguistics, flat trees are well-motivated as they allow underspecification of subtle and often time consuming attachment decisions. Indeed, it is sometimes doubted whether highly general Xbar schemata usefully scale to ``real'' language.\n\nIn previous work we developed methodologies for automatic feature-structure annotation of grammars extracted from treebanks. Automatic annotation of ``raw'' treebank grammars is difficult as annotation rules often need to identify subsequences in the RHSs of flat treebank rules as they explicitly encode head, complement and modifier relations. Xbar based CFG rules should substantially facilitate automatic feature-structure annotation of grammar rules.\n\nIn the present paper we conduct a number of experiments to explore a space of possible grammars based on a small fragment of the AP treebank resource. Starting with the original treebank fragment we automatically extract a CFG G. We then apply an automatic structure preserving grammar compaction step which generalises categories in the original treebank fragment and reduces the number of rules extracted, resulting in a generalised treebank fragment and in a compacted grammar Gc. The generalised fragment is then manually corrected to catch missed constituents (and the like) resulting in an automatically extracted, compacted and (effectively manually) corrected grammar Gc,m. Manual correction proceeds in the ``spirit'' of treebank grammars (we do not introduce Xbar analyses). We then explore how many of the manual correction steps on treebank trees can be achieved automatically. We develop, implement and test an automatic treebank ``grooming'' methodology which is applied to the generalised treebank fragment to yield a compacted and automatically corrected grammar Gc,a. Grammars Gc,m and Gc,a are very similar to compiled out ``flat'' LFG-82 style grammars. We explore regular expression based compaction (both manual and automatic) to relate Gc,m to a LFG-82 style grammar design. Finally, we manually recode a subsection of the generalised and manually corrected treebank fragment into ``vanilla-flavour'' XBar based trees. From these we extract a compacted, manually corrected, XBar based grammar Gc,m,x. We evaluate our grammars and methods using standard labelled bracketing measures and according to how well they perform under automatic feature-structure annotation tasks.
This study investigated the effect of a weak magnetic field (50 microT, 20 Hz sinusoidal, 5 s duration) on concurrent perceptions of visual stimuli. Subjects were seated between Helmholtz coils and gave post-exposure ratings for the affective content and arousing nature of presented images. They were blind as to the presence or absence of a simultaneously presented field. Skin conductance and arousal ratings did not show significant differences between experimental and control conditions, but the affective content rating did (P = 0.041), with the images viewed under field exposure being rated as having a more positive affect. Such measures might thus be useful as additional indicators of magnetic field detection. A post-hoc analysis of skin conductance profiles showed that 48% of subjects exhibited a lowering of skin conductance during field exposure, 34% exhibited no apparent reaction, and 17% exhibited an increase. Overall ratings given by each of the groups appeared to relate to these physiological profiles.
This paper presents empirical studies and closely corresponding theoretical models of the performance of a chart parser exhaustively parsing the Penn Treebank with the Treebank's own CFG grammar. We show how performance is dramatically affected by rule representation and tree transformations, but little by top-down vs. bottom-up strategies. We discuss grammatical saturation, including analysis of the strongly connected components of the phrasal nonterminals in the Treebank, and model how, as sentence length increases, the effective grammar rule size increases as regions of the grammar are unlocked, yielding super-cubic observed time behavior in some configurations.
Natural language processing technologies offer ease-of-use of computers for average users, and ease-of-access to on-line information. Natural language, however, is complex, and the traditional methods of parsing with a single grammar and parser may result in an inefficient and large system that is difficult to maintain and is fragile in dealing with language irregularities. This paper begins by reviewing an alternative effort in grammar decomposition (also known as grammar partitioning) for natural language parsing, which aims to alleviate these problems. We then propose a novel automatic approach for grammar partitioning, in comparison with a random method of partitioning. Our experiments show that syntactic GLR parsing is formidable for the Wall Street Journal corpus in the Penn Treebank when a single grammar is used. This is due to too many grammar rules for parsing table generation. However, grammar partitioning solves the problem and offers a viable alternative. Our results also show that our automatic grammar partitioning method based on the mutual information criterion fares better than a random partitioning method and exhibits efficiency in parsing as well as high parse coverage. 1
International audience
Se presentan los valores normativos de la adapta The norms of the Spanish adaptation for shows 9 ci6n espanola de los conjuntos 9-14 (segunda parte) del 14 (second part) of the International Affective Picture Intemational Affective Picture System (lAPS). Los resul System (lAPS) are presented. The results are highly tados muestran una alta consistencia con los obtenidos consistent with those obtained in the first part of the en la primera parte de la adaptaci6n espanola y con los Spanish adaptation and in the original USA version. The valores originales norteamericanos. La distrlbuclon de picture distribution in the bi-dimensional space, defined las diapositivas en el espacio bidimensional valencia by the ratings of valence and arousal, displays the typical arousal adopta la tipica forma de boomerang, obssrvan boomerang form. In addition, a more pronounced slope dose una menor inclinaci6n, junto con una mayor disper -and a smaller dispersion- is observed in the si6n, en el brazo que se extiende hacia el polo agrada unpleasantextreme of the boomerang than in the pleasant ble que en el brazo que se extiende hacia el polo des one. The correlations between the Noth-American and agradable. Las correlaciones entre las evaluaciones the Spanish values are all highly significant. Nevertheless, norteamericanas y espanolas son todas altamente sig the differences found in the first part of the study are nificativas. No obstante, se confirman las diferencias confirmed: Spanish people perceive the affective pictures, encontradas en la primera parte del trabajo, en el sen as a whole, as more arousing and less dominant than tido de que los aspafioles perciben las imagenes North-Americans. Similarly, our results confirm the gender afectivas, en su conjunto, con mayor nivel de activaci6n differences previously found: women rate the pictures y menor nivel de control que los norteamericanos. Asi as more arousing and less dominant than men. Although mismo, sa confirman las diferencias de genero encon no significant differences are observed in valence ratings, tradas anteriormente: las mujeres otorgan a las image as a whole, the pictures evaluated as more pleasant are nes un mayor nivel de acnvaclon y un nivel menor de clearly different for men and women. The implications of control que los varones. Aunque no exlsten diferenclas the results regarding the theoretical model underlying the significativas en las estimaciones globales de la valencia lAPS are highlighted. afectiva, las imagenes evaluadas como mas agradables por varones y mujeres son claramente diferentes. Las Key words: emotion, affective valence, arousal, implicaciones de los resultados con respecto al modelo dominance, cross-cultural differences, gender te6rico que subyace al lAPS son resaltadas. differences.
Hemispheric Effects of Concreteness in Pictures and Words Daniel J. Casasanto 1 (dcasasan@mail.med.upenn.edu) John Kounios 2 (jkounios@cattell.psych.upenn.edu) John A. Detre 1 (detre@mail.med.upenn.edu) Department of Neurology 1 and Psychology 2, University of Pennsylvania 3400 Spruce Street, Philadelphia, PA 19104 USA Introduction Functional Magnetic Resonance Imaging (fMRI) studies have demonstrated differently lateralized activation during episodic memory encoding for verbal and pictorial stimuli, as predicted by the material-specific model (1). Results are interpreted as consistent with Dual Coding Theory, which posits parallel verbal and imaginal systems by which a given stimulus may be represented (2), the neural substrates of which have also been shown to be differently lateralized (3). Whereas previous encoding studies have examined variation across material types, the present study investigated hemispheric effects within material types. FMRI was used to determine the laterality of encoding-related activation for verbal stimuli that varied in concreteness, and for pictorial stimuli that varied in verbalizability. Methods Task Design Verbal task. Nine healthy, right-handed native English speakers viewed blocks of serially presented English nouns, alternating with blocks of control strings (2500ms presentation, 500ms ITI). Noun stimuli comprised two sublists: 40 concrete and 40 abstract nouns (mean concreteness ratings 6.31 and 2.61, respectively, on a 1-to-7 point scale (Toglia & Battig, 1978)). Each task block comprised ten nouns: eight from one sublist and two from the other. Subjects classified nouns as concrete or abstract, responding via a left-or-right button press. Control blocks comprised ten strings of Ls or Rs, eliciting the same proportion of left and right button presses as the preceding block of nouns. Subjects were instructed to remember noun stimuli for a post-scan forced-choice recognition memory test. Pictorial task. For each of two face memory encoding tasks, eleven healthy, right-handed volunteers viewed blocks of unfamiliar face photographs, alternating with blocks of a repeatedly presented pixelated control image (six 40s task/control blocks, 10 stimuli per block, 3500ms presentation, 500ms ITI). For the first task, full-head photographs were shown, including hair, neck, and upper shoulders. In some cases, clothing and jewelry were visible. For the second task, the same set of face photographs was used, but each photograph was cropped so as to include the brow, eyes, nose, and mouth, but exclude ears, hair, and any extraneous objects. Subjects were instructed to remember the faces for a post-scan recognition test, and to attend the control images but not to memorize them. Scanning occurred during the encoding tasks but not during recognition testing. Image Acquisition and Data Analysis BOLD functional images were obtained at 1.5Tesla in 20 contiguous 5-mm-thick axial slices. Multisubject SPMs were constructed in Talairach space using the SPM99{t} random effects model. Cognitive subtraction revealed activation associated with encoding during blocks of predominantly concrete and predominantly abstract nouns, and during cropped-face and full-head encoding. For each task, activation exceeding a statistical threshold (∝=.05) was quantified in two a priori-defined regions of interest. ROIs comprised the inferior frontal gyrus (IFG) and fusiform gyri (FG), as these structures have demonstrated reliable material- specific effects in previous encoding studies. Hemispheric asymmetry of activation in each ROI was assessed using an asymmetry ratio [AR = (VoxelsR– VoxelsL)/(VoxelsR+VoxelsL)]. Results and Discussion Differing hemispheric effects shown previously across verbal and pictorial material types were demonstrated in the present study within material types. Average activation in the ROIs was bilateral during encoding of concrete nouns [AR(IFG)=-.06, ns; AR(FG)=.13, ns], which are amenable to both verbal an imaginal coding, but significantly left-lateralized during encoding of abstract nouns [AR(IFG)=-.22, P=.001; AR(FG)=-.80, P=.009], which are resistant to imaginal coding. Activation during encoding of full-head photographs was left-lateralized in the IFG [AR(IFG)=-.46, P=.001] and bilateral in the FG [AR(FG)=-.25, ns], but activation during encoding of cropped-faces, which are resistant to verbal coding, was significantly right- lateralized in both ROIs [AR(IFG)=.12, P=.001; AR(FG)=.71, P=.001]. Replication of these findings within verbal stimuli that vary in imageability and within pictorial stimuli that vary in verbalizability would suggest that hemispheric specialization during memory encoding heretofore described as material- specific might be more accurately described as code- specific. References Casasanto, D.J., et al. (2000) in Proceedings of the Cognitive Science Society 22, 77-82. Paivio, A. (1991) Canadian Journal of Psychology 45, Kounios, J. & Holcomb, P.J. (1994) J. Exp. Psych.: Learning, Memory, & Cognition 20, 804-823.
Treebanks are of two types according to their annotation schemata: phrase-structure Treebanks such as the English Penn Treebank [8] and dependency Treebanks such as the Czech dependency Treebank [6]. Long before Treebanks were developed and widely used for natural language processing, there had been much discussion of comparison between dependency grammars and context-free phrase-structure grammars [5]. In this paper, we address the relationship between dependency structures and phrase structures from a practical perspective; namely, the exploration of different algorithms that convert dependency structures to phrase structures and the evaluation of their performance against an existing Treebank. This work not only provides ways to convert Treebanks from one type of representation to the other, but also clarifies the differences in representational coverage of the two approaches.
This paper describes a framework for examining the effects of the cognitive complexity of tasks on language production and learner perceptions of task difficulty, and for motivating sequencing decisions in task-based syllabuses. Results of a study of the relationship between task complexity, difficulty, and production show that increasing the cognitive complexity of a direction-giving map task significantly affects speaker-information-giver production (more lexical variety on a complex version and greater fluency on a simple version) and hearer-information-receiver interaction (more confirmation checks on a complex version). Cognitive complexity also significantly affects learner perceptions of difficulty (e.g. a complex version is rated significantly more stressful than a simple version). Task role significantly affects ratings of difficulty, though task sequencing (simple to complex versus the reverse sequence) does not. However, sequencing does affect the accuracy and fluency of speaker production. Implications of the findings for task-based syllabus design and further research into task complexity, difficulty, and production interactions are discussed.
Finding simple, non-recursive, base noun phrase is an important step for many natural language processing applications. This paper presents a new corpus-based approach using decision tree for that purpose. In contrast to previous methods for Base NP identification, we adopt a decision tree trained from Penn Treebank to identify Base NP. And a self-learning mechanism is further integrated into our model. Experimental results show good performances using our method. The method can also be applied to processing of any other language.
CL Research's word-sense disambiguation (WSD) system is part of the DIMAP dictionary software, designed to use any full dictionary as the basis for unsupervised disambiguation. Official SENSEV AL-2 results were generated using WordNet, and separately using the New Oxford Dictionary of English (NODE). The disambiguation functionality exploits whatever information is made available by the lexical database. Special routines examined multiword units and contextual clues (both collocations, definition and example content words, and subject matter analyses); syntactic constraints have not yet been employed. The official coarsegrained precision was 0.367 for the lexical sample task and 0.460 for the all-words task (these are actually recall, with actual precision of 0.390 and 0.506 for the two tasks). NODE definitions were automatically mapped into WordNet, with precision of0.405 and 0.418 on 75 % and 70 % mapping for the lexical sample and all-words tasks, respectively, comparable to WordNet. Bug fixes and implementation of incomplete routines have increased the precision for the lexical sample to 0.429 (with many improvements still likely).
This paper introduces GLARF, a framework for predicate argument structure. We report on converting the Penn Treebank II into GLARF by automatic methods that achieved about 90% precision/recall on test sentences from the Penn Treebank. Plans for a corpus of hand-corrected output, extensions of GLARF to Japanese and applications for MT are also discussed.
Chunk parsing has focused on the recognition of partial constituent structures at the level of individual chunks. Little attention has been paid to the question of how such partial analyses can be combined into larger structures for complete utterances.The TüSBL parser extends current chunk parsing techniques by a tree-construction component that extends partial chunk parses to complete tree structures including recursive phrase structure as well as function-argument structure. TüSBL's tree construction algorithm relies on techniques from memory-based learning that allow similarity-based classification of a given input structure relative to a pre-stored set of tree instances from a fully annotated treebank.A quantitative evaluation of TüSBL has been conducted using a semi-automatically constructed treebank of German that consists of appr. 67,000 fully annotated sentences. The basic PARSEVAL measures were used although they were developed for parsers that have as their main goal a complete analysis that spans the entire input. This runs counter to the basic philosophy underlying TüSBL, which has as its main goal robustness of partially analyzed structures.
This document describes the syntactic bracketing guidelines for the Penn Korean Treebank, which is an online corpus of Korean texts annotated with morphological and syntactic information. The corpus consists of around 54,000 words and 5,000 sentences. The Treebank uses a phrase structure style of annotation, making head/phrasal node distinctions, argument/adjunct distinctions, and identifying empty arguments and traces for moved constituents. This document is organized as follows. In section 2, the basic syntactic ingredients of a clause structure are presented. Some notational conventions are introduced in section 3, including different types of syntactic tags, such as head level tags, phrase level tags and function tags used in the Treebank. In section 4, the bracketing guidelines for various types of clauses are discussed, including simple clauses, subordinate clauses, and clauses with coordination. Several types of subcategorizaion frames found in the Treebank are then presented in section 5, followed by bracketing guidelines for various linguistic phenomena in sections 6 to 21, including guidelines for annotating punctuation. The document ends with guidelines for handling some bracketing ambiguities and for handling some confusing examples.
The Japanese language, as spoken by native Japanese, has undergone a tremendous change in recent years-more noticeably in its spoken than in its written aspect. Some people brought up in the good old days deplore this, attributing it to the ignorance or negligence of decorum in speech among the younger generation. But the fact is that more and more adults are finding themselves unwittingly committing errors in usage, which they were formerly trained at school to avoid by all means. Among these deviations from linguistic norms, there are some which look likely to be established as perfectly acceptable usage, no matter whether one favors or disfavors them. Among them the most easily observable are: (1) recurrence of a rising intonation in mid-sentences, (2) omission of a morpheme, a word, or even a phrase, which was once considered indispensable in correct usage, (3) recurrence of what looks like an "empty" (semantically meaningless) word, (4) recurrence of what amounts almost to a cliche, and (5) prevalence of "feminine" (often infantile) language over strong "masculine" language, particularly in dialogs. What these phenomena reflect is, in the view of this paper-writer a kind of enervation in the verbal culture of the Japanese in general.
GlossLexer is a multi-user sign language lexical database integrating digital video that has been designed to support the compilation process for specialist dictionaries from data collection to production. Sign entries are identified by HamNoSys notations as well as glosses, but the user always has immediate access to video clips showing the signs as uttered by the informants.
In this paper we discuss the need for corpora with a variety of annotations to provide suitable resources to evaluate different Natural Language Processing systems and to compare them. A supervised machine learning technique is presented for translating corpora between syntactic formalisms and is applied to the task of translating the Penn Treebank annotation into a Categorial Grammar annotation. It is compared with a current alternative approach and results indicate annotation of broader coverage using a more compact grammar.
Parsing a natural language with its substantial structural complexity and ambiguity has turned out to be a puzzler. While the most of attempts in this area so far has relied on hand-generated parsers, difficulties inherent in the manual construction of natural language grammar lead up to efforts to induce the grammar automatically. Our approach to the automatic grammar induction presented in this paper has resulted in design and implementation of the system GRIND (Grammar Induction), which is capable to learn a sequence of context- -dependent parse actions from a given corpus of labelled derivation trees. To this end, GRIND combines two established methods of machine learning: transformation-based learning (TBL) and inductive logic programming (ILP). Being trained and tested on corpus SUSANNE, GRIND reached the accuracy of 96 % and the recall of 68 %. Keywords: grammar induction, inductive logic programming, transformation- -based learning 1
Chunk parsing has focused on the recognition of partial constituent structures at the level of individual chunks. Little attention has been paid to the question of how such partial analyses can be combined into larger structures for complete utterances. Such larger structures are not only desirable for a deeper syntactic analysis. They also constitute a necessary prerequisite for assigning function-argument structure.The present paper offers a similarity-based algorithm for assigning functional labels such as subject, object, head, complement, etc. to complete syntactic structures on the basis of prechunked input.The evaluation of the algorithm has concentrated on measuring the quality of functional labels. It was performed on a German and an English treebank using two different annotation schemes at the level of function-argument structure. The results of 89.73 % correct functional labels for German and 90.40 % for English validate the general approach.
We present a stochastic parsing system consisting of a Lexical-Functional Grammar (LFG), a constraint-based parser and a stochastic disambiguation model. We report on the results of applying this system to parsing the UPenn Wall Street Journal (WSJ) treebank. The model combines full and partial parsing techniques to reach full grammar coverage on unseen data. The treebank annotations are used to provide partially labeled data for discriminative statistical estimation using exponential models. Disambiguation performance is evaluated by measuring matches of predicate-argument relations on two distinct test sets. On a gold standard of manually annotated f-structures for a subset of the WSJ treebank, this evaluation reaches 79% F-score. An evaluation on a gold standard of dependency relations for Brown corpus data achieves 76% F-score.
Slang, as a social dialect, is a component of the language of anation. Language is a mirror of the social development and the importantmedium of culture. Slang, being a peculiar linguistic phenomenon, reflectedand is still reflecting dynamically the Russian society and culture. Slangexpressions, though not in conformity with the linguistic norms, do exist inoral Russian with peculiar linguistic features. Compared with standardRussian, slang expressions are coarse in rhetoric, but they stay in thesociety, the language and the culture, full of life.
One of the primary tasks of Information Extraction is recognizing all of the different guises in which a particular type of event can appear. For instance, a meeting between two dignitaries can be referred to as A meets B or A and B meet, or a meeting between A and B took place/was held/opened/convened/finished/dragged on or A had/presided over a meeting/conference with B