Papers reviewed and determined not to be word norm studies. Use the flag icon to report errors or suggest re-inclusion.
18265 papers
A series of studies was conducted to analyse the relationships between rating behaviour, rater goals and contextual variables, namely unit or class climate. Cleveland and Murphy (1992) suggested that errors and inter-rater disagreements in performance ratings could be understood in terms of differences in unit climate and the goals pursued by raters. In two studies, substantial correlations were found between self-rated goals and evaluations of instructor's performance; our second study provided evidence that goals measured before raters have an opportunity to observe performance are related to ratings obtained after observing instructor performance. A third study investigated Murphy and Cleveland's (1995) proposal that contextual variables, specifically organisational climate, affect rating behaviour. To test this hypothesis, data reflecting perceptions of the climate of college level courses and ratings of instructor performance were collected. Ratings of participative and co-operative climates showed a strong relationship with student ratings of instructors' performance. Average correlations between climate and ratings are substantially larger than those between rater goals and ratings; the mediation hypotheses tested in this study were not supported. Results of the three studies are discussed in relation to past research.
This paper presents a Java-based hyperbolic-style browser designed to render RDF files as structured ontological maps. The program was motivated by the need to browse the content of a web-accessible ontology server: WEB KB-2. The ontology server contains descriptions of over 74,500 object types derived from the WordNet 1.7 lexical database and can be accessed using RDF syntax. Such a structure creates complications for hyperbolic-style displays. In WEB KB-2 there are 140 stable ontology link types and a hyperbolic display needs to filter and iconify the view so different link relations can be distinguished in multi-link views. Our browsing tool, OntoRama, is therefore motivated by two possibly interfering aims: the first to display up to 10 times the number of nodes in a hyperbolic-style view than using a conventional graphics display; secondly, to render the ontology with multiple links comprehensible in that view.
In this paper we will present work carried out lately on the 50,000 words Italian Spontaneous Speech Corpus called AVIP, under national project API, made available for free download from the website of the coordinator, the University of Naples. We will concentrate on the tuning of the parser for Italian which had been previously used to parse 100,000 words corpus of written Italian within the National Treebank initiative coordinated by ILC in Pisa. The parser receives as input the adequately transformed orthographic transcription of the dialogues making up the corpus, in which pauses, hesitations and other disfluencies have been turned into most likely corresponding punctiation marks, interjections or truncation of the word underlying the uttered segment.\nThe most interesting phenomenon we will discuss is without any doubts "overlapping", i.e. a speech event in which two people speak at the same time by uttering actual words or in some cases nonwords, when one of the speakers, usually the one which is not the current turntaker, interrupts the current speaker.\nThis phenomenon takes place at a certain point in time where it has to be anchored to the speech signal but in order to be fully parsed and subsequently semantically interpreted, it needs to be referred semantically to a following turn.
This article describes a corpus-based investigation of quantifier scope preferences. Following recent work on multimodular grammar frameworks in theoretical linguistics and a long history of combining multiple information sources in natural language processing, scope is treated as a distinct module of grammar from syntax. This module incorporates multiple sources of evidence regarding the most likely scope reading for a sentence and is entirely data-driven. The experiments discussed in this article evaluate the performance of our models in predicting the most likely scope reading for a particular sentence, using Penn Treebank data both with and without syntactic annotation. We wish to focus attention on the issue of determining scope preferences, which has largely been ignored in theoretical linguistics, and to explore different models of the interaction between syntax and quantifier scope.
We use the grammatical relations (GRs) described in Carroll et al. (1998) to compare a number of parsing algorithms. A first ranking of the parsers is provided by comparing the extracted GRs to a gold standard GR annotation of 500 Susanne sentences: this required an implementation of GR extraction software for Penn Treebank style parsers. In addition, we perform an experiment using the extracted GRs as input to the Lappin and Leass (1994) anaphora resolution algorithm. This produces a second ranking of the parsers, and we investigate the number of errors that are caused by the incorrect 'GRs.
This paper reports on experiments in classifying the semantic role annotations assigned to prepositional phrases in both the Penn Treebank and FrameNet. In both cases, experiments are done to see how the prepositions can be classified given the dataset's role inventory, using standard word-sense disambiguation features. In addition to using traditional word collocations, the experiments incorporate class-based collocations in the form of WordNet hypernyms. For Treebank, the word collocations achieve slightly better performance: 78.5% versus 77.4% when separate classifiers are used per preposition. When using a single classifier for all of the prepositions together, the combined approach yields a significant gain at 85.8% accuracy versus 81.3% for word-only collocations. For FrameNet, the combined use of both collocation types achieves better performance for the individual classifiers: 70.3% versus 68.5%. However, classification using a single classifier is not effective due to confusion among the fine-grained roles.
Some studies with children have shown that there is no semantic priming at short stimulus onset asynchrony (SOA) in lexical decision and naming tasks for homographs. The predictions of spreading activation theories might explain this missing effect. There may be differences in children's and adults' memory structures. We have explored this hypothesis. The development of memory structure representations for homographs was measured by a Pathfinder algorithm. In Experiment 1, the three dependent variables were: the number of links in the network, closeness measures (C), and distances between nodes. Results revealed developmental differences in network structure representations in adults and children. In Experiment 2, results revealed that these differences were not due to the cohort effect. In Experiment 3, the relationship between associative strength, as measured by associative norms, and distances, as measured by Pathfinder algorithm, was explored. The results of these three experiments and empirical research from semantic priming experiments show that these differences in memory structure representations could be one of the sources of the missing semantic priming effect in children.
In this paper, we describe an approach to annotate the propositions in the Penn Chinese Treebank. We describe how diathesis alternation patterns can be used to make coarse sense distinctions for Chinese verbs as a necessary step in annotating the predicate-structure of Chinese verbs. We then discuss the representation scheme we use to label the semantic arguments and adjuncts of the predicates. We discuss several complications for this type of annotation and describe our solutions. We then discuss how a lexical database with predicate-argument structure information can be used to ensure consistent annotation. Finally, we discuss possible applications for this resource.
The paper presents a maximum entropy Chinese character-based parser trained on the Chinese Treebank ("CTB" henceforth). Word-based parse trees in CTB are first converted into character-based trees, where word-level part-of-speech (POS) tags become constituent labels and character-level tags are derived from word-level POS tags. A maximum entropy parser is then trained on the character-based corpus. The parser does word-segmentation, POS-tagging and parsing in a unified framework. An average label F-measure 81.4% and word-segmentation F-measure 96.0% are achieved by the parser. Our results show that word-level POS tags can improve significantly word-segmentation, but higher-level syntactic strutures are of little use to word segmentation in the maximum entropy parser. A word-dictionary helps to improve both word-segmentation and parsing accuracy.
In this paper we present a proposal to extend WordNet-like lexical databases by adding phrasets, i.e. sets of free combinations of words which are recurrently used to express a concept (let's call them recurrent free phrases). Phrasets are a useful source of information for different NLP tasks, and particularly in a multilingual environment to manage lexical gaps. Two experiments are presented to check the possibility of acquiring recurrent free phrases from dictionaries and corpora.
Intensified aquaculture has strong impact on fish health by stress and infectious diseases and has stimulated the interest in the orchestration of cytokines and growth factors, particularly their influence by environmental factors, however, only scarce data are available on the GH/IGF-system, central physiological system for development and tissue shaping. Most recently, the capability of the host to cope with tissue damage has been postulated as critical for survival. Thus, the present study assessed the combined impacts of estrogens and bacterial infection on the insulin-like growth factors (IGF) and tumor-necrosis factor (TNF)-α. Juvenile rainbow trout were exposed to 2 different concentrations of 17β-estradiol (E2) and infected with Yersinia ruckeri. Gene expressions of IGF-I, IGF-II and TNF-α were measured in liver, head kidney and spleen and all 4 estrogen receptors (ERα1, ERα2, ERβ1 and ERβ2) known in rainbow trout were measured in liver. After 5 weeks of E2 treatment, hepatic up-regulation of ERα1 and ERα2, but down-regulation of ERß1 and ERß2 were observed in those groups receiving E2-enriched food. In liver, the results further indicate a suppressive effect of Yersinia-infection regardless of E2-treatment on day 3, but not of E2-treatment on IGF-I whilst TNF-α gene expression was not influenced by Yersinia-infection but was reduced after 5 weeks of E2-treatment. In spleen, the results show a stimulatory effect of Yersinia-infection, but not of E2-treatment on both, IGF-I and TNF-α gene expressions. In head kidney, E2 strongly suppressed both, IGF-I and TNF-α. To summarise, the treatment effects were tissue- and treatment-specific and point to a relevant role of IGF-I in infection.
The article studies the particular features of the dynamics of lexical norms in Ukrainian language on the \nmaterials of dictionaries and mass-media. \nIt analyses the process of vocabulary enrichment by new lexical units, the phenomena of semantic \ntransformation and stylistic transposition.
BACKGROUND: Human affective responses appear to be regulated by limbic and paralimbic circuits. However, much less is known about the neurochemical systems engaged in this regulation. The mu-opioid neurotransmitter system is distributed in, and thought to regulate the function of, brain regions centrally implicated in affective processing. OBJECTIVE: To examine the involvement of mu-opioid neurotransmission in the regulation of affective states in healthy human volunteers. DESIGN: Measures of mu-opioid receptor availability in vivo were obtained with positron emission tomography and the mu-opioid receptor selective radiotracer [11C]carfentanil during a neutral state and during a sustained sadness state. Subtraction analyses of the binding potential maps were then performed within subjects, between conditions, on a voxel-by-voxel basis. SETTING: Imaging center at a university medical center. PARTICIPANTS: Fourteen healthy female volunteers. Intervention Sustained neutral and sadness states, randomized and counterbalanced in order, elicited by the cued recall of an autobiographical event associated with that emotion. MAIN OUTCOME MEASURES: Changes in mu-opioid receptor availability and negative and positive affect ratings between conditions. Increases or reductions in the in vivo receptor measure reflect deactivation or activation of neurotransmitter release, respectively. RESULTS: The sustained sadness condition was associated with a statistically significant deactivation in mu-opioid neurotransmission in the rostral anterior cingulate, ventral pallidum, amygdala, and inferior temporal cortex. This deactivation was reflected by increases in mu-opioid receptor availability in vivo. The deactivation of mu-opioid neurotransmission in the rostral anterior cingulate, ventral pallidum, and amygdala was correlated with the increases in negative affect ratings and the reductions in positive affect ratings during the sustained sadness state. CONCLUSIONS: These data demonstrate dynamic changes in mu-opioid neurotransmission in response to an experimentally induced negative affective state. The direction and localization of these responses confirms the role of the mu-opioid receptor system in the physiological regulation of affective experiences in humans.
Relationships between decoding and other reading abilities were investigated in experienced readers. Based on the distribution of decoding scores in a norming study with 5444 university students, 63 participants were selected, 30 high scorers and 33 low scorers. They were compared on measures of reading and related abilities: phonological awareness, nonword naming, spelling, working memory, print exposure, vocabulary, reading comprehension, reading rate, and oral reading of pseudotext. Significant differences were observed in phonological awareness, spelling, and accuracy but not speed of nonword and pseudotext reading, as well as print exposure and reading rate. ^ Two explanations of the relationship between decoding and comprehension differences were tested, each assuming a limited, common pool of cognitive resources that can be devoted to the component tasks of reading, ranging from word recognition through text comprehension. The Threshold Hypothesis emphasizes the importance of text content and structure, and predicts that when text difficulty exceeds some threshold, all processes that rely on working memory will be taxed to the point that efficiency of decoding limits comprehension outcomes. The Local Effects Hypothesis supposes that the basis of the difficulty lies in particular words in the text, predicting that text containing difficult-to-decode words will be comprehended less well than easily decoded text because more shared resources are diverted to the word-recognition process. Participants read two texts of equal difficulty, differing only in decoding demands. Consistent with the Threshold Hypothesis, no effect of decoding difficulty or participants' decoding skill was observed. ^ In the combined sample of participants, response time measures of basic skills (nonword reading and phonological awareness) accounted for 15% of the variance in reading rate. Pseudotext oral reading time explained an additional 13%, attributed chiefly to fluency in visual parsing of text. The time measures involving basic skills explained 12% of the variance in reading comprehension, and after controlling for general verbal ability, nonword reading time still accounted for significant variance in comprehension. The implications of phonological awareness and decoding differences for reading efficiency and comprehension are discussed in terms of phonological representations of lexical items that may vary in quality even among practiced readers. ^
The present study was designed to investigate the neuroendocrine modifications during affective states. In particular, we investigate if the pleasantness of the stimuli has a different effect on neuroendocrine responses. To address this issue, we compared the effects of pleasant, neutral, and unpleasant pictures on catecholamine, adrenocorticotrophic hormone (ACTH), cortisol, and prolactin plasma levels. Ten male participants were submitted to three experimental sessions, each on one of the three experimental days, a week apart in a counterbalanced order. Although in the subjective arousal rating, pleasant (erotic pictures) and unpleasant stimuli (pictures of mutilated bodies) receive the same high score, a different neuroendocrine pattern was obtained: unpleasant stimuli elicited a decrease in prolactin concentration and increases in noradrenaline, cortisol, and ACTH levels, whereas pleasant slide set viewing induced an increase in prolactin levels. The results suggest that the neuroendocrine system responds selectively to affective motivationally relevant pictures.
This study presents results from a corpus-based analysis of the expression of attitude, emotion, certainty and doubt (stance) in a large corpus of British and American conversation. Stance marker frequencies were assessed through an automated procedure for identifying stanced lexical items occur-ring in particular grammatical frames. The frequencies were analyzed with a multi-variate statistical procedure known as factor analysis which identifies co-occurrence patterns (factors). These factors can be understood to be the most salient moods of stance. Three factors were identified as character-istic: 1) informal AFFECT (American dialect-based), 2) boulomaic planning (American work-based) versus small talk (British dialect-based), and 3) hedged opinion (British dialect-based). Social norms were identified by examining the factors in light of discourse context and interpersonal relationships among speakers. Cross-cultural misunderstandings seemed particularly likely in work contexts, where Americans preferred boulomaic verbs (want, need), and British preferred evidentials (know, maybe). Differences in informal adult conversations are also potentially important, where Americans used many more affect markers (such as love, crazy). More work on pragmatic or functional domains using multi-variate analysis is proposed in the conclusion.
Past research has found that individual differences in both attitudinal and situational variables may be associated with males’ likelihood of acquaintance rape (LAR). The present research was conducted to examine the predictive value of both attitudinal and situational factors on males’ likelihood of forcing a female acquaintance to have non-consensual sexual intercourse. In Study 1, male and female respondents (Rs) were presented with a scenario depicting a hypothetical sexual interaction between the respondent and a newly acquainted member of the opposite sex. As the encounter progressed from one sexual activity to the next, Rs made three ratings regarding their own and partner’s intent to engage in each successive activity. The scenario ended with the female refusing further activity and males’ affect ratings, adherence to attitudes conducive of rape, and LAR were measured. Males’ initial perceptions of female sexual intent (to later engage in sexual intercourse) best predicted LAR. Study 2 was conducted to examine the role of female sexual communication on perceptions of consent to sexual intercourse. Rs were presented with a scenario similar to that in Study 1, but at each stage were requested to rate the extent to which the female had consented to engage in each sexual activity. Males completed the same affect and attitudinal measures. The results of Study 2 again suggested that males’ initial perception of female consent to (later) engage in sexual intercourse best predicted LAR. The present research suggests that further investigation into the role of situational factors and males’ initial perception of sexual intent and consent in the aetiology of acquaintance rape is required.
We investigated the reliability and validity of a video-based method of measuring the magnitude of children’s emotion-modulated startle response when electromyographic (EMG) measurement is not feasible. Thirty-one children between the ages of 4 and 7 years were videotaped while watching short video clips designed to elicit happiness or fear. Embedded in the audio track of the video clips were acoustic startle probes. A coding system was developed to quantify from the video record the strength of the eye-blink startle response to the probes. EMG measurement of the eye blink was obtained simultaneously. Intercoder reliability for the video coding was high (Cohen’sκ = .90). The average within-subjects probe-by-probe correlation between the EMG- and video-based methods was .84. Group-level correlations between the methods were also strong, and there was some evidence of emotion modulation of the startle response with both the EMG- and the video-derived data. Although the video method cannot be used to assess the latency, probability, or duration of startle blinks, the findings indicate that it can serve as a valid proxy of EMG in the assessment of the magnitude of emotion-modulated startle in studies of children conducted outside of a laboratory setting, where traditional psychophysiological methods are not feasible.
A course that relies on open-source software for teaching introductory computer programming and Web development to psychology graduate and advanced undergraduate students is described. The rationale, content, learning goals and outcomes of the course are described, along with the specific software used. The advantages of relying on open-source solutions rather than commercial software for implementing such a course are discussed.
We describe an algorithm for recovering non-local dependencies in syntactic dependency structures. The pattern-matching approach proposed by Johnson (2002) for a similar task for phrase structure trees is extended with machine learning techniques. The algorithm is essentially a classifier that predicts a non-local dependency given a connected fragment of a dependency structure and a set of structural features for this fragment. Evaluating the algorithm on the Penn Treebank shows an improvement of both precision and recall, compared to the results presented in (Johnson, 2002).
We investigate the performance of the Structured Language Model when one of its components is modeled by a connectionist model. Using a connectionist model and a distributed representation of the items in the history makes the component able to use much longer contexts than possible with currently used interpolated or backoff models, both because of the inherent capability of the connectionist model to fight the data sparseness problem, and because of the only sub-linear growth in the model size when increasing the context length. Experiments show significant improvement in perplexity and moderate reduction in word error rate over the baseline SLM results on the UPENN treebank and Wall Street Journal (WSJ) corpora respectively. The results also show that the probability distribution obtained by our model is much less correlated to regular N-grams than the baseline SLM model.
Metaphor is always a subject of central interest in the study of stylistics. But traditional stylistics has long focused only on lexical metaphor whereas systemic functional linguistics seldom extends to the stylistic function of grammatical metaphor, a term derived from the notion the form of the grammar relates naturally to the meanings that are being encoded (Halliday 1994: xvii). This paper seeks to go beyond and explore the stylistic value of grammatical metaphor. In functional stylistics, style is defined as a kind of prominence by systemically deviating the norms of the standard language. A natural corollary is that grammatical metaphor, as a systemic deviation of thecongruent lexi-cogrammatical forms, is of a certain stylistic value. Given such a working hypothesis, this paper first conducts a statistic analysis of the distribution of ideational, interpersonal and textual metaphors in 10 EST (English for science and technology) texts and 10 ENR (English news report) texts. The results reveal that the incongruity and deflection distribution patterns of grammatical metaphor are important criteria to differentiate one style from another. This paper proceeds to explore some properties of grammatical metaphor as a stylistic feature, including relativity, dynamism and fuzziness. In systemic functional linguistics, grammatical metaphor is defined in terms of forms, which are in turn defined in the three interrelated time frames of semogenesis: phylogenesis, ontogenesis and logogenesis. This suggests that congruent forms are relative, dynamic and fuzzy by definition and consequently that semogenesis is a process of metaphoricalization and demetaphoricalization. Now it may be rather difficult or even impossible to reconstruct the congruent counterparts for some metaphorical expressions. But this does not necessarily hinder the speaker/writer from using grammatical metaphor to achieve the desired stylistic effect. Given a more comprehensive statistic study of a larger corpus, we may have a better understanding of the stylistic value of grammatical metaphor and in turn a better understanding of the nature of style and a more solid prediction of the trend of language development.
Previous research (Aarts & Dijksterhuis, 2003) has shown that mental representations of situational norms (e.g., behaving quietly in libraries) and corresponding overt behaviors are capable of being automatically activated. Two experiments extended this line of research by investigating the conditional role of the tendency to conform to social norms in these effects. Participants explored a picture of a library and were given the goal to visit this library or not. Accessibility of representations of normative behavior was assessed in a lexical decision task. In the first experiment, individual differences in conformity to social norms were measured, whereas in the second experiment conformity was primed. Results indicated that the goal to visit the environment caused participants to automatically access representations of normative behavior. Importantly, in both experiments conformity was shown to moderate these accessibility effects: Automatic access to representations of normative behavior emerged when conformity tendencies were active.
This paper presents Advanced Glossing, a proposal for a general glossing format designed for language documentation, and a specific setup for the Shoebox-program that implements Advanced Glossing to a large extent. Advanced Glossing (AG) goes beyond the traditional Interlinear Morphemic Translation, keeping syntactic and morphological information apart from each other in separate glossing tables. AG provides specific lines for different kinds of annotation – phonetic, phonological, orthographical, prosodic, categorial, structural, relational, and semantic, and it allows for gradual and successive, incomplete, and partial filling in case that some information may be irrelevant, unknown or uncertain. The implementation of AG in Shoebox sets up several databases. Each documented text is represented as a file of syntactic glossings. The morphological glossings are kept in a separate database. As an additional feature interaction with lexical databases is possible. The implementation makes use of the interlinearizing automatism provided by Shoebox, thus obtaining the table format for the alignment of lines in cells, and for semi-automatic filling-in of information in glossing tables which has been extracted from databases
Since Eloise Jelinek has been interested in the issues of negation, focus and information structure, to the research of which she has contributed substantially, we want to use this nice occasion and present here partial results of an analysis of the Topic-Focus articulation (TFA) of Czech and of the impact of these results on inquiries into coreferrence in coherent discourse. In Czech linguistics, TFA has been systematically explored thanks to the classical Prague School of functional and structural linguistics. As reflecting the ‘given – new ’ strategy in discourse, TFA has been considered to belong to the main objects of linguistic study. Continuing the results gained by V. Mathesius, J. Firbas and others since the 1920s, the explicit linguistic descriptive framework characterized in Sgall et al. (1986), Hajičová (1993), Hajičová E., Partee B. and P. Sgall (1998) includes a possibility to describe TFA not only as concerning the intrinsic dynamics of the process of communication, patterned in the utterance (sentence occurrence), but also as constituting the structure of the sentence itself, i.e. grammar. Within this framework, TFA is understood as one of the basic aspects of (underlying) sentence structure, which characterizes the sentence as a unit of the interactive system of language; TFA thus is seen as a manifestation of the sentence being anchored in the context.
This paper reports on the use of two distinct evaluation metrics for assessing a stochastic parsing model consisting of a broad-coverage Lexical-Functional Grammar (LFG), an efficient constraint-based parser and a stochastic disambiguation model. The first evaluation metric measures matches of predicate-argument relations in LFG f-structures (henceforth the LFG annotation scheme) to a gold standard of manually annotated f-structures for a subset of the UPenn Wall Street Journal treebank. The other metric maps predicate-argument relations in LFG f-structures to dependency relations (henceforth DR annotations) as proposed by Carroll et al. (Carroll et al., 1999). For evaluation, these relations are matched against Carroll et al.'s gold standard which was manually annnotated on a subset of the Brown corpus. The parser plus stochastic disambiguator gives an F-measure of 79% (LFG) or 73% (DR) on the WSJ test set. This shows that the two evaluation schemes are similar in spirit, although accuracy is impaired systematically by mapping one annotation scheme to the other. A systematic loss of accuracy is incurred also by corpus variation: Training the stochastic disambiguation model on WSJ data and testing on Carroll et al.'s Brown corpus data yields an F-score of 74% (DR) for dependency-relation match. A variant of this measure comparable to the measure reported by Carroll et al. yields an F-measure of 76%. We examine divergences between annotation schemes aiming at a future improvement of methods for assessing parser quality.
Codeswitching is an important socio-cultural phenomenon in many multilingual and bilingual communities. In Malaysia, where there is a diverse linguistic panorama of speech communities, codeswitching is the norm. Malaysian speakers of English often codeswitch between Malay and English. The use of native words and expressions in Malaysian English is not only restricted to spoken English, but also found in written English. This article, based on a small-scale corpus study of Malaysian English, attempts to describe the use of Malay lexical items in Malaysian English by analysing the occurrence of these native items in texts written by Malaysian speakers of English. It also aims to explain the motivation behind their usage. The study suggests that Malaysian speakers of English use native lexical items in English texts as an intentional codeswitching strategy, motivated by the effects that the use of the native items has on the interpretation of the message.
This paper describes log-linear parsing models for Combinatory Categorial Grammar (CCG). Log-linear models can easily encode the long-range dependencies inherent in coordination and extraction phenomena, which CCG was designed to handle. Log-linear models have previously been applied to statistical parsing, under the assumption that all possible parses for a sentence can be enumerated. Enumerating all parses is infeasible for large grammars; however, dynamic programming over a packed chart can be used to efficiently estimate the model parameters. We describe a parellelised implementation which runs on a Beowulf cluster and allows the complete WSJ Penn Treebank to be used for estimation.
Digitizing and annotating texts and field recordings Given that several initiatives worldwide currently explore the new field of documentation of endangered languages, the E-MELD project proposes to survey and unite procedures, techniques and results in order to achieve its main goal, ''the formulation and promulgation of best practice in linguistic markup of texts and lexicons''. In this context, this year's workshop deals with the processing of recorded texts. I assume the most valuable contribution I could make to the workshop is to show the procedures and methods used in the Awetí Language Documentation Project. The procedures applied in the Awetí Project are not necessarily representative of all the projects in the DOBES program, and they may very well fall short in several respects of being best practice, but I hope they might provide a good and concrete starting point for comparison, criticism and further discussion. The procedures to be exposed include: * taping with digital devices, * digitizing (preliminarily in the field, later definitely by the TIDEL-team at the Max Planck Institute in Nijmegen), * segmenting and transcribing, using the transcriber computer program, * translating (on paper, or while transcribing), * adding more specific annotation, using the Shoebox program, * converting the annotation to the ELAN-format developed by the TIDEL-team, and doing annotation with ELAN. Focus will be on the different types of annotation. Especially, I will present, justify and discuss Advanced Glossing, a text annotation format developed by H.-H. Lieb and myself designed for language documentation. It will be shown how Advanced Glossing can be applied using the Shoebox program. The Shoebox setup used in the Awetí Project will be shown in greater detail, including lexical databases and semi-automatic interaction between different database types (jumping, interlinearization). ( Freie Universität Berlin and Museu Paraense Emílio Goeldi, with funding from the Volkswagen Foundation.)
In this paper we show how the trees in the Penn treebank can\nbe associated automatically with simple quasi-logical forms. Our approach is based on combining two independent strands of work: the first is the observation that there is a close correspondence between quasi-logical forms and LFG f-structures [van Genabith and Crouch, 1996]; the second is the development of an automatic f-structure annotation algorithm for the Penn treebank [Cahill et al, 2002a; Cahill\net al, 2002b]. We compare our approach with that of [Liakata and Pulman, 2002].
The sphere of language has become a privileged domain in which to interrogate the causes and effects of social injustice. (1) What is the right language to resist rape? Why is the language women use during rape frequently considered to be the wrong language? Such questions have a particular urgency in the context of women who have been raped by a man that they know. Why is their language the subject of particular scrutiny when they come before the law? How is it that the things these women say during rape can be used to turn violence into consensual sex? What relation between rape and language subtends the possibility of this transformation? Sharon Marcus, in the most influential reflection on the relation between rape and language, characterises feminist engagements with this question in the following manner: Whose words count in a rape trial? Whose 'no' can ever mean 'no'? How do rape trials condone men's misinterpretations of women's words? How do rape trials consolidate men's subjective accounts in objective 'norms of truth' and deprive women's subject accounts of cognitive value? Feminists have also insisted on the importance of naming rape as violence and of collectively narrating stories of rape. (2) This passage offers the central coordinates of the conventional feminist understanding of language in the context of rape. In particular, the primary coordinate here is words: words in general, which can count or not count, specific words like 'no' and category words which have the power to name. On this view, the problem with language in the context of rape is a problem of words, which are deprived of their power to designate the world according to women's experience. Australian criminologist Patricia Easteal articulates this position as she reflects on the relation between rape and After attending a meeting with other legal officers, she remarks that: [E]verytime... anyone present mentioned a judge or a lawyer, the pronouns 'he' and 'him' were used. I would contend that the use of language accurately reflects the reality that men continue to hold most of these positions in the criminal justice system. It goes deeper though. It mirrors the power that males have and use to maintain a stranglehold on the institutions and structures of Australian culture. And, the language in turn affects or even directs how we see the rest of our reality, including sexual assault and the law. (3) In this passage the problem of language is presented through a problem with words: words as a general category and specific grammatical categories like pronouns (4) which accumulate into the broad concept of masculine language. Here, language is understood as an aggregate of words (lexical items) which contain a self-evident signifying power. That is, words (connoting a reality) emit and effect their meanings, in a manner shorn of grammar, genre or context. (5) In other words, language is presented here in a theological formation: that is, in terms of its powers of naming. One of those things that language names (from a view) is sexual assault, and this act of naming will have a determining effect. Strategies for contesting this effect of words are not offered here, but would presumably include the use of non-gender specific pronouns and words that would reflect realities distinct to those inscribed by male power. (6) Dale Spender offers a famous solution to this linguistic-political problem when she argues that: 'Women need a word which renames male violence and misogyny and which asserts their blameless nature, a word which places the responsibility for rape where it belongs--on the dominant group'. (7) Spender's conclusion here implies that words can function as the symbols of political arguments: in this instance, new words about rape, invented by women, could (through condensation) denote already completed argumentative positions that clearly designate responsibility for sexual violence. …
ABSTRACT. A recent development in Chinese renders many occurrences of the construction [Adjunct PP + Verb + N[P.sub.2]] into [Verb + N2 + N[P.sub.1]]. A crucial difference between the two constructions is that in the former the verb and its N[P.sub.2] object can be separated. In the latter no separation is permitted, and N[P.sub.2] consists of only its head [N.sub.2]; that is, the verb and [N.sub.2], have formed a V-N COMPOUND. This paper attempts to account for such V-N compounding in Chinese. We suppose that the lexical structure representation of verbs in the Chinese V-N compound is similar to that of denominal verbs (Hale & Keyser 1993b). Thus, the formation of the Chinese V-N compound can be derived most simply by head movement, which is both morphologically driven and constrained by the Minimal Link Condition (Chomsky 1994). * INTRODUCTION. A recent development in Chinese renders many cases of [Adjunct PP (i.e. P + N[P.sub.1]) + Verb + N[P.sub.2]] into [Verb + [N.sub.2] + N[P.sub.1]], especially when the adjunct PP is locative (Hua 1997, Wang 1997, Xing 1997, Liu 1998a,b, Wang 1998), as shown by 1a-b, 2a-b and 3a-b. (1) (1) a. women xiang Niuyue qian ju we to New York move home [right arrow] b. women qian ju Niuyue (2) we move home New York 'We move to New York' (2) a. ta zai Hafu zhi jiao he at Harvard University engage teaching [right arrow] b. ta zhi jiao Hafu he engage teaching Harvard University He teaches at Harvard University' (3) a. Mali zai Jianada liu xue Mary in Canada engage study [right arrow] b. Mali liu xue Jianada Mary engage study Canada 'Mary studies in Canada' The crucial difference between the two constructions is that in [Adjunct PP + Verb + N[P.sub.2]] the verb and its N[P.sub.2] object can be separated by an aspect marker, a measure phrase, or a modifier of N[P.sub.2] (Li & Thompson 1981), as in 4a, 5a, and 6a. On the other hand, in [Verb + [N.sub.2] + N[P.sub.1]] no such separation is permitted, and N[P.sub.2] consists of its head noun [N.sub.2] only, as shown in 4b, 5b, and 6b. (4) a. women xiang Niuyue qian-le ju we to New York move-Asp home 'We have moved to New York' b. *women qian-le ju Niuyue we move-ASP home New York (5) a. ta zai Hafu zhi-le yi nian jiao he at Harvard University engage-ASP one year teaching 'He has taught at Harvard University for one year' b. *ta zhi-le yi nian jiao Hafu he engage-ASP one year teaching Harvard University (6) a. Mali zai Jianada liu-guo liang ci xue Mary in Canada engage-ASP two CL study 'Mary has studied in Canada twice' b. *Mali liu-guo liang ci xue Jianada Mary engage-ASP two CL study Canada In sum, the verb and [N.sub.2] in 1b, 2b, and 3b have formed a V-N COMPOUND. The V-N compound has, in fact, long existed in Chinese (Chao 1968), but until recently very few of them could take an NP object without being treated as ungrammatical or unacceptable (Li & Thompson 1981). By contrast, [[P + N[P.sub.1]] + Verb + N[P.sub.2]] has been regarded as the norm and has been strongly required or preferred (Hua 1997, Chow 2000). Since the 1970s, however, many cases of [[P + N[P.sub.1]] + Verb + N[P.sub.2]] have been transformed into [V-[N.sub.2] + N[P.sub.1]] (Xing 1997, Diao 1998). (3) Now the [V-[N.sub.2] + N[P.sub.1]] construction has become so common that it is no longer treated as ungrammatical or unacceptable but as a legitimate type of predicate structure (Gao 1998, Liu 1998a,b, Wang 1998). …
1006 Reviews 'national' or Parisian counterpart, and to give a clear idea of its distinctive place in the French media landscape. Martin manages to give this overview in a very readable manner and without being superficial. He acknowledges and draws on the excellent work that has been done on individual titles, periods, and geographical areas. A particularly welcome aspect is the significant space devoted to considering newspapers as eco? nomic and social entities, highlighting not just the Citizen Hersants and the starjour? nalists but also the networks of correspondents in the smallest ofvillages, the typographers, the delivery drivers, the sellers, and the readers. Especially fascinating fromthe perspective of social history is the analysis of the evolution, content, and role of the 'avis de deces' rubric: starting as simple quasi-administrative announcements, often appearing after the funeral, these came to be used as a substitute for the individual 'faire-part', then as a signifierof social status. They were also a major source of income fornewspapers. Martin is sensitive throughout to the impact of new technologies, up to and including the Internet, and notes that the regional press has often pioneered their use in France. The volume is impressively useable: there is an accurate general index and a separate index of newspaper titles, running to over seven pages; a useful chronology; a detailed table of contents; and an annotated summary bibliography to complement the abundant and detailed notes. These tools will help a range of readers make the most of a volume that achieves its purpose and invites furtherstudy. University of Leeds Paul Rowe La Neologie en francais contemporain: examen du concept et analyse de productions neologiquesrecentes. By Jean-Francois Sablayrolles. Paris: Champion. 2000. 588 pp.?86.90. ISBN 2-7453-0275-2. This is a scholarly and thought-provoking contribution to a field which has in the past suffered from either too narrow an academic approach, or from being subject to merely anecdotal treatment in amusing collections of neologisms. Jean-Francois Sablayrolles is ambitious, and largely successful, in his attempt to link broad theoret? ical discussion to a significant body of data. After a concise history of the notion of 'neologism' in Greek, Latin, and French, he reviews the differentapproaches to the subject by French linguists and then summarizes how twentieth-century theoretical models, from the structuralists to generativists and the most recent work of Melcu'k, have dealt with the processes of lexical creativity. He notes that, generally speaking, they have been assigned a very secondary and marginal role. In the second part of the book Sablayrolles proposes his own definitions of neologisme and neologie, and examines the types of unit and process that these involve. Perennial issues such as the role of dictionaries, upon which linguists have to rely, albeit often grudgingly, for their data, and the problem of differentiating between polysemy and homonymy, are given a fresh airing. More original are the brief dis? cussion of links between politico-cultural ideology and attitudes to neologisms, and speculation on the possibility of calculating the lifespan of a neologism. In the third part of the book the author analyses and compares the data that he has gathered from his six corpora, and ends with a discussion of the differentfunctions of neologisms. These range from their attention-catching use in newspaper headlines to their role in political polemics and their largely ludic function in the work of the writers R. Jorif and Ph. Meyer. Somewhat problematic is his inclusion of a corpus scolaire, drawn from the written work of secondary-school students. Many of these examples are non-standard verb forms such as ils croivent and il a acqueri. One can argue that these are not lexical, and possibly not new. Unlike the rest of his data, these forms are probably, as he concedes, neither 'voulus' nor 'conscients' (p. 317). Surely they MLRy 98.4, 2003 1007 are either part of a system which just happens to be differentfrom the norm or an attempt to conjugate a lexical item which is simply alien to the system? Theoretically at least, Sablayrolles appears to give the status of neologism to all new forms, what? ever their source or motivation. (Perhaps intentional, conscious creation...
In this article we focus on valency, which belongs to the core phenomena being captured in the underlying level of the Prague Dependency Treebank (PDT). We present a summary of the basic principles of the applied theoretical framework including proposals for suitable refinement relevant to NLP. The current status of description of valency behavior of verbs, nouns and adjectives is outlined. We present two branches of manual creation of a valency lexicon: (i) the PDT-VALLEX created during the annotation of the PDT and used primarily to obtain consistent annotation, and (ii) the Complex Valency Lexicon VALLEX, where the whole verbal lexemes are processed and other syntactically relevant information is assigned to particular valency frames.
This empirically based study aims at establishing the frequency, distribution, tenacity and the nature of errors produced by more than 700 advanced Swedish learners of French, at different university levels. The errors are classified in grammatical and lexical categories. Four grammatical categories (articles, nouns, pronouns, verbs) have been selected for a thorough examination, and examples of different types of errors are presented and commented upon. The error analysis follows, principally, the steps suggested by Corder (1974) and implies the discussion of such concepts as error, norm and usage.The verbs present the most frequent and tenacious errors constituting at each level about a third of all errors. The proportions of the different types of errors change according to task and level. The errors concerning tenses predominate in the translation tests, those concerning conjugations in the diagnostic test. As for pronouns, they cause 40 % of all errors in the diagnostic test (fill-in/transformation tests), 7 %-8 % in the translation tests. This reduction is partly due to the fact that many pronouns are necessarily regarded as part of the construction of a verb and classed as such. Put together, the errors concerning gender, orthography and vocabulary constitute 40 % at the advanced levels. The study includes a correlational analysis aiming at confirming or confuting some hypotheses regarding the interrelation, on the one hand between the results of the different parts of the two tests given during the first term of the university studies (vocabulary, grammar, metalingual knowledge), on the other hand beween the results and some external factors such as marks in French (secondary school), stay in French-speaking country, length of previous studies of French, etc. Among other things, it is shown that the students’ results in grammar are systematically better than those in vocabulary and that six years of previous studies of French do not generally lead to better results than three years’ studies.
MLRy 98.4, 2003 1007 are either part of a system which just happens to be differentfrom the norm or an attempt to conjugate a lexical item which is simply alien to the system? Theoretically at least, Sablayrolles appears to give the status of neologism to all new forms, what? ever their source or motivation. (Perhaps intentional, conscious creation could be a useful criterion to constrain an otherwise over-generous approach to the concept?) A detailed list of contents helps to guide the reader through this tightly structured book, although one could have wished for a fuller index. (It is restricted to some of the technical terms used in the text.) The bibliography, too, is incomplete, curiously excluding reference to the works by Jorifand Meyer which provide the data for two of his corpora. The appendix, however, is generous in listing all of the neologisms discussed, with details of their function and context. This alone makes forfascinating reading. Particularly ingenious examples are la markethique as an ever-topical social issue, and multiconjugal, referringto a serial spouse, both from Le Monde. Queen Mary, London Hilary Wise A New Life ofDante. By Stephen Bemrose. Exeter: University of Exeter Press. 2000. xxi + 249pp.?42-5o(pbk?i4.99). ISBN0-85989-583-1 (pbk0-85989-584-x). This book aims to provide an account of Dante's life combined with discussion of his writings that is accessible to both university students and non-specialist readers. The eleven chapters take Dante's life in a series of chronological sequences, from childhood and the meeting with Beatrice, to his involvement in politics and exile, and on to his post-exilic years. In most chapters, the chronological segments encompass Dante's literary activities, and the relevant works are usually given detailed com? mentary. Chapter 8 provides a separate survey of the Comedy. The volume, which includes a bibliography and guide to further reading, is written in a fluent and often lively style, even if there is the occasional moment when the tone is a little stilted ('about which more anon'; 'So are all great men of learning'). At a documentary level, the writing of a biography of Dante is a forbidding task, not least because many of the sources are provided by Dante himself, or else mediated often in highly partial form by later chroniclers, commentators, and biographers. On the whole, Bemrose finds his way through the material with a straightforward but sure-footed approach that gives considerable emphasis to Dante's socio-political and intellectual context. Dante emerges as a thinker,a political animal, and a consummate literary artist. The socio-political emphasis in particular makes fora portrait of Dante which, in its general outlines, is at times reminiscent of Bruni's 'documentary' life. Bemrose adds nothing new to Petrocchi's authoritative modern life, does not find space for Padoan's more recent attempts to backdate the Inferno, and offersno real new evidence or arguments to resolve issues of dating and attribution. Bemrose is not, however, to be faulted in these respects, forthe book does fulfil(and often admirably) its own aims and it is very well tailored to its readership, especially at an undergraduate level. One of the greatest strengths of the volume is its acute awareness of the knowledge gaps in its intended readers: Bemrose repeatedly offershelpful points of clarification on general contexts (especially intellectual and historical), specific issues (Guelphs and Ghibellines), and important terminology (e.g. plenary indulgence, curia, vernacular) that are often taken for granted in other general and introductory works on Dante. What is more, the summaries that he gives of Dante's works, es? pecially the Comedy, the De vulgari eloquentia, and the Convivio, are clear, accurate, judicious, and often highly readable. Almost inevitably, of course, any general account of this kind will elicit quibbles and prompt specialists to point out omissions: some may well have expected more 1008 Reviews discussion of the Vita nuova and its relationship to the Comedy; others may have wished for a fuller account of the philosophical and scientific content of the Rime petrose (which is surprising, given Bemrose's expertise in such matters); still others a more focused and critically informed discussion of allegory...