Papers reviewed and determined not to be word norm studies. Use the flag icon to report errors or suggest re-inclusion.
18265 papers
In this paper, we describe the application of a bidirectional dependency parser trained on the Turin University Treebank.
This paper presents recent extensions to Poliqarp, an open source tool for indexing and searching morphosyntactically annotated corpora, which turn it into a tool for indexing and searching certain kinds of treebanks, complementary to existing treebank search engines. In particular, the paper discusses the motivation for such a new tool, the extended query syntax of Poliqarp and implementation and efficiency issues.
Several incompatible syntactic annotation schemes are currently used by parsers and corpora in biomedical information extraction. The recently introduced Stanford dependency scheme has been suggested to be a suitable unifying syntax formalism. In this paper, we present a step towards such unification by creating a conversion from the Link Grammar to the Stanford scheme. Further, we create a version of the BioInfer corpus with syntactic annotation in this scheme. We present an application-oriented evaluation of the transformation and assess the suitability of the scheme and our conversion to the unification of the syntactic annotations of BioInfer and the GENIA Treebank.
Collocation is of great importance in dictionary compilation and natural language processing.Collocation extraction is one of the principal applications of corpus linguistics.Automatic extraction of bi-grams as candidate collocations is studied on Penn Treebank using the criteria of log likelihood,chi square and mutual information as association measure.The experimental results show the feasibility of the statistical methods.On the other hand,collocations extracted show different characteristics because of the different distribution assumptions by the three criteria.
Databases of hierarchically annotated text occupy a central place in linguistic research and language technology development. We describe a new approach to tree query which we call "Query by Annotation". Users express a query by annotating a tree, and the annotation is compiled into an expression in a path language. The result trees are overlaid with the original query, permitting the user to see why they match. Since queries and results are annotated trees, users can easily refine and resubmit their queries. The approach to Query by Annotation is motivated and exemplified using databases of linguistic trees, or treebanks.
Understanding user interests from text documents can provide support to personalized information recommendation services. Typically, these services automatically infer the user profile, a structured model of the user interests, from documents that were already deemed relevant by the user. Traditional keyword-based approaches are unable to capture the semantics of the user interests. This work proposes the integration of linguistic knowledge in the process of learning semantic user profiles that capture concepts concerning user interests. The proposed strategy consists of two steps. The first one is based on a word sense disambiguation technique that exploits the lexical database WordNet to select, among all the possible meanings (senses) of a polysemous word, the correct one. In the second step, a naïve Bayes approach learns semantic sensebased user profiles as binary text classifiers (userlikes and user-dislikes) from disambiguated documents. Experiments have been conducted to compare the performance obtained by keyword-based profiles to that obtained by sense-based profiles. Both the classification accuracy and the effectiveness of the ranking imposed by the two different kinds of profile on the documents to be recommended have been considered. The main outcome is that the classification accuracy is increased with no improvement on the ranking. The conclusion is that the integration of linguistic knowledge in the learning process improves the classification of those documents whose classification score is close to the likes / dislikes threshold (the items for which the classification is highly uncertain). 1
In this paper, we address the issue of improving a Chinese chunking system with rich lexicalized information. A method that incorporates statistical information based on distributional similarity between words obtained from large unlabeled corpus and morphological knowledge into a state-of-the-art CRF-based chunking model is proposed to tackle the data sparseness problem given limited amount of labeled training data. Evaluations are performed on the latest release of Chinese Treebank, and experimental results show that our method outperforms the chunking models based on features over word and automatically assigned POS tags when using the same amount of training data.
This paper proposes a novel Chinese syntactic parsing model based on semantic class, which is a variant of normal lexicalized statistical model. It attempts to make use of the syntactic and semantic similarity between Chinese words and then produces a more knowledgeable estimate of the probability of grammar rules. A simple but effective unsupervised method is designed to determine the proper semantic class of given words. Semantic class is used to improve the performance of parsing model. We evaluate our methods on the widely used Penn Chinese Treebank. Experimental results show that it outperforms a famous lexicalized model significantly on appropriate semantic class levels.
In the paper, we describe methods for exploitation of a new lexical database of valency frames (VerbaLex) in relation to Transparent Intensional Logic (TIL). We present a detailed description of the Complex Valency Frames (CVF) as they appear in VerbaLex including basic ontology of the VerbaLex semantic roles.
This paper presents the first steps towards a statistical syntactic analyzer for Basque. The system is based on a syntactically dependency annotated treebank and an adaptation of the deterministic syntactic analyzer of Nivre et al. (2007), which relies on a shift/reduce deterministic analyzer together with a machine learning module that determines which one of 4 analysis options to take, giving a unique syntactic dependency analysis of an input sentence. The results are near to those obtained by similar systems.
The lexicon of the modern Norwegian bokmal standard needs a better description and documentation than what is the situation today. The article presents a plan for building a modern lexical database based on a balanced corpus of 40 million words of modern bokmal. This base should serve as a source for a traditional scientific dictionary as well as a dictionary for language technological applications.
Psychocomputational Models of Human Language Acquisition (PsychoCompLA-2007) William Gregory Sakas (sakas@hunter.cuny.edu) Department of Computer Science, Hunter College Ph.D. Programs in Linguistics and Computer Science, The Graduate Center City University of New York 695 Park Ave, New York, NY 10021 USA David Guy Brizan (dbrizan@gc.cuny.edu) Ph.D. Program in Computer Science, The Graduate Center City University of New York 365 Fifth Avenue, New York, NY 10016 USA Keywords: language acquisition; syntax acquisition; language learning; language change; computational; linguistics; psycholinguistics; psychology; statistical; innateness. important question. One effective line of investigation is to computationally model the acquisition process and determine interrelationships between a model and linguistic or psycholinguistic theory, and/or correlations between a model's performance and data from linguistic environments that children are exposed to. Workshop Topic and History The workshop is devoted to psychocomputational models of language acquisition. By psychocomputational, we mean computational models that are compatible with research in psycholinguistics, developmental psychology and/or linguistics. Although there has been a significant amount of presented research targeted at modeling the acquisition of word categories, morphology and phonology, research aimed at modeling syntax acquisition has just begun to emerge. This is the third meeting of the Psychocomputational Models of Human Language Acquisition workshop following PsychoCompLA-2004, held in Geneva, Switzerland as part of the 20th International Conference on Computational Linguistics (COLING 2004) and PsychoCompLA-2005 as part of the 43rd Annual Meeting of the Association for Computational Linguistics (ACL-2005) held in Ann Arbor, Michigan where the workshop shared a joint session with the Ninth Conference on Computational Natural Language Learning (CoNLL-2005). Invited Presentations Statistical language learning: Computational and maturational constraints Elissa Newport, University of Rochester, USA The next challenges in unsupervised language acquisition: Dependencies and complex sentences Shimon Edelman, Cornell University, USA Learnable representations of languages: Something old and something new Alex Clark, Royal Holloway University of London, UK Workshop Description The workshop will present research and foster discussion centered around psychologically-motivated computational models of language acquisition, with an emphasis on the acquisition of syntax. In recent decades there has been a thriving research agenda that applies computational learning techniques to emerging natural language technologies and many meetings, conferences and workshops in which to present such research. However, there have been only a few (but growing number of) venues in which psychocomputational models of how humans acquire their native language(s) are the primary focus. Indirect evidence and the poverty of the stimulus Terry Regier, University of Chicago, USA Lexical learning and lexical diffusion Charles D. Yang, University of Pennsylvania, USA Bootstrapping bootstrapping Damir Cavar, Zadar University, Croatia and University of Indiana, USA Transformational networks Bob Frank, John Hopkins University, USA Psychocomputational models of language acquisition are of particular interest in light of recent results in developmental psychology that suggest that very young infants are adept at detecting statistical patterns in an audible input stream. Though, how children might plausibly apply statistical 'machinery' to the task of grammar acquisition, with or without an innate language component, remains an open and The great (Penn Treebank) robbery: When statistics is not enough Sandiway Fong, University of Arizona, USA, and Robert C. Berwick, MIT, USA (joint work with Partha Niyogi, University of Chicago, USA)
this paper, is not a language-specific one, in spite of the fact that most of the languages have their own repository for marking the role of the addressee in communicative utterances. In our opinion this linguistic phenomenon needs its adequate treatment because of two main reasons: 1. The vocative is supposed to be present on two levels: syntax and pragmatics. Therefore it needs more elaborate interpretation on the interface side, which, in HPSG, is more developed for morphology /syntax and syntax/semantics than syntax/pragmatics; 2. It will be useful for HPSG-oriented implementations, especially treebanks and dialogue systems. The paper is structured as follows: in the next section the status of the vocative in Bulgarian is discussed. In section 3 we propose our ideas on a unified treatment of vocatives. In section 4 the HPSG model is given. Section 5 outlines the conclusions and future work. 2 The Status of the Vocative in Bulgarian Vocatives are assumed to be restricted to the second person usage only. Usually they subsume the following two subtypes: calls (hey you) and addresses (Madam) [Levinson 1987, p. 71]. Bulgarian vocative role is usually treated within the opposition: vocative form (a remnant of the case paradigm) vs. base nominative form, i.e. with respect to the presence or loss of the special vocative inflections. Hence, The work reported here is done within the BulTreeBank project. The project is funded by the Volkswagen Stiftung, Federal Republic of Germany under the Programme "Cooperation with Natural and Engineering Scientists in Central and Eastern Europe" contract I/76 887. The authors wish to thank the Seminar fur Sprachwissenschaft of the Eberhard-Karls-Universitat, Tubingen, for hosting the writing of this paper, and the Internationales Zentr...
This laboratory experiment examined the effects of paired sensory cues that indicate the location of targets on target acquisition performance, the recall of information presented in concurrent visual and auditory communications, and perceived workload. The multimodal cueing techniques assessed in this study were Visual+Spatial Language, Visual+3-D Audio, Visual+Tactile, and Spatial Language+Tactile. A unimodal visual only cue was included as a baseline. Except for reaction times to cues, no significant differences were found between the multimodal cue conditions and the Visual Only mode in primary and secondary task performance or subjective workload. Reaction times were faster in the Visual+3-D Audio and the Visual+Tactile conditions than in modes that included a spatial language cue. Reaction times to the visual+spatial language cue were faster than the spatial language+tactile cue, but no significant differences were found between the Visual+Spatial Language and the Visual Only modes. Adding the 3-D audio cue to the visual cue significantly improved reaction time beyond that of the Visual Only condition, but no significant difference was found between the Visual Only and the Visual+Tactile modes. Reaction times to cues were slower when communications were presented visually, but no interaction was found between communications modality and cue condition on this measure. Communications modality, however, did have a different effect on subjective ratings of effort in the Visual+Tactile mode than in the other cue conditions. In the Visual+Tactile mode, ratings of effort were significantly lower when communications were presented auditorily than when they were presented visually, but communications modality did not appear to affect ratings of effort in the other cue conditions.
Activity within the visual cortex can be influenced by the emotional salience of a stimulus, but it is not clear whether such cortical activity is modulated by the affective status of the individual. This study used functional magnetic resonance imaging (fMRI) to examine the relationship between affect ratings on the Positive and Negative Affect Schedule and activity within the occipital cortex of 13 normal-weight women while viewing images of high calorie and low calorie foods. Regression analyses revealed that when participants viewed high calorie foods, Positive Affect correlated significantly with activity within the lingual gyrus and calcarine cortex, whereas Negative Affect was unrelated to visual cortex activity. In contrast, during presentations of low calorie foods, affect ratings, regardless of valence, were unrelated to occipital cortex activity. These findings suggest a mechanism whereby positive affective state may affect the early stages of sensory processing, possibly influencing subsequent perceptual experience of a stimulus.
Parsing unrestricted text is useful for many language technology applications but requires parsing methods that are both robust and efficient. MaltParser is a language-independent system for data-driven dependency parsing that can be used to induce a parser for a new language from a treebank sample in a simple yet flexible manner. Experimental evaluation confirms that MaltParser can achieve robust, efficient and accurate parsing for a wide range of languages without language-specific enhancements and with rather limited amounts of training data.
In the domain of video content retrieval, we present an approach for selecting words and phrases from highly imperfect automatically generated transcripts. Extracted terms are ranked according to their descriptiveness and presented to the user in a multimedia browser interface. We use sense querying from the WordNet lexical database for our method of text selection and ranking. Evaluation of 679 video summarization tasks from 442 users shows that the method of ranking and emphasizing terms according to descriptiveness results in higher accuracy responses in less time compared to the baseline of no ranking.
This paper tests three factors that have been held to be responsible for the variable stress behavior of noun-noun constructs in English: argument structure, semantics, and analogy. In a large-scale investigation of some 4500 compounds extracted from the CELEX lexical database (Baayen et al. 1995), we show that traditional claims about noun-noun stress cannot be upheld. Argument structure plays a role only with synthetic compounds ending in the agentive suffix - er. The semantic categories and relations assumed in the literature to trigger rightward stress do not show the expected effects. As an alternative to the rule-based approaches, the data were modeled computationally and probabilistically using a memory-based analogical algorithm (TiMBL 5.1) and logistic regression, respectively. It turns out that probabilistic models and the analogical algorithm are more successful in predicting stress assignment correctly than any of the rules proposed in the literature. Furthermore, the results of the analogical modeling suggest that the left and right constituent are the most important factor in compound stress assignment. This is in line with recent findings on the semi-regular behavior of compounds in other languages.
An approach for identifying the human source of a text by leveraging the significance of synonyms in language is presented. While others have attempted to identify authors in the past, they have focused on purely statistical approaches such as word length distribution, number of distinct words, and language models. We claim that an author's choice of synonyms is idiosyncratic and can be used in determining the identity of an author, which we demonstrate via our algorithm for recognizing authors. This algorithm uses synonym sets from the WordNet lexical database to give more weight to words that have many common synonyms. The results of this method applied to the task of identifying the authors of classic literature show that there is a correlation between an author's synonym choice and the author's identity. With this new author recognition technology, we may now explore new avenues of intelligent and meaningful interaction with users.
This paper presents a new approach to part-of-speech (POS) tagging in which the basic unit being tagged is a contiguous sequence of words rather than a single word. We run experiments on two different tagsets: the UPENN treebank and a treebank annotated with more ambiguous tags that have a semantic component. We show that the phrase-based system alone is a respectable tagger that exceeds the performance of the ME tagger on the ambiguous tagset. Moreover, when a log-linear model is built using features from both phrase-and word-based techniques, the tagging accuracy improved on both of our data sets yielding the highest reported performance to date on the more ambiguous tagset.
Assessments of the quality of parts of syntactic grammars of natural languages are useful for the validation of their construction. We extended a grammar of French determiners that takes the form of a recursive transition network and evaluated its quality. The result of the application of this local grammar gives deeper syntactic information than chunking or information available in treebanks. We performed the evaluation by comparison with a corpus independently annotated with information on determiners. We obtained 85 % precision and 93 % recall on text not tagged for parts of speech.
Emotional processes modulate the size of the eyeblink startle reflex in a picture-viewing paradigm, but it is unclear whether emotional processes are responsible for blink modulation in human conditioning. Experiment 1 involved an aversive differential conditioning phase followed by an extinction phase in which acoustic startle probes were presented during CS+, CS-, and intertrial intervals. Valence ratings and affective priming showed the CS+ was unpleasant postacquisition. Blink startle magnitude was larger during CS+ than during CS-. Experiment 2 used the same design in two groups trained with pleasant or unpleasant pictorial USs. Ratings and affective priming indicated that the CS+ had become pleasant or unpleasant in the respective group. Regardless of CS valence, blink startle was larger during CS+ than CS- in both groups. Thus, startle was not modulated by CS valence.
“The History of Error”:Hardy's Critics and the Self Unseen Jill Richards (bio) For one thing, many fine poems that have lyric moments are not entirely lyrical; many largely narrative poems are not entirely narrative; many personal reflections or meditations in verse hover across the frontiers of lyricism. ———Thomas Hardy, in the "Preface" to Select Poems of William Barnes1 Recalling a Gothic arch straying hazardously over its bounds, or the abrupt blankness of a wall where pattern anticipates a window, Hardy believed his poetry to reveal architectural moments of "cunning irregularity."2 For Hardy, the "unforeseen" arises as poetic conventions yield to the jarring peculiarities of the words themselves. Hardy's language radically displaces the speaking "I," so that past and present subjects are distinct and often contradictory rather than continuous with the speech that renders them. Playing upon the Victorian convention of a speaker in a location, Hardy eschews topographical coherence to redefine the poetic "I" as a subject misplaced in time, continually looking backward yet unseeing in the present moment. These lyric voices are not of the Victorian tradition from which they arise, nor do they anticipate a strictly modernist destruction of selfhood. Hardy's subjects are not fractured in a traditional sense, but are rather a response to the aberrances of their language. Distortion in poetic voice echoes the gaps of narrative sequence and idiom that so often appear in the poems. It is then neither language nor narrative voice that lends us a grammar of Hardy's poetry, but the way these structures stray out of their bounds, unfolding against one another. Yet Hardy's language has often gained the criticism that it hovers across poetic norms through ignorance and clumsiness not "cunning irregularity." The critical reception of Hardy's poetry has attracted notice for such rancor, one stemming from a misguided elitism that spurred even Lytton Strachey to uncharacteristically sour remarks.3 From the outset, F. R. Leavis found Hardy's unconventional language to be full of "gauche unshrinking mismarriages" of words, of the "prosaic banal, the stilted literary, the colloquial" jumbled together in a seemingly random process.4 What Leavis called a "style out of stylelessnesss" had been less vehemently articulated by William Archer as a [End Page 117] lack of "local and historical perspective in language, seeing all the words in the dictionary on one plane, so to speak, and regarding them all as equally available."5 For the earliest reviewers, Hardy's poetry was to be valued in spite of its language: the best poems were those that managed to shrug off such a handicap, so that a select canon of works might emerge unscathed from their marred diction and syntax. More recently, critics have found themselves praising Hardy's language for the awkwardness that was once condemned. Such purposeful moments of gracelessness are read as "authentic," and thus closer to the speech of a rural working-class, one that sits opposed to a purely literary language.6 In the end, Archer and Leavis seem closer to the truth after all, albeit inadvertently. Hardy's language acts on many different literary and historical planes indeed; it includes classes of language whose interactions are far too systematic to be simply "authentic." Ralph Elliott charts these linguistic features in Thomas Hardy's English (1984) proving Hardy's lexical irregularities to fall into patterns with specific literary purpose.7 Elaborating on the findings of Elliott in a more theoretical bent, Dennis Taylor's Hardy's Literary Language and Victorian Philosophy turns against critical tradition to show how Hardy's language is contextualized by its past and future, so that "awkwardness reflects those points where language is changing and where language is seeking new possibilities of precision" (p. 378). This is not double language—a pairing of learned and local, scholarly and sincere. Instead, it is one that works on historical levels, extending to use dialect falling in and out of usage, words old and new, familiar and coined. Thus, what Leavis calls the "mismarriage" of words keeps language from settling in a crystallized moment in history. This interaction is at once reaching out towards a past, acknowledging a philological history, and then acting...
To date, work on Non-Local Dependencies (NLDs) has focused almost exclusively on English and it is an open research question how well these approaches migrate to other languages. This paper surveys non-local dependency constructions in Chinese as represented in the Penn Chinese Treebank (CTB) and provides an approach for generating proper predicate-argument-modifier structures including NLDs from surface contextfree phrase structure trees. Our approach recovers non-local dependencies at the level of Lexical-Functional Grammar f-structures, using automatically acquired subcategorisation frames and f-structure paths linking antecedents and traces in NLDs. Currently our algorithm achieves 92.2 % f-score for trace insertion and 84.3 % for antecedent recovery evaluating on gold-standard CTB trees, and 64.7 % and 54.7%, respectively, on CTBtrained state-of-the-art parser output trees. 1
Discourse segmentation is the task of determining minimal non-overlapping units of discourse called elementary discourse units (EDUs). It can be further subdivided into sentence segmentation and sentence-level discourse segmentation. This paper addresses the latter, more challenging subtask, which takes a sentence and outputs the EDUs for that particular sentence. (1) Saturday, he amended his remarks to say that he would continue to abide by the cease-fire if the U.S. ends its financial support for the Contras. (1a) Saturday, he amended his remarks (1b) to say (1c) that he would continue to abide by the cease-fire (1d) if the U.S. ends its financial support for the Contras. In example (1), a sentence from a Wall Street Journal article taken from the Penn TreeBank corpus is further segmented into four EDUs, (1a), (1b), (1c) and (1d) (RST, 2002). Discourse segmentation, clearly, is not as easy as sentence boundary detection. The lack of consensus with regards to what constitutes an elementary discourse unit adds to the difficulty. Building a rule based discourse segmenter can be a tedious task since these rules would have to be based on the underlying grammar of the particular parser that is to be used. Therefore, we adopted a neural network model for automatically building a discourse segmenter from an underlying corpus of segmented text. We chose to use part-of-speech tags, syntactic information, discourse cues and punctuation. Our ultimate goal is to build a discourse parser that uses this discourse segmenter.
Proceedings of the Sixth International Workshop on Treebanks and \nLinguistic Theories. \nEditors: Koenraad De Smedt, Jan Hajič and Sandra Kübler. \nNEALT Proceedings Series, Vol. 1 (2007), 85-96. \n© 2007 The editors and contributors. \nPublished by \nNorthern European Association for Language \nTechnology (NEALT) \nhttp://omilia.uio.no/nealt. \nElectronically published at \nTartu University Library (Estonia) \nhttp://hdl.handle.net/10062/4476.
The is the reported phenomenon of increased spatial abilities after listening to that composer's music. However, subsequent research suggests that the Mozart effect may be an artifactual consequence of heightened arousal and mood rather than the music of Mozart per se (e.g., Thompson, Schellenberg, & Husain, 2001). The present study considers if performance improvements in a scored computer game are consistent with the mood and arousal hypothesis. Indeed, the use of a computer game as the experimental vehicle makes this work notably the most ecologically valid study of the Mozart effect to date. Specifically, in this work, ratings of musical preference as well as the game performance of individuals listening to different types of music are compared. If arousal and mood are the real we hypothesized that the performance level of participants would increase when listening to the selections they most enjoy. Results supported this hypothesis. ********** The Mozart Effect (e.g., Rauscher, Shaw, & Ky, 1993) is the reported phenomenon that listening to Mozart would temporarily increase spatial reasoning ability by the equivalent of 8-9 points on the Stanford-Binet. Rauscher, Shaw, and Ky explain the Mozart Effect by suggesting that exposure to musical compositions that are structurally complex excites certain cortical firing patterns comparable to those activated when completing spatial-temporal tasks. The Mozart Effect has also been seized upon by the media, and even distorted into the claim that simply listening to Mozart would make people smarter. Indeed, the idea that passively listening to Mozart might increase IQ scores has sparked the development of many educational books and music products (see McKelvie & Low, 2002). As Nantals and Schellenberg (1999) note, one Governor even budgeted for a compact disc or cassette for each infant born in his state. Despite issues with face validity, the Mozart Effect has been seriously discussed in such prestigious publications as Science and Nature, and still frequents the pages of respected psychology journals. At times, there have been problems replicating the basic but it has been suggested by Rauscher, Shaw, and Ky (1998) that inconsistent results by other researchers can be attributed to methodological differences. However, Nantais and Schellenberg (1999) had no difficulty replicating the basic finding: That is, they found a significant increase in performance on spatial-temporal tasks for subjects that heard a musical piece; but, there was no marked difference between those that heard Mozart or those who heard Schubert. Likewise, other researchers (e.g., Ashby, Isben, & Turken, 1999; Steele, Bass, & Crook, 1999) also observed that changes in mood can have a significant effect on cognitive performance, and that the original experimental conditions (e.g., listening to Mozart, relaxation music, or silence) likely each have an affect on mood and arousal. As such, the argument emerged that observed performance differences may occur due to improvements in mood and arousal rather than from neurophysiological priming. Consistent, Thompson, Schellenberg, and Husain (2001) reported that individuals that listened to Mozart performed better on spatial tasks, but also scored higher on positive mood and arousal ratings. Subjects that scored low on mood and arousal showed no effect of the music. By examining participant's spatial abilities after listening to a Mozart sonata (expected to produce positive mood), and an adagio by Albinoni (a sad piece), they were able to provide additional support for the arousal and mood hypothesis. In summary, the most current explanation for the Mozart Effect would suggest that an individual's mood/preference for a particular piece of music should correlate with any cognitive gains (Steele, 2000). In fact, if arousal and mood produce the effect, then equally pleasant stimuli other than music should have the same result. …
This article presents a transfer-based statistical model for Chinese to Taiwanese sign-language (TSL) translation. Two sets of probabilistic context-free grammars (PCFGs) are derived from a Chinese Treebank and a bilingual parallel corpus. In this approach, a three-stage translation model is proposed. First, the input Chinese sentence is parsed into possible phrase structure trees (PSTs) based on the Chinese PCFGs. Second, the Chinese PSTs are then transferred into TSL PSTs according to the transfer probabilities between the context-free grammar (CFG) rules of Chinese and TSL derived from the bilingual parallel corpus. Finally, the TSL PSTs are used to generate the possible translation results. The Viterbi algorithm is adopted to obtain the best translation result via the three-stage translation. For evaluation, three objective evaluation metrics including AER, Top-N, and BLUE and one subjective evaluation metric using MOS were used. Experimental results show that the proposed approach outperforms the IBM Model 3 in the task of Chinese to sign-language translation.
In this report an unsupervised and knowledge-based algorithm for concept sense disambiguation in concept maps is proposed. Concept maps are graphical tools for organizing and representing knowledge, based on concepts and labeled interconnections among them, forming propositions. The disambiguation process is carried combining Magnini’s domain, context information and the gloss. It’s supported in the Spanish WordNet lexical database and the lexical relations hypernyms-hyponyms, meronyms-holonyms and instance.
As the Introduction to this volume observes, sixteenth-century France is marked by ‘une vaste réflexion sur le bien dire’. This not only impacted upon theory and practice across the different literary genres using French but also promoted considerable debate on the form and basis of the emerging standard form of the vernacular. Given the wide-ranging nature of this réflexion, covering its different manifestations in one volume poses an almost insuperable challenge, but the twenty-eight contributions to the colloquium collected here certainly address an ambitiously broad span of topics and add usefully to our understanding of cultural developments in a period of major change. The papers are organized into three general sub-sections, ‘Interroger la norme’, ‘Évolutions de la norme’ and ‘Normes et société’. However, such is the fluid nature of the subject matter treated in certain papers that their classification under one or other of these headings can sometimes seem of doubtful appropriateness. The focus of the first sub-section is predominantly literary. The various contributions address the creation or adaptation of norms across a considerable number of different genres some of which are perhaps rather less familiar, for instance, oracular writings (Dubois), accounts of pilgrimages (Gomez-Géraud) and Jesuit letter-writing (Laborie). Particularly interesting is the close study by Duché of the influential approach which Nicolas Herberay adopted for translation. Herberay, an acknowledged master of French prose (‘un vray Cicero françois’, according to Jean Martin), wrote with a female as well as a male readership in mind, developing a prose style that was eloquent and natural that would set an example for bien dire in this area. The second sub-section begins with a cogent overview (Baddeley) of a familiar field, developments in orthography and the interplay between orthography and spelling, and is followed by a series of studies which explore revealingly topics such as the evolving relationship between poetics and grammar (Monferran), the increasing limitation on the use of metaphor in literary works (Cernogora) and developments in historiography (Dumontet). Perhaps the most interesting paper is the examination of the fortunes of the alexandrine in the early part of the century (Halévy). Particular attention is given to the writings of Jean Lemaire de Belges and Geoffroy Tory both of whom, on the basis of fanciful argumentation, sought to invest the alexandrine with special prestige and nationalistic symbolism matching the terza rima in Italian. Their exercises in myth-making were to contribute indirectly, it is argued, to the rapid rise in the alexandrine's use from around 1555. The final sub-section of the volume contains contributions that more particularly address linguistic issues. Notable amongst these are two items: a re-evaluation of the system of vers mesurés devised by Baïf which is seen as an attempt not only to reproduce the metrical patterns of ancient Greek but also to contribute towards the norms of spoken French by reflecting the élite ‘usage des Bons’ (Vignes); and a meticulous examination by Morin of change in the pronunciation norms presented by Peletier du Mans in his earlier works (1550, 1555) as against his 1581 Euvres poëtiques, the new norm correlating with that presented later in the works of Lanoue (1596) and La Touche (1696). Alongside these are a number of other attractive essays including a study of the linguistic norms in the speeches made at the formal opening of the Paris Parlement, with eloquence and high rhetoric dominating over practicality and clarity between 1560 and 1600 before a reversal occurred in the early seventeenth century (Petey-Girard), and an investigation of sixteenth-century liminaires (any text preceding a written work) composed by women where a complex set of norms operate involving humility, simplicity of style, the practice of dedicating the work to another woman and, in the light of the lack of image for the female writer, an attempt to ‘socialiser l'auteur’ (Gauthier). Completing the text is an Index Nominum and a table of contents. The diversity and scholarly depth of the volume should ensure that all seiziémistes will derive benefit from a close reading.
Reviewed by: Indian and British English: A handbook of usage and pronunciationby Paroo Nihalni, R. K. Tongue, Priya Hosali, and Jonathan Crowther Niladri Sekhar Dash Indian and British English: A handbook of usage and pronunciation. 2ndedn. By Paroo Nihalni, R. K. Tongue, Priya Hosali, and Jonathan Crowther. New Delhi: Oxford University Press, 2004. Pp. x, 260. ISBN 0195666569. $15.95. The present handbook is divided into two main parts. The first part (‘Lexicon of usage’) is designed to provide English users with information about the way in which certain words, idioms, collocations, phrases, and similar expressions of English used in India differ from British Standard English (BSE)—a model that has the closest affinity to Indian English. This part includes a thousand English words, which are used in a distinctive manner by large numbers of educated Indian speakers of English irrespective of their place, profession, education, gender, or other sociolinguistic factors. The words included in the handbook are selected from the speech or writing samples of the persons (such as university and school teachers, journalists, and radio commentators) who are likely to influence the English use of Indian learners. The handbook also contains many European words that have been Indianized over the years. Thus, it serves as a handy resource for Indian speakers of English, illustrating the many, often quite subtle, ways in which Indian English differs from standard British English usage, and where these differences are regarded as acceptable or substandard in the subcontinent. Examples in the handbook, which supplement the texts, are helpful to Indian users of English who are uncertain about the ‘correctness’ of their speech and writing, and serve those scholars who want to explore the differences between Indian and British uses of English. The book also has the potential to address special problems faced by learners of English, who are often impeded by the difficulties of recognizing finer nuances of meaning and usage. The second part of the handbook includes a brief report on the development of the pronunciation dictionary in India and abroad, followed by insightful discussions on standards of pronunciation in second/ foreign language teaching, the phonological systems of the British Received Pronunciation (BRP) and Educated Indian English (EIE), and the role of supraseg-mental properties (i.e. word stress, sentence stress, rhythm, intonation, etc.) in Indian English. The introduction contains a list of keywords for phonetic symbols used in the following part, ‘Dictionary of pronunciation’. Two types of pronunciation (Indian Recommended Pronunciation and the BRP) are supplied for more than two thousand words collected from the original lexical database of Michael West’s General service list of English wordstogether with a few additions compiled from the language resources available to the compilers. Each entry of the dictionary is tagged with relevant phonological information. This second edition also includes additional information on lexical collocation (the tendency of words to be used together in fixed phrases). In essence, the handbook not only serves as an invaluable reference guide for students and teachers of English, but also makes a valuable contribution for applied linguists, lexicographers, journalists, and scholars who write in Indian English. [End Page 465] Niladri Sekhar Dash Indian Statistical Institute, Kolkata Copyright © 2007 Linguistic Society of America
The paper aims at the complexity of syntactic network and the feasibility that the complex network work as a means of linguistic studies.The paper proposes the method how to build a syntactic based on dependency treebank and investigates the complexity of Chinese syntactic dependency network based on two Chinese treebanks with different genres.The results show that syntactic networks have similar average path length and diameter with the random networks,but cluster coefficients of syntactic networks are much greater than that of random networks,and degree distributions of syntactic networks also obey the power law.The paper reveals that two syntactic networks have the same diameter,but with different average degree,path length,cluster coefficients and power exponent.
Objective To validate the translated Chinese version of Pain Assessment in Advanced Dementia Scale (C-PAINAD) for its clinical application in assessing the discomfort level of severely cognitive impaired patients. Methods In developing the C-PAINAD, both clinical and academic experts were engaged in an iterative translation process for ensuring its semantic equivalence with its original language as well as it comprehensibility for clinical application. In establishing C-PAINAD's inter-rater reliability, its applicability for clinical use is further assessed in 11 severely cognitive impaired patients. The assessment tools included C-PAINAD, Discomfort Visual Analog Scale (DVA), and Philadelphia Geriatric Center Affect Rating Scale (PGCAR). Correlations, ANOVA and factor analysis were undertaken to examine the reliability and validity of C-PAINAD. Results The scores of C-PAINAD were not in normal distribution but clustered around Zero. C-PAINAD was positively correlated with DVA and negative affect but mildly and negatively correlated with positive affect. It was able to detect different pain level under different condition. One factor was extracted and the percentage of variance was 51.20%. C-PAINAD was proved to have satisfactory reliability and validity(Cronbach's α=0.66). Conclusions C-PAINAD was a simple, reliable and effective pain assessment instrument for measuring pain level of non-communicative patients like advanced dementia. Further research is deemed necessary to conduct with pre - and post-test of analgesic prescriptions prior to wider clinical application.
Semantic differential techniques are a useful, well-validated tool to assess affective processing of stimuli and determine how that processing is impacted by various demographic factors, such as gender. In this paper, we explore differences in connotative word processing between men and women as measured by Osgood's semantic differential and what those differences imply about affective processing in the two genders. We recruited 94 young participants (47 men, 47 women, ages 18-39) using an online survey and collected their affective ratings of 120 words on three rating tasks: Evaluation (E), Potency (P), and Activity (A). With these data, we explored the theoretical and mathematical overlap between Osgood's affective meaning factor structure and other models of emotional processing commonly used in gender analyses. We then used Osgood's three-dimensional structure to assess gender-related differences in three affective classes of words (words with connotation that is Positive, Neutral, or Negative for each task) and found that there was no significant difference between the genders when rating Positive words and Neutral words on each of the three rating tasks. However, young women consistently rated Negative words more negatively than young men did on all three of the independent dimensions. This confirms the importance of taking gender effects into account when measuring emotional processing. Our results further indicate there may be differences between Osgood's structure and other models of affective processing that should be further explored.
Models of social evaluation aim to capture the information people use to form first impressions of unfamiliar others. However, little is currently known about the relationship between perceived traits across gender. In Study 1, we asked viewers to provide ratings of key social dimensions (dominance, trustworthiness, etc.) for multiple images of 40 unfamiliar identities. We observed clear sex differences in the perception of dominance-with negative evaluations of high dominance in unfamiliar females but not males. In Study 2, we used the social evaluation context to investigate the key predictions about the importance of pictorial information in familiar and unfamiliar face processing. We compared the consistency of ratings attributed to different images of the same identities and demonstrated that ratings of images depicting the same familiar identity are more tightly clustered than those of unfamiliar identities. Such results imply a shift from image rating to person rating with increased familiarity, a finding which generalises results previously observed in studies of identification.
Enhanced conditionability has been proposed as a crucial factor in the etiology and maintenance of panic disorder (PD). To test this assumption, the authors of the current study examined the acquisition and extinction of conditioned responses to aversive stimuli in PD. Thirty-nine PD patients and 33 healthy control participants took part in a differential aversive conditioning experiment. A highly annoying but not painful electrical stimulus served as the unconditioned stimulus (US), and two neutral pictures were used as either the paired conditioned stimulus (CS+) or the unpaired conditioned stimulus (CS-). Results indicate that PD patients do not show larger conditioned responses during acquisition than control participants. However, in contrast to control participants, PD patients exhibited larger skin conductance responses to CS+ stimuli during extinction and maintained a more negative evaluation of them, as indicated by valence ratings obtained several times throughout the experiment. This suggests that PD patients show enhanced conditionability with respect to extinction.