Papers reviewed and determined not to be word norm studies. Use the flag icon to report errors or suggest re-inclusion.
18265 papers
Some of the speech databases and large spoken language corpora that have been collected during the last fifteen years have been (at least partly) annotated with a broad phonetic transcription. Such phonetic transcriptions are often validated in terms of their resemblance to a handcrafted reference transcription. However, there are at least two methodological issues questioning this validation method. First, no reference transcription can fully represent the phonetic truth. This calls into question the status of such a transcription as a single reference for the quality of other phonetic transcriptions. Second, phonetic transcriptions are often generated to serve various purposes, none of which are considered when the transcriptions are compared to a reference transcription that was not made with the same purpose in mind. Since phonetic transcriptions are often used for the development of automatic speech recognition (ASR) systems, and since the relationship between ASR performance and a transcription’s resemblance to a reference transcription does not seem to be straightforward, we verified whether phonetic transcriptions that are to be used for ASR development can be justifiably validated in terms of their similarity to a purpose-independent reference transcription. To this end, we validated canonical representations and manually verified broad phonetic transcriptions of read speech and spontaneous telephone dialogues in terms of their resemblance to a handcrafted reference transcription on the one hand, and in terms of their suitability for ASR development on the other hand. Whereas the manually verified phonetic transcriptions resembled the reference transcription much closer than the canonical representations, the use of both transcription types yielded similar recognition results. The difference between the outcomes of the two validation methods has two implications. First, ASR developers can save themselves the effort of collecting expensive reference transcriptions in order to validate phonetic transcriptions of speech databases or spoken language corpora. Second, phonetic transcriptions should preferably be validated in terms of the application they will serve because a higher resemblance to a purpose-independent reference transcription is no guarantee for a transcription to be better suited for ASR development.
The authors examine personality variables and interview format as potential antecedents of impression management (IM) behaviors in simulated selection interviews. The means by which these variables affect ratings of interview performance is also investigated. The altruism facet of agreeableness predicted defensive IM behaviors, the vulnerability facet of emotional stability predicted self- and other-focused behaviors, and interview format (behavior description vs. situational questions) predicted self-focused and defensive behaviors. Consistent with theory and research on situational strength, antecedent—IM relations were consistently weaker in a strong situation in which interviewees had an incentive to manage their impressions. There was also evidence that IM partially mediated the effects of personality and interview format on interview performance in the weak situation.
Application of egalitarian and prioritarian accounts of health resource allocation in low-income countries have both been criticized for implying distribution outcomes that allow decreasing/undermining health gains and for tolerating unacceptable standards of health care and health status that result from such allocation schemes. Insufficient health care and severe deprivation of health resources are difficult to accept even when justified by aggregative efficiency or legitimized by fair deliberative process in pursuing equality and priority oriented outcomes. I affirm the sufficientarian argument that, given extreme scarcity of public health resources in low-income countries, neither health status equality between populations nor priority for the worse off is normatively adequate. Nevertheless, the threshold norm alone need not be the sole consideration when a country's total health budget is extremely scarce. Threshold considerations are necessary in developing a theory of fair distribution of health resources that is sensitive to the lexically prior norm of sufficiency. Based on the intuition that shares must not be taken away from those who barely achieve a minimal level of health, I argue that assessments based on standards of minimal physical/mental health must be developed to evaluate the sufficiency of the total resources of health systems in low-income countries prior to pursuing equality, priority, and efficiency based resource allocation. I also begin to examine how threshold sensitive health resource assessment could be used in the Philippines.
PURPOSE: This study examined the effect of socioeconomic status (SES) on the early lexical performance of African American children. METHOD: Thirty African American toddlers (30 to 40 months old) from low-SES (n = 15) and middle-SES (n = 15) backgrounds participated in the study. Their lexical-semantic performance was examined on 2 norm-referenced standardized tests of vocabulary, a measure of lexical diversity (number of different words) derived from language samples, and a fast mapping task that examined novel word learning. RESULTS: Toddlers from low-SES homes performed significantly poorer than those from middle-SES homes on standardized receptive and expressive vocabulary tests and on the number of different words used in spontaneous speech. No significant SES group differences were observed in their ability to learn novel word meanings on a fast mapping task. CONCLUSION: The influence of socioeconomic background on African American children's lexical semantic tasks varies with the type of measure used.
This paper explores a parsimonious approach to Data-Oriented Parsing. While allowing, in principle, all possible subtrees of trees in the treebank to be productive elements, our approach aims at nding a manageable subset of these trees that can accurately describe empirical distributions over phrase-structure trees. The proposed algorithm leads to computationally much more tracktable parsers, as well as linguistically more informative grammars. The parser is evaluated on the OVIS and WSJ corpora, and shows improvements on efciency, parse accuracy and testset likelihood. 1 Data-Oriented Parsing Data-Oriented Parsing (DOP) is a framework for statistical parsing and language modeling originally proposed by Scha (1990). Some of its innovations, although radical at the time, are now widely accepted: the use of fragments from the trees in an annotated corpus as the symbolic grammar (now known as itreebank grammarsi, Charniak, 1996) and inclusion of all statistical dependencies between nodes in the trees for disambiguation (the iallsubtrees approachi, Collins & Duffy, 2002). The best known instantiations of the DOPframework are due to Bod (1998; 2001; 2003), using the Probabilistic Tree Substitution Grammar (PTSG) formalism. Bod has advocated a maximalist approach to DOP, inducing grammars that contain all subtrees of all parse trees in the treebank, and using them to parse unknown sentences where all of these subtrees can potentially contribute to the most probable parse. Although Bod’s empirical results have been excellent, his maximalism poses important computational challenges that, although not necessarily unsolvable, threaten both the scalability to larger treebanks and the cognitive plausibility of the models. In this paper I explore a different approach to DOP, that I will call iParsimonious Data-Oriented Parsingi (P-DOP). This approach remains true to Scha’s original program, by allowing, in principle, all possible subtrees of trees in the treebank to be the productive elements. But unlike Bod’s approach, P-DOP aims at nding a succinct subset of such elementary trees, chosen such that it can still accurately describe observed distributions over phrasestructure trees. I will demonstrate that P-DOP leads to computationally more tracktable parsers, as well as linguistically more informative grammars. Moreover, as P-DOP is formulated as an enrichment of the treebank Probabilistic Context-free Grammar (PCFG), it allows for much easier comparison to alternative approaches to statistical parsing (Collins, 1997; Charniak, 1997; Johnson, 1998; Klein and Manning, 2003; Petrov et al., 2006).
Proceedings of the Sixth International Workshop on Treebanks and \nLinguistic Theories. \nEditors: Koenraad De Smedt, Jan Hajič and Sandra Kübler. \nNEALT Proceedings Series, Vol. 1 (2007), 1-6. \n© 2007 The editors and contributors. \nPublished by \nNorthern European Association for Language \nTechnology (NEALT) \nhttp://omilia.uio.no/nealt. \nElectronically published at \nTartu University Library (Estonia) \nhttp://hdl.handle.net/10062/4476.
Proceedings of the 16th Nordic Conference \nof Computational Linguistics NODALIDA-2007. \nEditors: Joakim Nivre, Heiki-Jaan Kaalep, Kadri Muischnek and Mare Koit. \nUniversity of Tartu, Tartu, 2007. \nISBN 978-9985-4-0513-0 (online) \nISBN 978-9985-4-0514-7 (CD-ROM) \npp. 81-88.
Proceedings of the Sixth International Workshop on Treebanks and \nLinguistic Theories. \nEditors: Koenraad De Smedt, Jan Hajič and Sandra Kübler. \nNEALT Proceedings Series, Vol. 1 (2007), 189-200. \n© 2007 The editors and contributors. \nPublished by \nNorthern European Association for Language \nTechnology (NEALT) \nhttp://omilia.uio.no/nealt. \nElectronically published at \nTartu University Library (Estonia) \nhttp://hdl.handle.net/10062/4476.
We show that phrase structures in Penn Treebank style parses are not optimal for syntaxbased machine translation. We exploit a series of binarization methods to restructure the Penn Treebank style trees such that syntactified phrases smaller than Penn Treebank constituents can be acquired and exploited in translation. We find that by employing the EM algorithm for determining the binarization of a parse tree among a set of alternative binarizations gives us the best translation result. 1
We present two methods to address the problem of sparsity in the FrameNet lexical database. The first method is based on the idea that a word that belongs to a frame is ``similar'' to the other words in that frame. We measure the similarity using a WordNet-based variant of the Lesk metric. The second method uses the sequence of synsets in WordNet hypernym trees as feature vectors that can be used to train a classifier to determine whether a word belongs to a frame or not. The extended dictionary produced by the second method was used in a system for FrameNet-based semantic analysis and gave an improvement in recall. We believe that the methods are useful for bootstrapping FrameNets for new languages.
The lack of a large annotated systemic functional grammar (SFG) corpus has posed a significant challenge for the development of the theory. Automating SFG annotation is challenging because the theory uses a minimal constituency model, allocating as much of the work as possible to a set of hierarchically organised features.
The aim of this paper is to investigate whether a treebank grammar can be used to automatically classify and annotate German phrases contained in a MT lexicon. Phrases from the lexicon appear in their citation form and may differ structurally from the phrase tokens found in the corpus. We describe the grammar extraction process for a formalism called Tree-Generating Binary Grammar and evaluate the performance of subsets of the obtained grammar on a set of four types of lexical phrases.
Reviewed by: Treebanks: Building and using parsed corpora ed. by Anne Abeillé Philip Resnik Treebanks: Building and using parsed corpora. Ed. by Anne Abeillé. Dordrecht: Kluwer, 2003. Pp. 440. ISBN 1402013353. $74.95. Annotated corpora have been the fuel for a number of recent advances in the study of language, notably, although not exclusively, in computational linguistics. At the sentence level, corpus annotations range from shallow levels of linguistic representation, such as part-of-speech categories (Francis & Kučera 1982, Leech et al. 1994) or named entities (Strassel et al. 2003), through intermediate levels such as argument structure (Meyers et al. 2004, Palmer et al. 2005), to deeper levels of semantic representation such as semantic roles (Baker et al. 1998), word senses (Landes et al. 1998), events and temporal relations (Pustejovsky et al. 2003), or language-independent meaning representations (Farwell et al. 2004). When it comes to annotating sentences with linguistic representation, the sky is the limit (Meyers 2005). The ‘sweet spot’ in this range of annotations is occupied by treebanks, which is to say parsed (syntactically annotated) corpora. Unlike shallower annotations, syntactic parses capture hierarchical organization, a fundamental notion in virtually any modern theory of sentence structure. But unlike most deeper representations, syntactic parses can be created with high levels of inter-annotator reliability (though see Hovy et al. 2006 and references therein for recent progress in semantic treebanking). With respect to natural language processing applications, the impact of treebanks during the last decade has been remarkable, more than validating Marcus and colleagues’ (1993) premise that ‘significant, rapid progress can be made … by investigating those phenomena that occur most centrally in naturally occurring unconstrained materials and by attempting to automatically extract information about language from very large corpora’ (313). Beyond applications, syntactically annotated corpora have also begun to play an increasingly productive role in psycholinguistics, theoretical syntax, and language pedagogy (e.g. Corley et al. 2001, Jurafsky 2002, Dillon 2005, Meurers 2005, Resnik et al. 2005). In Treebanks: Building and using parsed corpora, Anne Abeille´ draws together a collection of fifteen short pieces focused primarily on the issues that come up in creating treebanks, demonstrated across an impressive variety of languages, along with six chapters on how treebanks are used. Although twenty-one chapters cannot be covered in detail in a short space, I present a brief walk through the chapters, followed by a discussion of the book as a whole. Abeillé’s introduction offers a very concise but clearly written primer on the main issues that come up in choosing representations, annotating corpora, and using treebanks in applications, folding in the obligatory pointers to the chapters that follow. In ‘The Penn Treebank: An overview’, Anne Taylor, Mitchell Marcus, and Beatrice Santorini provide a short, accessible description of the widely used Penn Treebank, extracting and updating the seminal article by Marcus and colleagues (1993) (which still remains the definitive source for in-depth discussion). In ‘Thoughts on two decades of drawing trees’, Geoffrey Sampson offers an engaging, personal discussion that combines elements of corpus description, linguistic analysis, and position paper. Unlike other chapters in the book, Sampson’s chapter offers a high-level look at corpus linguistics as an engineering and scientific discipline, and, contrary to some treebanking work, suggests that corpus annotation should make detail, accuracy, and explicitness a higher priority than the number of sentences annotated. Among the next thirteen chapters, nine provide detailed discussions of treebanking projects for specific languages, following a pattern that generally includes: (1) the goals and historical context of the project, (2) the selection of data to annotate (usually news text), (3) the annotation process (usually automatic analysis followed by human correction using project-specific tools), (4) details of representation (admitting wide variety, but usually an elaboration on syntactic constituency or grammatical dependency representation, with additional features to address language-specific [End Page 876] issues), (5) tools used for manual creation and correction of annotations (again admitting very wide variety), (6) a set of detailed, language-specific annotation choices that illustrate interesting and challenging aspects of the language under consideration, (7) a description of the project’s status, and (8) a brief evaluation and...
Students learned teaching principles either with or without (control group) the presentation of a classroom exemplar in video or text format. Across 2 experiments, the video group produced higher transfer scores and affective ratings than the other groups. Four weeks later, the video group recalled more information about the exemplar than the text group, but no treatment effects were found on transfer. Qualitative analyses (Experiment 2) showed that the video group produced a significantly larger number of modeled behaviors in the transfer test than the text (immediate) and control (immediate and delayed) groups. Results encourage using classroom video exemplars to promote students ’ affect and retention, but suggest that additional pedagogies are needed to promote longer term transfer of theory into practice.
This paper describes how a treebank of ungrammatical \nsentences can be created from a treebank of well-formed sentences. The treebank creation procedure involves the automatic introduction of frequently occurring grammatical errors into the sentences in an existing treebank, and the minimal transformation of the analyses in the treebank so \nthat they describe the newly created ill-formed sentences. \nSuch a treebank can be used to test how well a parser is able to ignore grammatical errors in texts (as people can), and can be used to induce a grammar capable of analysing such sentences. This paper also demonstrates the first of these uses.
In this paper, we report on the role of the Urdu grammar in the Parallel Grammar (ParGram) project (Butt, M., King, T. H., Niño, M.-E., & Segond, F. (1999). A grammar writer’s cookbook. CSLI Publications; Butt, M., Dyvik, H., King, T. H., Masuichi, H., & Rohrer, C. (2002). ‘The parallel grammar project’. In: Proceedings of COLING 2002, Workshop on grammar engineering and evaluation, pp. 1–7). The Urdu grammar was able to take advantage of standards in analyses set by the original grammars in order to speed development. However, novel constructions, such as correlatives and extensive complex predicates, resulted in expansions of the analysis feature space as well as extensions to the underlying parsing platform. These improvements are now available to all the project grammars.
This paper reports about our efforts in creating a tri-lingual parallel treebank. The focal points are consistency checking and all aspects of sub-sentential alignment. We discuss the alignment guidelines, the importance of quality checks, and special alignment problems. Then we look at alignment algorithms and alignment visualization tools and we compare our own TreeAligner with other alignment tools. Our constituent structure treebanks contain just over 1,000 sentences and around 18,000 tokens in each language.
782 SEER, 85, 4, OCTOBER 2007 television is largely Prague-based may have a more profound bearing on the use of language than has generally been appreciated. Not surprisingly, this study has many of the strengths and some of the weaknesses of a typical doctoral thesis. It offers a comprehensive summary and evaluation of existing research and provides very useful cross-references. It also highlights the complexity of language usage in a linguistic settingwhere stylisticallyand functionally divergent forms coexist, and where theprestigious 'standard' variant is not the spoken norm. Most importantly, it offers new statistical information to add to the existing body of data on morphological, phonological and lexical variation, and to substantiate claims that language choice always depends to a significant extent on the purpose of the dialogue and the formality of the situation.However, minor problems with editing and proof-reading detract from the overall quality of thework. Furthermore, the selection of television broadcasts inevitably contains a degree of subjectivity and is not indicative of the speech of the population as a whole. Finally, itwould appear that a lack of space may have prevented the author from developing some of hermore interesting ideas, such as the notion thatwomen may be treated differendy tomen in the television studio, and that this may be reflected in theiruse of language. In summary, despite some shortcomings, this is an original and stimulating study,which is of relevance to all scholars of language variation and change, and presents considerable scope for further research. School of Humanities, Languages and Social Sciences Tom Dickins Universityof Wolverhampton Pushkin, Alexander. 'TheGypsies' and Other NarrativePoems. Translated, with an introduction and notes, by Antony Wood. Engravings by Simon Brett. Angel Books, London, 2006. xl + 116pp. Notes.?14.95. Anyone who has ever attempted to translate nineteenth-century Russian verse into English should make a point of turning to the Afterword of Antony Wood's new book. Subtided 'Pushkin's Voice inEnglish', it is a pithy credo from one of the UK's leading translators of verse. One statement in particular should be writ large above any translator's desk: 'the rise of translation theory in recent decades has not been accompanied by the emergence of any sub stantial body of translation of Pushkin's verse that has impressed as verse in English' (p. 114). Wood makes clear how he intends to remedy thisdeficiency. To begin with he is uncontroversial. He will eschew alternating masculine and feminine rhymes as being too difficult to achieve inRussian. He will have recourse to half-rhymes, since rhymes are far easier to find in an inflected language than in an uninflected language. His other points, however, are more contentious. He is clearly no enthusiast for translations which reproduce exactly themetre and rhyme scheme of the original, considering that their effect 'tends to be self-conscious, self-satisfied,unengaged and disembodied' (p. in). Nor does he think that the number of lines of the original should necessarily be maintained. reviews 783 The Afterword is one of the items added toWood's earlier work 'The Bridegroom', with 'Count Nulin} and 'TheTale of the Golden CockereT,published by Angel Books in 2002 and reviewed in this journal (vol. 82, July 2004, no. 3). These three poems are reproduced here with slight amendments, one of which, fromThe Bridegroom,isparticularly felicitous.Whereas in 2002 we find in the tenth stanza of the poem 'and then a jet/Over Natasha's head', the revised translation reads 'then splash a/Dash of iton Natasha. The new book is some twice the length of the earlier book. The new translations are Pushkin's first 'problem' poema,The Gypsiesand the skazka,The Tale of the Dead Princess and theSevenChampions.True to his credo,Wood does not attempt to replicate Pushkin's iambic tetrameter throughout his transla tion of The Gypsies. His favoured departure from this involves removing the initial unstressed syllable and turning the line into trochaic tetrameter. There are numerous examples of the type 'Life resounds on every side' (p. 3 ). In addition there are variants of this variant, all scrupulously noted in the Afterword. These departures from Pushkin's metre are clearly no oversight and Wood shows considerable expertise in producing...
Proceedings of the Sixth International Workshop on Treebanks and \nLinguistic Theories. \nEditors: Koenraad De Smedt, Jan Hajič and Sandra Kübler. \nNEALT Proceedings Series, Vol. 1 (2007), 115-126. \n© 2007 The editors and contributors. \nPublished by \nNorthern European Association for Language \nTechnology (NEALT) \nhttp://omilia.uio.no/nealt. \nElectronically published at \nTartu University Library (Estonia) \nhttp://hdl.handle.net/10062/4476.
This paper aims to explore the norms, strategies and procedures of translating lexical doublets in Arabic literary discourse. Lexical doublets are sets of two (near-) synonyms connected with ﻮ ‘and’, ﺃﻮ ‘or’, or the zero article. The empirical basis material for this study consists of a three-part autobiography ( al-Ayyām, ‘The Days’) and a narrative ( Hadīth ´Īsā ibn Hishām, ‘´Īsā ibn Hishām’s Tale’). Findings show that patterns of repetition are shifted in the English translations, and various translation strategies are applied, the most common being grammatical transposition and reduction. A quantitative analysis of the translation of lexical doublets in three samples is also conducted. The samples are about 2500 words each, randomly selected from the three parts of the autobiography. The figures indicate that one translator (that of Part One) adopts a source text-oriented strategy while the other two translators prefer a shifting strategy. This may be seen as a useful indicator of the translations’ orientation towards either adequacy or acceptability (Toury 1995).
Computer games potentially offer a useful research tool for psychology but there has been little use made of them in assessing cognitive abilities. Two studies assessing the viability of a computer game-like test of cognitive processing speed are described. In Experiment 1, a computerized coding task that uses a mouse response method (McPherson & Burns, 2005) was the basis for a simple computer game-like test. In Experiment 2, dynamic game-like elements were added. Validity was assessed within a factor analytic framework using standardized abilities tests as marker tests. We conclude that computer game-like tests of processing speed may provide an alternative or supplementary tool for research and assessment. There is clearly potential to develop game-like tests for other cognitive abilities.
This work is about multimodal and expressive synthesis on virtual agents, based on the analysis of actions performed by human users. As input we consider the image sequence of the recorded human behavior. Computer vision and image processing techniques are incorporated in order to detect cues needed for expressivity features extraction. The multimodality of the approach lies in the fact that both facial and gestural aspects of the user’s behavior are analyzed and processed. The mimicry consists of perception, interpretation, planning and animation of the expressions shown by the human, resulting not in an exact duplicate rather than an expressive model of the user’s original behavior.
OBJECTIVE: The current literature highlights the research and clinical applications of parental report in investigating the status of language skills in young children. Since language acquisition norms for Maltese have not yet been established, this study attempts to obtain preliminary indications of developmental trends in early lexical development by adapting an established parent-completed vocabulary checklist for use with Maltese children. PATIENTS AND METHODS: The concurrent validity of this bilingual adaptation was examined relative to picture naming abilities and spontaneous vocabulary use in a cross-sectional cohort of 10 children aged between 12 and 30 months who were primarily exposed to Maltese. RESULTS: The results indicate a high and significant correlation between lexical production abilities as reported by parents completing the checklist and as measured through confrontation naming and conversational language use. Reported vocabulary measures indicate a steady increase in lexical production with age, with a sharp increment evident beyond the age of 24 months. CONCLUSION: These findings suggest that the preliminary version of the vocabulary checklist has potential for gauging early lexical growth and point towards the need for further research on a larger scale.
Proceedings of the Sixth International Workshop on Treebanks and \nLinguistic Theories. \nEditors: Koenraad De Smedt, Jan Hajič and Sandra Kübler. \nNEALT Proceedings Series, Vol. 1 (2007), 19-30. \n© 2007 The editors and contributors. \nPublished by \nNorthern European Association for Language \nTechnology (NEALT) \nhttp://omilia.uio.no/nealt. \nElectronically published at \nTartu University Library (Estonia) \nhttp://hdl.handle.net/10062/4476.
We modified the traditional (verbal) digit span task for administration via computer and the Internet. This online version collects data on the floor and ceiling of a subject’s span capacity, rather than generating a rough estimate of capacity based on a 50% success rate, as the traditional version does. We compared the two versions within adult subjects in two cohorts: college-age normal readers and college-age reading-disabled readers. To explore the reliability of the online version as a research tool, we employed the Bland-Altman approach to examine agreement between instruments. The online version yielded spans similar to those yielded by the traditional version, tending toward smaller values at the high end and larger values at the low end of span sizes, in the typical readers. It differentiated between the better and poorer readers reliably, and to the same extent as does the verbal version. The online version of the digit span task is comparable to the traditional version in assessing verbal span capacity; the code with which to implement the task is available at www.psychonomic.org/archive.
You have accessThe ASHA LeaderFeature1 Jul 2007Ethnographic and Sociolinguistic Aspects of Communication: Research-Praxis Relationships José G. Centeno, Raquel T. Anderson, M. Adelaida Restrepo, Peggy F. Jacobson, Jackie Guendouzi, Nicole Müller, Ana Inés Ansaldo and Karine Marcotte José G. Centeno Google Scholar More articles by this author, Raquel T. Anderson Google Scholar More articles by this author, M. Adelaida Restrepo Google Scholar More articles by this author, Peggy F. Jacobson Google Scholar More articles by this author, Jackie Guendouzi Google Scholar More articles by this author, Nicole Müller Google Scholar More articles by this author, Ana Inés Ansaldo Google Scholar More articles by this author and Karine Marcotte Google Scholar More articles by this author https://doi.org/10.1044/leader.FTR2.12092007.12 SectionsAbout ToolsAdd to favorites ShareFacebookTwitterLinked In Effective communication requires the integration of multiple factors, including linguistic, cultural, cognitive, and neurological variables. Ethnography and sociolinguistics may enhance our understanding of how those factors interact. Ethnography is the systematic, qualitative study of culture, including the cultural bases of linguistic skills and communicative contexts (Ochs & Schieffelin, 1995). Sociolinguistics, on the other hand, focuses on how language use is shaped by individual and societal forces (Coulmas, 1997). As examples, ethnographic research may examine discourse and vocabulary trends in a specific cultural group; sociolinguistic studies may focus on language input differences in bilingual development or age-related speech variation (Ball, 2005). Although the separation between ethnography and sociolinguistics is not always clear (Salzmann, 1993), the application of ethnographic and sociolinguistic principles to speech-language pathology research and practice is critical. Ethnographic and sociolinguistic descriptions point to key relationships in the inextricable links among culture, language, communication, and cognition. Language development, communication acts, and concomitant thought processes are affected by the cultural world in which we live (Centeno, 2007b). Ethnographic and sociolinguistic analysis expands our understanding of an individual’s communication history, language profile, and psycholinguistic processing (Ball, 2005; Centeno, 2007b), and has particular significance in our increasingly diverse clinical caseloads. Based on monolingual and bilingual speakers, current research and theory illustrate how approaches grounded in the ethnographic and sociolinguistic realities of language and communication can enrich experimental methodology, theory-building, and evidence-based practices in speech-language pathology. Child Language Analysis The analysis of spontaneous language samples is a critical tool for SLPs involved in research and pediatric practice. Knowledge of the cultural and sociolinguistic contexts in which children acquire and use language enhances the productive elicitation and accurate analysis of children’s language skills. This knowledge includes topics, conversational formats, and tasks that maximize language productivity, as well as analysis techniques that consider acquisitional variables. Language samples from Latino children in the United States can serve as an illustration. Latino children constitute the nation’s largest minority population receiving speech-language services in pediatric settings (e.g., Roseberry-McKibbin et al., 2005). These children also represent diverse social, cultural, educational, and linguistic backgrounds, which translate into considerable variability in cultural norms, literacy experiences, discourse styles, Spanish dialects, and levels of bilingualism (McCabe & Bliss, 2003; Zentella, 2005). To stimulate productivity in the collection of language samples, clinicians need to acknowledge language socialization practices consistent with a child’s developmental background. For example, the discourse of Mexican-American families frequently focuses on the family. Storytelling as entertainment also is common in Mexican-American homes (McCabe & Bliss, 2003); consequently, family-related topics in storytelling may lead to greater expressive output when used with Mexican-American children. Similarly, elicitation of appropriate language requires suitable techniques. Although there is limited research on language elicitation methods used with Latino children, some studies have pointed to effective strategies. For example, Latino children can be successful story retellers from preschool age, particularly when they have training or models (Fiestas & Peña, 2004; Gutiérrez-Clellen & Hofstetter, 1994; Restrepo, 1998). Story retelling can be more fruitful than spontaneous story production in eliciting language in both preschool and school-age Latino children (Castilla & Restrepo, 2004). In fact, story retelling combined with parent reports provide the best identification of Spanish-speaking children with language disorders (Restrepo, 1998; Restrepo et al., 2005). The accurate examination and diagnostic assessment of language samples must be grounded in the sociolinguistic contexts affecting language input during acquisition. Sound language analysis can help SLPs understand children’s developmental linguistic changes in monolingual and bilingual contexts and assess post-intervention linguistic outcomes. For example, sensitive measures of language growth in preschool Mexican-American Spanish-speaking children can include Spanish mean length of utterance (MLU) in words and subordination index (number of dependent clauses per sentence) obtained from story retellings. Further, preschool Spanish-speaking children receiving bilingual intervention have shown significant productivity in these two measures within the same school year as compared with children in English-only language interventions (Castilla & Restrepo, 2004). Language analysis also can detect cross-linguistic effects or grammatical changes caused by the unequal use of languages in bilingual environments. It is critical to differentiate between linguistic limitations caused by a disorder from linguistic features related to language use in bilingual communication. For example, Spanish sentence length—prior to the acquisition of English as a second language—can predict growth in English grammar in preschool children who speak Spanish (Castilla & Restrepo, 2004). In situations of language loss (attrition), the complexity of certain linguistic elements—such as verbs in Spanish—may weaken as children develop proficiency in English and use Spanish less frequently (Anderson, 2001, 2004). Additionally, language-disordered Spanish-English children in educational programs may demonstrate different attrition patterns in Spanish, their first language (L1). When examining the grammatical profiles of two bilingual Spanish-English children with language impairment, Restrepo (2003) found different patterns of L1 (Spanish) loss; one child exhibited growth in MLU while his utterances increased in errors. The other child demonstrated a decrease in MLU and a decrease in errors per utterance. Linking ethnographic and sociolinguistic factors to language sampling facilitates appropriate methodology and diagnostic interpretations of children’s grammatical development. Given the variability in acquisitional scenarios across sociocultural contexts, much research in language sampling in specific groups of monolingual and bilingual children is required before generalizations can be made. Typical Discourse Routines Sociolinguistic descriptions of language use provide plausible theoretical grounds to interpret psycholinguistic processing in speakers with expressive restrictions, as exemplified by verb use. Monolingual Spanish-speaking children, for example, oscillate between the present and the past in choice of dominant verb tense until age 5; the present appears to stabilize as the most frequently used tense in the spoken narratives of older children and adults (Sebastián & Slobin, 1994). Similarly, discourse analysis of Spanish conversational adult narratives revealed that—despite the frequent alternated use of the present, past, and imperfect tenses—the past-present alternation emerged as the most prominent tense shift (Silva-Corvalán, 1983). In speakers with compromised expressive resources, the early emergence and frequent use of simple verb forms in speech may combine to maximize access and production of such verbs. Such speakers include monolingual children with expressive delays, bilingual speakers experiencing L1 loss, and aphasic speakers with limited syntactic production in their oral expressions. Spanish-speaking adults and children with typical development and language impairment may use similar verb tenses in their narratives (Jacobson, 2006). Both groups tend to use the past (Yo caminé, “I walked”), imperfect (Yo caminaba, “I used to walk”), and present (Yo camino, “I walk”) tenses more frequently than any other verb tenses in story retelling tasks. Similarly, bilingual Spanish-English children and adults experiencing L1 (Spanish) loss show a greater use of simple verb tenses in their spoken discourse—the present, the present progressive (Yo estoy caminando, “I am walking”), and the past tense (Anderson, 2001; 2004; Silva-Corvalán, 1991). Also, monolingual Spanish-speaking individuals with Broca’s aphasia who experience agrammatism (impoverished syntax in their utterances) favor the present tense in their spontaneous discourse, and both the present and the past tenses in sentence repetition tasks (Centeno, 2007a; Centeno & Obler, 2001). This evidence supports a socio-cognitive approach to interpret verb use in speakers with expressive restrictions or disorders (Centeno, 2007a; Silva-Corvalán, 1991). Verb tenses acquired early by children and used frequently in unimpaired conversation (e.g., past and present) may have certain features—being so common as to be automatic, for instance—that enhance resistance to loss and errors. In addition, these tenses with their simple inflectional endings may be easier to process than more morphologically complex tenses (e.g., conditional: Yo caminaría, “I would walk”; present perfect: Yo he caminado, “I have walked”) (Centeno, 2007a; Centeno & Obler, 2001). To understand linguistic restrictions in speakers with expressive deficits, our analysis may be strengthened by considering the frequency of use in daily conversation and the linguistic complexity of expressive elements favored by speakers. Language Switching and Mixing Sociolinguistic description of bilingual discourse suggests the frequent use of code-switching and code-mixing (Bhatia & Ritchie, 1996). The former involves language switches occurring at sentence boundaries (e.g., I’m hungry pero no quiero comer, I’m hungry but I don’t want to eat); the latter includes language switches taking place within clause or sentence boundaries (e.g., Ella está very happy, She is very happy). Though the distinction between switching and mixing is controversial, both expressive devices constitute a trademark of proficient bilinguals (Bhatia & Ritchie, 1996). Effective control of language switching (LS) and language mixing (LM) is a requirement for bilinguals, especially for pragmatically appropriate language selection in monolingual and bilingual discourse. Brain damage may impair control mechanisms and lead to pathological LS and LM. Different approaches have been proposed to account for normal and pathological language switching. Among them, the lesion approach examines the impact of brain damage on LS, whereas the cognitive approach describes LS in terms of cognitive operations or processing requirements. More recently, the neurocognitive approach emphasizes how processing devices map onto neuroanatomical sites (Green, 1986, 2005). A neurocognitive model of control in bilingual language switching can be useful in analyzing and treating pathological switching in bilingual patients with aphasia. This model suggests that control of the bilingual language system, including language switching, may be affected when brain damage impairs the necessary cognitive operations. Ansaldo and Marcotte (2007) relied on this concept to plan treatment for a Spanish-English individual with aphasia.. The patient used switching as a strategy to overcome anomia (word retrieval problems) but did not have voluntary control over his LM and LS, even with monolingual partners. The patient, however, could translate better than he could name specific items. According to the neurocognitive model, the patient’s involuntary impairment in LM and LS affected the lexical level (i.e., anomia) and the L1-L2 control level (i.e., involuntary mixing and switching). The word-retrieval deficit, combined with possible problems in the mechanisms that control inhibition, resulted in switching from one language to the other. Translation and switching were integrated into a treatment program based on a neurocognitive strategy that would enhance voluntary switching. Translation was a useful compensatory technique for treatment because translational abilities are less impaired than other linguistic skills in bilinguals with aphasia (Paradis, 2004), as the patient exhibited. Switching aimed to increase control by systematically manipulating the impaired use of this behavior. Prior to treatment, the patient was tested in noun and verb naming, repetition, and translation of the same items from Spanish to English, and vice versa. A “Switch Back through Translation” (SBT) approach integrated translation and switching into treatment. SBT transformed pathological LS and LM into translation by cueing the patient to provide the closest equivalent of a noun or a verb in the other language whenever he erred on language selection. The increased awareness and control in switching improved his communication abilities, as the patient gradually learned to translate independently and shifted to the appropriate language. It also facilitated his word-finding abilities. A theory-based neurocognitive intervention that connects cognitive operations and discourse features may be useful in cases of impaired switching in bilinguals with aphasia. This strategy targets impaired neurocognitive factors (i.e., attention and control) while simultaneously relying on translation to address the impaired use of discourse features in bilinguals (i.e., LM and LS) to optimize communication. Ethnography and Dementia The tradition of ethnographic research focuses on describing a culture’s patterns of behavior, norms, and beliefs, among other characteristics, from the perspective of its members (Hymes, 1972a, b; 1974). Ethnographers observe and record patterns of social and communicative behaviors in relation to a specific situation or a specific stimulus. The ethnography of communication (EC) provides a systematic investigation of patterns in language use in interaction. It also provides a descriptive, analytical framework for the communication context and for the participants, their social roles and their impact on the interaction. A central tenet of this approach is that communication is an act that reveals a speech community’s attitudes and beliefs (Guendouzi & Müller, 2006). The clinical benefits of EC can be used in treatment of individuals with communication disorders, as shown in the interactions with patients with dementia. Beliefs about dementia may affect the way in which a society reacts to and cares for—or doesn’t care for—individuals with dementia. These beliefs then give rise to a culture of stereotypes that include negative views of aging. When we interact with people who have dementia, we may bring cultural expectations (e.g., the belief that all people with dementia are aggressive) to the interaction. Such cultural expectations may frame the way people, including clinicians and relatives, approach the person with dementia and influence how they communicate with that person. Indeed, it may be that our beliefs about dementia inadvertently cause us to interact in ways that are less than optimal for treatment. Acknowledging ethnographic and sociolinguistic factors broadens our interpretation of language development/impairment and psycholinguistic processing in young and adult speakers, and our understanding of practitioner/relative-client interactions. As these examples have shown, important relationships exist between ethnographic factors (e.g., language practices in Hispanic individuals) and sociolinguistic dimensions (e.g., discourse patterns); both areas have relevance to linguistic profiles (e.g., language delay, language attrition, and aphasia) and processing domains (i.e., linguistic and cognitive operations). Although the use of ethnography and sociolinguistics is not new in speech-language pathology (e.g., Simmons-Mackie & Damico, 1999; Washington & Craig, 1994; Westby, 1994), an increased application of interdisciplinary and innovative approaches to the study of communication disorders is needed (e.g., Centeno et al., 2007; Code, 2001; Silliman, 2007). Ethnographic and sociolinguistic analysis provides valuable insights into the complex interactions of culture, language, communication, and cognition. Understanding how these factors relate to research in our discipline can strengthen the development of sound experimental methodology, ecologically valid theoretical accounts, and realistic evidence-based practices. Given our increasingly diverse clinical caseloads, such strategies are imperative. This article is based on a presentation by the authors at the 2006 ASHA Convention. Ethnography of Communication: A Person-Centered Approach Ethnography of communication relies on systematic person-centered descriptions of patterns of linguistic form, pragmatic usage, and social function. Consider a visit by a graduate student to a nursing home to collect data for a project on dementia. Hymes (1972a, b) summarized the major ethnographic factors involved in analyzing a speech situation through the use of the anagram SPEAKING. Setting: The resident’s room in the nursing home Participants: A person with dementia and a graduate student in speech-language pathology Ends (goals): These are difficult to ascertain in the case of the person with dementia. However, the graduate student has both overt goals (e.g., to learn about the resident’s life, and to spend time visiting with him/her) and covert goals (e.g., to collect data in order to study and treat dementia) Acts sequence: The types of communication used (e.g., a question-answer format) Key: Whether the interaction is informal or formal Instrumentality: The mode of communication (e.g., conversation or sign language) Norms: Polite conversation Genre: Possibly a friendly chat or “small talk” An ethnographic approach allows the clinician to consider the communicative behaviors the patient or client manifests based upon his or her communication status and the situational and environmental factors that influence the interaction. This approach can be applied to all speakers interacting with people with communication disorders. References Anderson R. T. (2004). First language loss in Spanish-speaking children: Patterns of loss and implications for clinical practice.In Goldstein B. A. (Ed.), Bilingual language development and disorders in Spanish-English speakers (pp. 187–212). Baltimore, MD: Brookes. Google Scholar Anderson R. T. (2001). Loss of gender agreement in L1 attrition: Preliminary results.Bilingual Research Journal, 23, 389–408. CrossrefGoogle Scholar Ansaldo A. I., & Marcotte K. (2007). Language switching in the context of Spanish-English bilingual aphasia.In Centeno J. G., Anderson R. T., & Obler L. K., L. K. (Eds.), Communication disorders in Spanish speakers: Theoretical, research, and clinical aspects. Clevedon, UK: Multilingual Matters. Google Scholar Bhatia T. K., & Ritchie W. C. (1996). Bilingual language mixing, universal grammar, and second language acquisition.In Ritchie W. C. & Bhatia T. K. (Eds), Handbook of second language acquisition (pp. 627–688). San Diego, CA: Academic Press. Google Scholar Castilla A. P., & Restrepo M. A. (2004). L1 predictors of semantic and morphosyntactic development in English as a second language. Unpublished manuscript. Google Scholar Centeno J. G. (2007a). Canonical features in the inflectional morphology of Spanish-speaking individuals with agrammatic speech.Advances in Speech-Language Pathology, 9(2), 1–11. CrossrefGoogle Scholar Centeno J. G. (2007b). Considerations for an ethnopsycholinguistic framework for aphasia intervention.In Ardila A. & Ramos E. (Eds.), Speech and language disorders in bilingual adults. New York: Nova Science. Google Scholar Centeno J. G., & Obler L. K. (2001). Agrammatic verb errors in Spanish speakers and their normal discourse correlates.Journal of Neurolinguistics, 14, 349–363. CrossrefGoogle Scholar Code C. (2001). Multifactorial processes in recovery from aphasia: Developing the foundations for a multilevel framework.Brain and Language, 77, 25–44. CrossrefGoogle Scholar Coulmas F. (1997). Introduction.In Coulmas F. (Ed.), The handbook of sociolinguistics (pp. 1–12). Cambridge, MA: Blackwell. Google Scholar Fiestas C. E., & Peña E. (2004). Narrative discourse in bilingual children: Language and task effects.Language, Speech and Hearing Services in Schools, 35, 155–166. LinkGoogle Scholar Green D.W. (1986). Control, activation, and resource: A framework and a model for the control of speech in bilinguals.Brain and Language, 27, 210–223. CrossrefGoogle Scholar Green D. W. (2005) The neurocognition of recovery patterns in bilingual aphasics.In Kroll J. F. & de Groot A. M. B. (Eds), Handbook of bilingualism: Psycholinguistic approaches (pp. 516–530). New York: Oxford University Press. Google Scholar Gutiérrez-Clellen V. F., & Hofstetter R. (1994). Syntactic complexity in Spanish narratives: A developmental study.Journal of Speech and Hearing Research, 37, 645–654. LinkGoogle Scholar Hymes D. (1972a) Models of the interaction of language and social life.In Gumperz J. & Hymes D. (Eds.), Directions in sociolinguistics (pp. 35–71). New York: Holt, Rinehart and Winston. Google Scholar Hymes D. (1972b). Toward ethnographies of communication: The analysis of communicative events.In Giglioli P. P. (Ed.), Language and social context (pp. 21–44). Harmondsworth: Penguin Books. Google Scholar Hymes D. (1974). Foundations in An ethnographic University of Press. Google Scholar Jacobson in Spanish-English bilingual children: production and tasks. at the for Research in Child Language Google Scholar & L. Patterns of discourse. Google Scholar E., & B. B. The impact of language socialization on grammatical P. & B. (Eds.), The handbook of child language (pp. Cambridge, MA: Blackwell. Google Scholar M. (2004). A theory of CrossrefGoogle Scholar Restrepo M. A. of Spanish-speaking children with language of Language, and Hearing Research, LinkGoogle Scholar Restrepo M. A. for A program for bilingual children. at the Google Scholar Restrepo M. C. & Castilla A. P. of language elicitation techniques with Spanish-speaking children. in Google Scholar Roseberry-McKibbin & L. English language in school A and Hearing Services in Schools, LinkGoogle Scholar Language, culture, and An to linguistic Press. Google Scholar E., & D. (1994). of linguistic R. A. & D. (Eds.), in A developmental study (pp. Google Scholar E. R. research ASHA LinkGoogle Scholar C. and in oral Spanish CrossrefGoogle Scholar C. Spanish language attrition in a situation with W. and R. M. (Eds.), First language attrition (pp. Cambridge, UK: University Press. CrossrefGoogle Scholar Washington J. & (1994). forms during discourse of of Speech and Hearing Research, 37, LinkGoogle Scholar C. E. (1994). The of culture on and of oral and G. P. & K. G. (Eds.), Language in school-age children and principles and and Google Scholar A. C. (2005) on language and literacy in Latino families and A. C. (Ed.), on Language and literacy in Latino families and (pp. 1–12). New York: Press. Google Scholar José G. Centeno, is an in the Speech-Language and in the of Communication and at at Raquel T. Anderson, is an in the of Speech and Hearing at her at M. Adelaida Restrepo, is an in the Speech and Hearing at her at Peggy F. Jacobson, is an in the Speech-Language and in the of Communication and at research is by an her at Jackie Guendouzi, is an in the Speech-Language at the University of her at Nicole Müller, is an and in at the University of at her at Ana Inés is an at the et de de her at Karine is a speech-language and at the program of the de her at to in Jul &
Vernacularisation et traduction des textes pragmatiques en Afrique — La traduction des textes comportant des lacunes d'ordre grammatical, lexical, stylistique ou idiomatique présente habituellement des difficultés particulières, lesquelles sont amplifiées lorsqu'elles sont attribuables à la vernacularisation d'une langue étrangère. Dans les sociétés postcoloniales, l'absence ou la non-disponibilité des études linguistiques sur la plupart des langues locales rend ardue l'analyse des interférences entre ces dernières et les langues officielles étrangères. Cette situation, ajoutée à la grande diversité ethnolinguistique ambiante, ne facilite pas l'interprétation des textes produits par les personnes semi-lettrées. Le traducteur de ces textes se présente davantage comme un rédacteur qui, à partir de l'idée globale qui se dégage de l'original, conçoit et produit un texte répondant aux normes de la langue cible. L'évaluation d'un tel travail ne peut se faire qu'en comparant la finalité des deux textes.
In this paper we will describe VIT (Venice Italian Treebank), created at the University of Venice. We will focus on the syntactic-semantic features and on the quantitative analysis of the data of our treebank comparing them to other treebanks. In general, we will try to substantiate the claim that treebanking grammars or parsers is dramatically dependent on the chosen treebank; and eventually this process seems to be dependent either from substantial factors such as the adopted linguistic framework for structural description or, ultimately, the described language.
This article is about analysis of data obtained in repeated measures designs in psycholinguistics and related disciplines with items (words) nested within treatment (5 type of words). Statistics tested in a series of computer simulations are:F 1,F 2,F 1 &F 2,F′, minF′, plus two decision procedures, the one suggested by Forster and Dickinson (1976) and one suggested by the authors of this article. The most common test statistic,F 1 &F 2, turns out to be wrong, but all alternative statistics suggested in the literature have problems too. The two decision procedures perform much better, especially the new one, because it systematically takes into account the subject by treatment interaction and the degree of word variability.
Psychocomputational Models of Human Language Acquisition (PsychoCompLA-2007) William Gregory Sakas (sakas@hunter.cuny.edu) Department of Computer Science, Hunter College Ph.D. Programs in Linguistics and Computer Science, The Graduate Center City University of New York 695 Park Ave, New York, NY 10021 USA David Guy Brizan (dbrizan@gc.cuny.edu) Ph.D. Program in Computer Science, The Graduate Center City University of New York 365 Fifth Avenue, New York, NY 10016 USA Keywords: language acquisition; syntax acquisition; language learning; language change; computational; linguistics; psycholinguistics; psychology; statistical; innateness. important question. One effective line of investigation is to computationally model the acquisition process and determine interrelationships between a model and linguistic or psycholinguistic theory, and/or correlations between a model's performance and data from linguistic environments that children are exposed to. Workshop Topic and History The workshop is devoted to psychocomputational models of language acquisition. By psychocomputational, we mean computational models that are compatible with research in psycholinguistics, developmental psychology and/or linguistics. Although there has been a significant amount of presented research targeted at modeling the acquisition of word categories, morphology and phonology, research aimed at modeling syntax acquisition has just begun to emerge. This is the third meeting of the Psychocomputational Models of Human Language Acquisition workshop following PsychoCompLA-2004, held in Geneva, Switzerland as part of the 20th International Conference on Computational Linguistics (COLING 2004) and PsychoCompLA-2005 as part of the 43rd Annual Meeting of the Association for Computational Linguistics (ACL-2005) held in Ann Arbor, Michigan where the workshop shared a joint session with the Ninth Conference on Computational Natural Language Learning (CoNLL-2005). Invited Presentations Statistical language learning: Computational and maturational constraints Elissa Newport, University of Rochester, USA The next challenges in unsupervised language acquisition: Dependencies and complex sentences Shimon Edelman, Cornell University, USA Learnable representations of languages: Something old and something new Alex Clark, Royal Holloway University of London, UK Workshop Description The workshop will present research and foster discussion centered around psychologically-motivated computational models of language acquisition, with an emphasis on the acquisition of syntax. In recent decades there has been a thriving research agenda that applies computational learning techniques to emerging natural language technologies and many meetings, conferences and workshops in which to present such research. However, there have been only a few (but growing number of) venues in which psychocomputational models of how humans acquire their native language(s) are the primary focus. Indirect evidence and the poverty of the stimulus Terry Regier, University of Chicago, USA Lexical learning and lexical diffusion Charles D. Yang, University of Pennsylvania, USA Bootstrapping bootstrapping Damir Cavar, Zadar University, Croatia and University of Indiana, USA Transformational networks Bob Frank, John Hopkins University, USA Psychocomputational models of language acquisition are of particular interest in light of recent results in developmental psychology that suggest that very young infants are adept at detecting statistical patterns in an audible input stream. Though, how children might plausibly apply statistical 'machinery' to the task of grammar acquisition, with or without an innate language component, remains an open and The great (Penn Treebank) robbery: When statistics is not enough Sandiway Fong, University of Arizona, USA, and Robert C. Berwick, MIT, USA (joint work with Partha Niyogi, University of Chicago, USA)
We report work in progress on a complex system generating Czech sentences expressing the meaning of input syntactic-semantic structures. Such component is usually referred to as a realizer in the domain of Natural Language Generation. Existing realizers usually take advantage of a background linguistic theory. We introduce the Functional Generative Description, a framework of our choice conceived in 1960's by Petr Sgall. This language theory lays out foundations of the formalism in which our input syntactic-semantic structures are specified. The structure definition was further elaborated and refined during the annotation of the Prague Dependency Treebank, now available in its second version. A section of the paper is devoted to description of another theoretical framework suitable for the task of Natural Language Generation - the Meaning-Text Theory. We explore state-of-the-art realizers deployed in real life applications, describe common architecture of a generation system and highlight the strengths and weaknesses of our approach. Finally, preliminary output of our surface realizer is compared against a baseline solution.
We present an unsupervised linguistically-based approach to discourse relations recognition,\nwhich uses publicly available resources like manually annotated corpora (Discourse Graph\nBank, Penn Discourse TreeBank, RST-DT), as well as empirically derived data from “causally”\nannotated lexica like LCS, to produce a rule-based algorithm. In our approach we use\nthe subdivision of Discourse Relations into four subsets – CONTRAST, CAUSE, CONDITION,\nELABORATION, proposed by [1] in their paper where they report results obtained with a\nmachine-learning approach from a similar experiment against which we compare our results.\nOur approach is fully symbolic and is partially derived from the system called GETARUNS,\nfor text understanding, adapted to a specific task: recognition of Discourse Causal Relations\nin free text. We show that in order to achieve better accuracy both in the general task and in\nthe specific one, semantic information needs to be used besides syntactic structural information.\nOur approach outperforms results reported in previous papers
The current research explores the effects of exemplars on the stereotype representation of one's ingroup. Previous research demonstrated that exposure to an ingroup exemplar affects the stereotype one holds of one's ingroup (Coats & Smith, 1999). The primary purpose of the present study was to examine whether this effect is moderated by relative ingroup size. Participants were placed into either a minority or majority group situation and exposed to 1 of 2 dissimilar exemplars of their ingroup. Later, they rated their ingroup. Ratings of the ingroup differed between exemplar conditions in unexpected ways, indicating that the exemplar affected participants' stereotype of their ingroup. Furthermore, exemplars had a stronger effect on participants in the minority group than those in the majority group. Finally, relative ingroup size and, to a marginal extent, exemplars were found to affect ratings of ingroup variability.
Typically, personalized information recommendation services automatically infer a user profile, a structured model of the user interests, from documents the user already deemed as relevant. Traditional keyword-based approaches are unable to capture the semantics of the user interests. This work proposes a strategy consisting of two steps. The first one is a semantic indexing procedure based on a word sense disambiguation strategy which exploits the WordNet lexical database to select, among all the possible meanings (senses) of a polysemous word, the correct one. In the second step, semantically indexed documents are mined by a naive Bayes learning algorithm that infer semantic, sense-based user profiles. Two experimental sessions were carried out to compare the performance of keyword-based profiles to that of sense-based profiles. We measured both the classification accuracy and the effectiveness of the ranking imposed by the two different kinds of profile on the documents to be recommended. The main outcome of both experiments is that the classification accuracy is improved without improving the ranking. Personalized systems adapt their behavior to individual users by learning their preferences during the interaction in order to construct a user profile that can be later exploited in the search process. Traditional keyword-based approaches are primarily driven by a string-matching operation: If a string, or some morphological variant, is found in both the profile and the document, a match is made and the document is considered relevant. String matching suffers from problems of polysemy, the presence of multiple meanings for one word, and synonymy, multiple words having the same meaning. The result is that, due to synonymy, relevant information can be missed if the profile does not contain the exact
Models of social evaluation aim to capture the information people use to form first impressions of unfamiliar others. However, little is currently known about the relationship between perceived traits across gender. In Study 1, we asked viewers to provide ratings of key social dimensions (dominance, trustworthiness, etc.) for multiple images of 40 unfamiliar identities. We observed clear sex differences in the perception of dominance-with negative evaluations of high dominance in unfamiliar females but not males. In Study 2, we used the social evaluation context to investigate the key predictions about the importance of pictorial information in familiar and unfamiliar face processing. We compared the consistency of ratings attributed to different images of the same identities and demonstrated that ratings of images depicting the same familiar identity are more tightly clustered than those of unfamiliar identities. Such results imply a shift from image rating to person rating with increased familiarity, a finding which generalises results previously observed in studies of identification.
Semantic differential techniques are a useful, well-validated tool to assess affective processing of stimuli and determine how that processing is impacted by various demographic factors, such as gender. In this paper, we explore differences in connotative word processing between men and women as measured by Osgood's semantic differential and what those differences imply about affective processing in the two genders. We recruited 94 young participants (47 men, 47 women, ages 18-39) using an online survey and collected their affective ratings of 120 words on three rating tasks: Evaluation (E), Potency (P), and Activity (A). With these data, we explored the theoretical and mathematical overlap between Osgood's affective meaning factor structure and other models of emotional processing commonly used in gender analyses. We then used Osgood's three-dimensional structure to assess gender-related differences in three affective classes of words (words with connotation that is Positive, Neutral, or Negative for each task) and found that there was no significant difference between the genders when rating Positive words and Neutral words on each of the three rating tasks. However, young women consistently rated Negative words more negatively than young men did on all three of the independent dimensions. This confirms the importance of taking gender effects into account when measuring emotional processing. Our results further indicate there may be differences between Osgood's structure and other models of affective processing that should be further explored.
Objective To validate the translated Chinese version of Pain Assessment in Advanced Dementia Scale (C-PAINAD) for its clinical application in assessing the discomfort level of severely cognitive impaired patients. Methods In developing the C-PAINAD, both clinical and academic experts were engaged in an iterative translation process for ensuring its semantic equivalence with its original language as well as it comprehensibility for clinical application. In establishing C-PAINAD's inter-rater reliability, its applicability for clinical use is further assessed in 11 severely cognitive impaired patients. The assessment tools included C-PAINAD, Discomfort Visual Analog Scale (DVA), and Philadelphia Geriatric Center Affect Rating Scale (PGCAR). Correlations, ANOVA and factor analysis were undertaken to examine the reliability and validity of C-PAINAD. Results The scores of C-PAINAD were not in normal distribution but clustered around Zero. C-PAINAD was positively correlated with DVA and negative affect but mildly and negatively correlated with positive affect. It was able to detect different pain level under different condition. One factor was extracted and the percentage of variance was 51.20%. C-PAINAD was proved to have satisfactory reliability and validity(Cronbach's α=0.66). Conclusions C-PAINAD was a simple, reliable and effective pain assessment instrument for measuring pain level of non-communicative patients like advanced dementia. Further research is deemed necessary to conduct with pre - and post-test of analgesic prescriptions prior to wider clinical application.
Reviewed by: Indian and British English: A handbook of usage and pronunciationby Paroo Nihalni, R. K. Tongue, Priya Hosali, and Jonathan Crowther Niladri Sekhar Dash Indian and British English: A handbook of usage and pronunciation. 2ndedn. By Paroo Nihalni, R. K. Tongue, Priya Hosali, and Jonathan Crowther. New Delhi: Oxford University Press, 2004. Pp. x, 260. ISBN 0195666569. $15.95. The present handbook is divided into two main parts. The first part (‘Lexicon of usage’) is designed to provide English users with information about the way in which certain words, idioms, collocations, phrases, and similar expressions of English used in India differ from British Standard English (BSE)—a model that has the closest affinity to Indian English. This part includes a thousand English words, which are used in a distinctive manner by large numbers of educated Indian speakers of English irrespective of their place, profession, education, gender, or other sociolinguistic factors. The words included in the handbook are selected from the speech or writing samples of the persons (such as university and school teachers, journalists, and radio commentators) who are likely to influence the English use of Indian learners. The handbook also contains many European words that have been Indianized over the years. Thus, it serves as a handy resource for Indian speakers of English, illustrating the many, often quite subtle, ways in which Indian English differs from standard British English usage, and where these differences are regarded as acceptable or substandard in the subcontinent. Examples in the handbook, which supplement the texts, are helpful to Indian users of English who are uncertain about the ‘correctness’ of their speech and writing, and serve those scholars who want to explore the differences between Indian and British uses of English. The book also has the potential to address special problems faced by learners of English, who are often impeded by the difficulties of recognizing finer nuances of meaning and usage. The second part of the handbook includes a brief report on the development of the pronunciation dictionary in India and abroad, followed by insightful discussions on standards of pronunciation in second/ foreign language teaching, the phonological systems of the British Received Pronunciation (BRP) and Educated Indian English (EIE), and the role of supraseg-mental properties (i.e. word stress, sentence stress, rhythm, intonation, etc.) in Indian English. The introduction contains a list of keywords for phonetic symbols used in the following part, ‘Dictionary of pronunciation’. Two types of pronunciation (Indian Recommended Pronunciation and the BRP) are supplied for more than two thousand words collected from the original lexical database of Michael West’s General service list of English wordstogether with a few additions compiled from the language resources available to the compilers. Each entry of the dictionary is tagged with relevant phonological information. This second edition also includes additional information on lexical collocation (the tendency of words to be used together in fixed phrases). In essence, the handbook not only serves as an invaluable reference guide for students and teachers of English, but also makes a valuable contribution for applied linguists, lexicographers, journalists, and scholars who write in Indian English. [End Page 465] Niladri Sekhar Dash Indian Statistical Institute, Kolkata Copyright © 2007 Linguistic Society of America
As the Introduction to this volume observes, sixteenth-century France is marked by ‘une vaste réflexion sur le bien dire’. This not only impacted upon theory and practice across the different literary genres using French but also promoted considerable debate on the form and basis of the emerging standard form of the vernacular. Given the wide-ranging nature of this réflexion, covering its different manifestations in one volume poses an almost insuperable challenge, but the twenty-eight contributions to the colloquium collected here certainly address an ambitiously broad span of topics and add usefully to our understanding of cultural developments in a period of major change. The papers are organized into three general sub-sections, ‘Interroger la norme’, ‘Évolutions de la norme’ and ‘Normes et société’. However, such is the fluid nature of the subject matter treated in certain papers that their classification under one or other of these headings can sometimes seem of doubtful appropriateness. The focus of the first sub-section is predominantly literary. The various contributions address the creation or adaptation of norms across a considerable number of different genres some of which are perhaps rather less familiar, for instance, oracular writings (Dubois), accounts of pilgrimages (Gomez-Géraud) and Jesuit letter-writing (Laborie). Particularly interesting is the close study by Duché of the influential approach which Nicolas Herberay adopted for translation. Herberay, an acknowledged master of French prose (‘un vray Cicero françois’, according to Jean Martin), wrote with a female as well as a male readership in mind, developing a prose style that was eloquent and natural that would set an example for bien dire in this area. The second sub-section begins with a cogent overview (Baddeley) of a familiar field, developments in orthography and the interplay between orthography and spelling, and is followed by a series of studies which explore revealingly topics such as the evolving relationship between poetics and grammar (Monferran), the increasing limitation on the use of metaphor in literary works (Cernogora) and developments in historiography (Dumontet). Perhaps the most interesting paper is the examination of the fortunes of the alexandrine in the early part of the century (Halévy). Particular attention is given to the writings of Jean Lemaire de Belges and Geoffroy Tory both of whom, on the basis of fanciful argumentation, sought to invest the alexandrine with special prestige and nationalistic symbolism matching the terza rima in Italian. Their exercises in myth-making were to contribute indirectly, it is argued, to the rapid rise in the alexandrine's use from around 1555. The final sub-section of the volume contains contributions that more particularly address linguistic issues. Notable amongst these are two items: a re-evaluation of the system of vers mesurés devised by Baïf which is seen as an attempt not only to reproduce the metrical patterns of ancient Greek but also to contribute towards the norms of spoken French by reflecting the élite ‘usage des Bons’ (Vignes); and a meticulous examination by Morin of change in the pronunciation norms presented by Peletier du Mans in his earlier works (1550, 1555) as against his 1581 Euvres poëtiques, the new norm correlating with that presented later in the works of Lanoue (1596) and La Touche (1696). Alongside these are a number of other attractive essays including a study of the linguistic norms in the speeches made at the formal opening of the Paris Parlement, with eloquence and high rhetoric dominating over practicality and clarity between 1560 and 1600 before a reversal occurred in the early seventeenth century (Petey-Girard), and an investigation of sixteenth-century liminaires (any text preceding a written work) composed by women where a complex set of norms operate involving humility, simplicity of style, the practice of dedicating the work to another woman and, in the light of the lack of image for the female writer, an attempt to ‘socialiser l'auteur’ (Gauthier). Completing the text is an Index Nominum and a table of contents. The diversity and scholarly depth of the volume should ensure that all seiziémistes will derive benefit from a close reading.
Abstract In this chapter, we discuss the development and use of picture stimuli incorporated in the International Affective Picture System (IAPS, pronounced “eye-aps”; Lang, Bradley, & Cuthbert, 2005), a large set of emotionally evocative color photographs that includes pleasure, arousal, and dominance ratings made by men and women. The IAPS is currently used in experimental investigations of emotion and attention worldwide, providing experimental control in the selection of emotional stimuli, facilitating the comparison of results across different studies, and encouraging replication within and across psychological and neuroscience research laboratories. Numerous studies in our laboratory over the past 15 years have explored subjective, psychophysiological, behavioral, and neurophysiological reactions when viewing these affective stimuli. Basic findings from these studies, which will be informative for researchers considering or using the IAPS stimuli, are briefly summarized in this chapter.
Parsing unrestricted text is useful for many language technology applications but requires parsing methods that are both robust and efficient. MaltParser is a language-independent system for data-driven dependency parsing that can be used to induce a parser for a new language from a treebank sample in a simple yet flexible manner. Experimental evaluation confirms that MaltParser can achieve robust, efficient and accurate parsing for a wide range of languages without language-specific enhancements and with rather limited amounts of training data.
Databases of hierarchically annotated text occupy a central place in linguistic research and language technology development. We describe a new approach to tree query which we call "Query by Annotation". Users express a query by annotating a tree, and the annotation is compiled into an expression in a path language. The result trees are overlaid with the original query, permitting the user to see why they match. Since queries and results are annotated trees, users can easily refine and resubmit their queries. The approach to Query by Annotation is motivated and exemplified using databases of linguistic trees, or treebanks.
Proceedings of the Sixth International Workshop on Treebanks and \nLinguistic Theories. \nEditors: Koenraad De Smedt, Jan Hajič and Sandra Kübler. \nNEALT Proceedings Series, Vol. 1 (2007), 127-138. \n© 2007 The editors and contributors. \nPublished by \nNorthern European Association for Language \nTechnology (NEALT) \nhttp://omilia.uio.no/nealt. \nElectronically published at \nTartu University Library (Estonia) \nhttp://hdl.handle.net/10062/4476.
Function labels enrich constituency parse tree nodes with information about their abstract syntactic and semantic roles. A common way to obtain function-labeled trees is to use a two-stage architecture where first a statistical parser produces the constituent structure and then a second \ncomponent such as a classifier adds the missing function tags. In order to achieve optimal results, training \nexamples for machine-learning-based classifiers should be as similar as possible to the instances seen during prediction. However, the method which has been used so far to obtain training examples for the function labeling classifier suffers from a serious drawback: the training examples come from perfect treebank trees, whereas test \nexamples are derived from parser-produced, imperfect trees. \nWe show that extracting training instances from the reparsed training part of the treebank results in better training material as measured by similarity to test instances. We show that our training method achieves statistically significantly higher f-scores on the function labeling task for the English Penn Treebank. Currently our method achieves 91.47% f-score on the section 23 of WSJ, the highest score reported in the literature so far.
We present several improvements to unlexicalized parsing with hierarchically state-split PCFGs. First, we present a novel coarse-to-fine method in which a grammar’s own hierarchical projections are used for incremental pruning, including a method for efficiently computing projections of a grammar without a treebank. In our experiments, hierarchical pruning greatly accelerates parsing with no loss in empirical accuracy. Second, we compare various inference procedures for state-split PCFGs from the standpoint of risk minimization, paying particular attention to their practical tradeoffs. Finally, we present multilingual experiments which show that parsing with hierarchical state-splitting is fast and accurate in multiple languages and domains, even without any language-specific tuning. 1