Papers reviewed and determined not to be word norm studies. Use the flag icon to report errors or suggest re-inclusion.
18265 papers
The aim of this study is to show how clusteranalysis can shed light on very complexvariation in a transitional dialect zone ineastern Finland. In the course of history thisarea has been on the border between Sweden andRussia and the population has clearly been oftwo kinds: the Savo people and the Karelians.It is a well-known fact that there is variationamong these dialects, but the spread and extentof the variation has not been demonstrated previously.The idiolects of the area were studied in thelight of ten phonological and morphologicalfeatures. The material consisted of recordingsof 198 idiolects, totalling around 195 hoursand representing 19 parishes. The variation wasanalysed using hierarchical cluster analysis.While the analysis showed the extent of thevariation between idiolects and parishes, italso demonstrated how the effects of the oldparishes, borders and settlements are stillvisible in the dialects. On the parish level,the data formed clear clusters that correspondwith the main dialects in the area and itssurroundings. On the idiolect level, however,the speakers from the surrounding areas formedfairly homogenous clusters but the idiolectsfrom the Savonlinna area were spread acrossalmost all clusters.
Cited several times. E.g. 1. Marco Kuhlmann & Joakim Nivre: Mildly non-projective dependency structures. In the Proceedings of the COLING/ACL on Main conference poster sessions, p. 507--514. In series COLING-ACL '06. Sydney, Australia, 2006. 2. Carlos Gómez-Rodriguez and Joakim Nivre: A transition-based for 2-Planar Dependency Structures. In Proceedings of the 48th Annual Meeting of the Association for Computational Linguistics, pages 1492--1501, Uppsala, Sweden, 11-16 July 2010. ACL 3. Marco Kuhlmann. Dependency Structures and Lexicalized Grammars. An Algebraic Approach. LNAI 6270. FoLLI Publications on Logic, Language and Information. Springer 2010. 4.
The linguistic annotation of natural language corpora is one of the main areas of computational linguistics. Much energy has been devoted to building large syntactically annotated corpora, which are also called treebanks after the phrase-structure trees they contain. For some time now, functional information, i.e. information on whether a constituent functions as e. g. subject, object or adverbial, has also been included in the annotation. Yet, while at first glance this may not seem to be a venture too complicated, matters are not always as easy as they seem.
This paper discusses an automatic, data-driven approach to treebank error detection. The approach adapts the use of so-called variation n-grams as defined in Dickinson and Meurers (2003) for the detection of inconsistent part-of-speech annotations to syntactic annotation. The underlying idea is to define a consistency test for the mapping from recurring strings to their syntactic annotation. The paper illustrates with a case study based on the WSJ treebank that the method successfully detects inconsistencies in syntactic category annotation. Since such inconsistencies are typically introduced by humans, our method works best for large corpora that have been annotated manually or semi-automatically, which is generally the case for current syntactic and other high-level annotation. \n \nOur work serves two main purposes for treebank improvement. It is a means for finding erroneous variation in a corpus, which can then be corrected. And it provides feedback for the development of empirically adequate standards for syntactic annotation, showing which distinctions are difficult to maintain over an entire corpus. Additionally, as a method for comparing syntactic annotation, our work could have uses for interannotator agreement testing and parser evaluation.
The Syntagmatic Paradigmatic model (SP; Dennis & Harrington 2001, Dennis submitted) and the Pooled Adjacent Context model (PAC; Redington, Chater & Finch 1998) are compared on their ability to extract syntactic, semantic and associative information from a corpus of text.On a measure of syntactic class (and subclass) information based on the WordNet lexical database (Miller 1990), the models performed similarly with a small advantage for the PAC model.On a measure of semantic structure based on the similarities produced by Latent Semantic Analysis (LSA; Landauer & Dumais 1997), the models performed equivalently with a small advantage for the SP model.On a measure of associative information based on the free association norms of Nelson, McEvoy & Schreiber (1999), the SP model shows a substantive advantage over the PAC model producing more than twice as many associates.
This paper deals with up-translation - a process of lexical data transformation from any source format to the XML document. Relevant aspects of the XML format and many related technologies are surveyed first. Then, information content enhancement of existing lexical resources is discussed. The last part brings information about up-translation ofthe Dictionary of Literary Czech Language and the way of efficient storage and retrieval of data.
Treebanks are widely recognised as a necessary source of information in NLP as well as in Linguistics studies. In this paper we present and justify methodological principles and syntactic criteria to build a Treebank for Spanish: annotating only explicit information, constituents and syntactic functions and being theory independent. Previous work is also presented in order to account for taken decisions. The annotation process will be done in different steps so that each one of them is the input of the next. We present the basic guidelines of syntactic annotation and the boundaries of the work to be done in a first step: annotation of low constituents and surface functions. Moreover, some semantic information (subject type) is likely to be included.
One of the major purposes of annotated corpora is their potential for use as databases for linguistic research. An important design criterion for corpora specifically intended for this use is the need to encode a plurality of types of information, some of which are clearly interrelated. This need can conflict
Oxymonads are a morphologically well-characterized and highly diverse lineage of protists. They are, however, under sampled at a molecular level. It has recently been demonstrated that a genus of oxymonads, Pyrsonympha, is phylogenetically related to the excavate taxon Trimastix. Here, we addressed issues of internal oxymonad evolution. Pyrsonympha and Dinenympha are shown, by fluorescent in situ hybridization and phylogenetic evidence, to be separate genera and not morphotypes of the same organism. We demonstrated that three genera of oxymonads, Dinenympha, Pyrsonympha, and Oxymonas are each monophyletic and together form a clade which excludes other known eukaryotes. We have presented a taxonomic scheme of oxymonads taking into account their sisterhood with Trimastix and speculated on morphological evolution of oxymonads, particularly of their attachment apparatuses. Our biogeographical analysis with Japanese and Canadian Pyrsonympha and Dinenympha suggests that these genera diverged before the separation of termites that inhabit Eastern Asia and Western North America.
^j;=5!? ITHIN the rich corpus of metrical psalms comtm itt^ta * posed during Spain's Golden Age, Fray Luis de Aivi bl T 0_Le6n's versions are universally accorded the high;@ ^ VV |@ est praise. As heir to a literary tradition that ex,.A s iGne * tended back to the late Middle Ages and early.f4Li J Renaissance, the Salamancan scholar and poet revolutionized Spain's engagement with the Psalter, establishing the lira or estrofa alirada as the dominant verse form for vernacular psalm translations (Rivers 112; Nufiez 357), and making close lexical parallelism and philological accuracy, rather than interpretive digression, the norm for most of his followers. As may be expected from the great Augustinian's role as el primer poeta humanista espaniol en lengua vulgar (A. Blecua 97), Fray Luis was widely imitated, especially among disciples of his own order, and questions of authorship and dating of the many psalm versions attributed to him continue to trouble literary historians (Nufiez 357-8; J. M. Blecua Poesia completa, 41-2). Jose Manuel Blecua, in his 1990 edition of the Poesia completa, based on all extant manuscripts, includes as genuine the following poems: Psalm 1 Beatus vir, 4 Cum invocarem, 6 ne in furore, 9 Salvum me fac, 12 Usquequo, Domine (2 versions), 17 Diligam te, 18 Coeli enarrant, 24 Ad te, Domine, levavi,
We present a new approach to topological parsing of German which is corpus-based and built on a simple model of probabilistic CFG parsing. The topological field model of German provides a linguistically motivated, flat macro structure for complex sentences. Besides the practical aspect of developing a robust and accurate topological parser for hybrid shallow and deep NLP, we investigate to what extent topological structures can be handled by context-free probabilistic models. We discuss experiments with systematic variants of a topological treebank grammar, which yield competitive results.
This study investigated the human eyeblink startle reflex as a measure of alcohol cue reactivity. Alcohol-dependent participants early (n = 36) and late (n = 34) in abstinence received presentations of alcohol and water cues. Consistent with previous research, greater salivation and higher ratings of urge to drink occurred in response to the alcohol cues. Differential salivary and urge responding to alcohol versus water cues did not vary as a function of abstinence duration. Of special interest was the finding that startle response magnitudes were relatively elevated to alcohol cues, but only in individuals early in abstinence. Affective ratings of alcohol cues suggested that alcohol cues were perceived as aversive. Methodological and theoretical implications of the findings are discussed.
This paper describes a lexicalized tree adjoining grammar (LTAG) based parsing system for Korean which combines corpus-based morphological analysis and tagging with a statistical parser. Part of the challenge of statistical parsing for Korean comes from the fact that Korean has free word order and a complex morphological system. The parser uses an LTAG grammar which is automatically extracted using LexTract (Xia et al., 2000) from the Penn Korean TreeBank (Han et al., 2002). The morphological tagger/analyzer is also trained on the TreeBank. The tagger/analyzer obtained the correctly disambiguated morphological analysis of words with 95.78/95.39% precision/recall when tested on a test set of 3,717 previously unseen words. The parser obtained an accuracy of 75.7% when tested on the same test set (of 425 sentences). These performance results are better than an existing off-the-shelf Korean morphological analyzer and parser run on the same data
Many recent statistical parsers rely on a preprocessing step which uses hand-written, corpus-specific rules to augment the training data with extra information. For example, head-finding rules are used to augment node labels with lexical heads. In this paper, we provide machinery to reduce the amount of human effort needed to adapt existing models to new corpora: first, we propose a flexible notation for specifying these rules that would allow them to be shared by different models; second, we report on an experiment to see whether we can use Expectation-Maximization to automatically fine-tune a set of hand-written rules to a particular corpus.
We present an algorithm which translates the Penn Treebank into a corpus of Combinatory Categorial Grammar (CCG) derivations. To do this we have needed to make several systematic changes to the Treebank which have to effect of cleaning up a number of errors and inconsistencies. This process has yielded a cleaner treebank that can potentially be used in any framework. We also show how unary type-changing rules for certain types of modifiers can be introduced in a CCG grammar to ensure a compact lexicon without augmenting the generative power of the system. We demonstrate how the combination of preprocessing and type-changing rules minimizes the lexical coverage problem. 1.
You have accessThe ASHA LeaderFeature1 Nov 2002AAC, Literacy and Bilingualism Ovetta L. Harrison-Harris Ovetta L. Harrison-Harris Google Scholar https://doi.org/10.1044/leader.FTR2.07202002.4 SectionsAbout ToolsAdd to favorites ShareFacebookTwitterLinked In Children who use augmentative and alternative communication (AAC) have historically been challenged in their attainment of literacy skills. These challenges are even greater for AAC users who are bilingual. AAC users in the United States comprise large numbers of individuals from culturally and linguistically diverse backgrounds. Current demographic trends indicate that linguistic diversity will continue to intensify. During the 12 years between 1986 and 1998, the number of U. S. children who were identified as limited English proficient increased from 1.6 million to 9.9 million (see Tucker 1999). It is estimated that, by the year 2050, 40% of school-aged children in the United States will come from homes where English is not the first language. Individuals who use AAC systems surely will be represented in this group. The fact that many children in the United States, including those who use AAC systems, live amidst a sea of languages has captured national attention and has influenced our educational system. The new thrust to achieve educational equality represents a historic change. Many bilingual or monolingual schools that taught in languages other than English existed before World War II. For example, many German-only schools could be found in the northern Midwest. Afterward, a pattern of English-only instruction dominated our education system. As recognition of the cultural and linguistic diversity of the United States grew, a need to provide effective and appropriate education for bilingual children arose. Educators, parents, and researchers have challenged the notion of an English-only education for children from linguistically diverse backgrounds. Research supports the notion of education for limited-English-proficient children, including those relying on AAC systems, to be introduced in their first language, providing a transition to stronger second-language usage. This is logical given the fact that literacy attainment depends on language. Language learning, including reading and writing, is always culturally based. Reading and writing involve particular ways of using and thinking about written language that go beyond finding meaning in text and include the construction of sociocultural viewpoints or ways of understanding the world around us. It is important to realize the sociocultural and communicative nature of literacy, because of the possible therapeutic impact when working with bilingual AAC users. Writing, similarly, is a contextualized social event. It is a transactional, circular process created from a person's linguistic resources and interaction with past experiences. Viewing literacy learning as it is socially constructed through language provides a nice perspective of the need to educate linguistically and culturally diverse children from their first-language knowledge base. Research Challenges The challenges of literacy attainment for both monolingual and bilingual children who use AAC have become an area of focus for special educators, speech-language pathologists, parents, and researchers. Although research in reading and writing development of AAC users has increased steadily over the past 10 years, only recently have researchers turned their attention to reading and writing development of bilingual AAC users. Many of these children are unsuccessful in developing literacy, yet there is increasing recognition that this group is capable of developing sophisticated reading and writing skills. Bilingual AAC users who are highly successful in developing these skills make tremendous gains in overall language development and in use of their AAC systems. Acquisition of more vocabulary and the ability to compose text are just two advantages that literacy attainment brings to their receptive and expressive language development. Major focus has been brought to the topic of literacy attainment for bilingual AAC users because of its particular importance for this population. Attainment of literacy allows bilingual and monolingual AAC users, like all students, to be able to prepare messages to be used at a later time, produce exact messages, and learn vocabulary with which they can spontaneously spell out messages. But Light and McNaughton (l993) give three reasons why literacy development holds additional importance for AAC communicators. First, their face-to-face communication skills are often severely limited. Communication can be quite slow. Often the able-bodied message receiver doesn't have time to participate in communication interaction with an AAC communicator. Research shows that, in interactions between a person who is using an AAC system and a speaking person, the speaking person often dominates the interaction, and the person using the AAC system may not have opportunities to initiate topics or converse fully. Literacy gives an AAC communicator the opportunity to overcome many of the restrictions of face-to-face interaction, especially those imposed by slow AAC systems. Through writing it is possible for individuals to communicate more fully, to express themselves in more detail, and to circumvent some of the time limitations that they would normally experience in face-to-face interactions. The second aspect of school literacy importance for individuals who use AAC systems is that those who are preliterate are often limited to an ideographic literacy system. Some of these graphic systems force AAC communicators to use a closed vocabulary set and do not allow them to generate words to communicate new ideas. For example, an AAC communicator may operate a system composed of just 50 pictures or 100–200 ideographic symbols. They do not have access to the many thousands of concepts and ideas that they need in order to communicate fully and effectively. The use of orthographic literacy skills can be one way to open up access to a full range of concepts and vocabulary to students who use AAC. The literate AAC communicator, using traditional orthography, may spell words that are not printed on their communication boards or indicate first letters of words to which they don't have access on their communication system. In this way, they can use literacy skills to communicate in face-to-face interactions. The literacy development of augmentative communicators also may provide them with a means to participate in society by using written communication (as others also use written communication) to express opinions and give information. Using literacy as others do may help the bilingual AAC communicator advocate for bilingual education and acquire a sense of belonging to society as well as a stronger sense of value. The third way that literacy development carries added importance for bilingual AAC users involves vocational opportunities. In North America, there are very few individuals who use AAC systems who are competitively employed. The number holding white-collar jobs is few. The range of job opportunities available to individuals who have physical disabilities in general is restricted. AAC communicators are not usually employed in jobs requiring manual labor. Thus, they may need highly developed literacy skills for jobs involving, for example, data entry or word processing. Given limited vocational opportunities, the role of literacy in job preparation for bilingual AAC communicators is critical. Yvonne's Story AAC users must rely on innovative and sometimes creative strategies to learn to read, write, and monitor their understanding of what they are reading. Literacy-learning strategies for bilingual AAC users have not received as much attention as those of monolingual users. Some of the unique struggles and successes of literacy attainment can be seen in the story of Yvonne, a young Puerto Rican AAC user. Yvonne provides a wonderful example of the importance of first-language support and the use of specific literacy-learning strategies for bilingual AAC users. Yvonne is a 10-year-old girl with cerebral palsy of the spastic quadriplegic variety. She is nonambulatory and limited-speaking secondary to cerebral palsy. Her hearing and vision are within normal limits. During my initial contact with Yvonne, her intellectual functioning had not been formally determined. Yvonne's family immigrated to the United States one year before my initial contact with them. She is an only child. The primary language of the home is Spanish. Her father had limited English proficiency and her mother spoke no English at the time of my initial contact, although over the course of the school year they gained more proficiency. Another important characteristic of this family was the fact that the parents decided not to have any other children in order to devote total attention to Yvonne's education and health needs. Although no extended family lived in the area, they resided in a supportive neighborhood with other Puerto Ricans. Yvonne communicated primarily through use of an eye-gaze communication board. She used Mayer-Johnson Symbols and usually had a maximum of six symbols on her board. Other methods of communication included a smile/frown, yes/no response. A smile meant yes and a frown meant no. Yvonne also communicated by directing her eyes toward people or items that she wanted. Yvonne was not reading or writing very much in English when we first met. She may have recognized some English words that she encountered daily such as the names of her school, teacher, and classmates, and she had limited environmental vocabulary. I was not sure of her exact reading proficiency in Spanish; however, she did not demonstrate the ability to independently read upper-elementary-graded text w ritten in Spanish and answer basic content questions. Her listening comprehension for stories read to her in Spanish was good. We were not able to assess written language use because the classroom lacked the technology for text composition. Yvonne had a strong desire to learn to read more proficiently. Yvonne was a student in a general elementary school located in western Massachusetts. Her classroom was nongraded, but the students, all classified as special needs, were of comparable ages to those of fourth grade. The room was self-contained and designated by the school system as a special education classroom. The special need categories included physically and cognitively impaired. Half of the class comprised other Puerto Rican children. My role was that of AAC literacy consultant, but I also brought my expertise in the area of multiculturalism in speech-language pathology. My initial meeting with Yvonne occurred early in the school year, in her classroom with the classroom teacher and instructional aide. Yvonne immediately greeted me with a welcoming smile because she appeared to know that I was there especially to help her learn. During my initial meeting I was able to informally assess that Yvonne had good cognitive skills. She used her voice to initiate communication to bring attention to matters of need or interest. She laughed appropriately at jokes, her eyes followed speakers in a conversation, and she spontaneously used her eyes to appropriately answer yes/no questions. All of the conversations around her and directed to her by her teacher were in English. Yvonne obviously acquired some English proficiency, although she may not have understood everything. I had formal training in Spanish and worked some years earlier in a predominately Mexican-American school district in Southern California where I used the language daily. Although I lacked confidence in my use of Spanish, I greeted Yvonne and introduced myself in Spanish. Approaching her using Spanish set a tone for Yvonne that I was supportive of her background and language usage. She recognized that I needed help using the dialect of Spanish that she was familiar with as a primary way of communicating with her. We learned quickly to work together around the use of a language system. Honoring her first language was important to our working together. Another important factor was Yvonne's desire and willingness to learn English, which contributed significantly to her rapid acquisition of stronger English proficiency. On my second day of visiting the classroom, I was extremely pleased to meet the school SLP assigned to Yvonne. This wonderfully competent, energetic clinician just happened to be bilingual in English and Spanish. With a bilingual SLP and my knowledge of literacy-learning techniques for AAC users, Yvonne blossomed over the course of that academic year in her English proficiency and particularly in her ability to read and spell. A Successful Technique I first introduced a spelling/word-level reading technique to Yvonne that proved to be highly successful and allowed her to gain 10–12 new words in reading recognition and spelling each week. Upper-elementary-aged bilingual AAC users with profiles similar to Yvonne should start with whole-word-level reading aimed at teaching recognition of entire words such as swim, pool, the, or cap. Instruction of whole words leads to success in reading phrases and simple sentences quickly. Phonetic instruction should occur as well. The Words on the Wall technique, which can be used with monolingual as well as bilingual AAC users, begins by the teacher selecting approximately 3–5 new words that the student needs to learn. These should be words relevant to familiar situations and not spelling words from a spelling book. For example, Yvonne went swimming each week in school and thus, during her first week, she learned the words swimming, towel, pool, water, and splash. These words were initially introduced in Spanish only. The next step in this technique is to make the word accessible by writing it in large print on a sentence strip and attaching it to the wall. The word may initially be paired with a symbol, with the symbol being phased out over time leaving just the written word. The student and the teacher define the word and talk about events involving the target word. After all of the target words are discussed and displayed on the wall, the teacher asks the student to identify each word one at a time as in a spelling test. Yvonne used eye gaze to identify her target words. During the next day or week, depending on how well the student masters each set of words, introduce more words (1–3 a day). Leave all words on the wall for the school year, increasing the number of words each week. Review old and new words. After enough words are mastered, have the student begin to read simple sentences. Introduce words such as a and the to allow formation of sentences. The school SLP delivered all of the training to Yvonne in Spanish first and followed it with English only after she knew that Yvonne understood the word in Spanish. Because this literacy-learning technique is based primarily at the word level, it is easier to transition from the Spanish to the English word. The school SLP also kept in close contact with Yvonne's parents, phoning them and sending home each week the word that Yvonne was working on. Yvonne's communication reflected her increased vocabulary. A board in Spanish was sent home and used with her parents and an English board was used at school initially. As Yvonne's parents gained more English proficiency, they requested to have the English communication board as well. During the school year we piloted different types of high-tech AAC devices and switches with Yvonne. We also explored technology for writing purposes during this year. Assessment Words on the Wall lends itself to a Maze Reading Assessment technique once a student has acquired reading of simple sentences. This technique involves the deletion of target words in a sentence leaving a blank space. The student should be provided with three alternative words in random order at each blank (correct choice, incorrect choice of the same part of speech, incorrect choice of a different part of speech). For example: The boy ate a ______ (truck, this, banana). This technique can be used easily with many AAC users. Yvonne's eye gazed to her chosen word using this technique. The scale of reading proficiency most often used for informal reading assessments such as this is 90% accuracy indicating that the student is reading at an independent level, 60%–80% accuracy relating to a level where more instruction is needed, and below 60% is equivalent to a frustration level. For Yvonne, the Maze technique was delivered in English because she already had mastered the words on the wall and read simple sentences in English. The Words on the Wall technique and a Maze Reading Assessment Technique are two techniques that can be culturally and linguistically sensitive and used well with AAC users. Voice output is not required for these techniques, and the words are derived from the students' existing linguistic bases or contextual experiences. Other techniques also can be used to facilitate literacy development with bilingual AAC users. Techniques that contextualize instruction in the experiences of the home and first language are desirable. For young bilingual literacy-language learners, it is important to use interactive learning techniques that involve the teacher, peers, and the AAC user. Techniques that allow students to demonstrate competence in using language and literacy throughout the school day in all instructional activities are greatly beneficial. Techniques that use narratives such as storytelling, listening to stories, or writing are good for content development. These narratives should be delivered in the language that will allow the child to gain academic skill while learning English. My first year with Yvonne was a successful one. She gained approximately 10 new words a week over the course of the school year. For AAC users similar to Yvonne in age and cognitive ability, this is an expected rate of growth. There is no typical rate of growth for all AAC users because this population is so diverse in skill and ability. The Next Year I returned to visit Yvonne the next year when she had been promoted to a new class and school. The successful learning environment that she had previously experienced had come to an abrupt end. There was a lack of continuity with her education from the previous year. Yvonne was in a new school with all new staff. There was no Spanish language support. The literacy-learning methods had been abandoned. Communication with Yvonne was a problem. There was limited communication between the school and home. I spent the first few days in Yvonne's classroom as a participant observer and quickly assessed the social and literacy-learning needs of everyone involved in Yvonne's schooling. The goals of my intervention with Yvonne during this second school year included elimination of the communication problem between the school and the family and establishment of better trust and communication, reestablishment of appropriate instructional methods, eliminating AAC barriers, and supporting cultural identity through literacy lessons/interactions and development of a more efficient communication system. The lack of Spanish support and having to demonstrate and convince the new teachers of Yvonne's literacy-learning capabilities resulted in lost time in her development. Strong first-language support and knowledge of specific literacy-learning techniques for bilingual AAC users led to a successful outcome for Yvonne. She enjoys reading and had a strong desire to continue reading and learning English. This was compatible to the wishes of her parents. Like Yvonne, not all bilingual AAC users have significant difficulties learning to read and write; however, many of them do. Therefore, it becomes important to communication disorders specialists to identify variables of language that are predictive of later reading difficulties. Researchers and other professionals from different fields of study are combining their interests to close the knowledge/information gap that exists between what is already known about bilingual AAC users' acquisition and development and the information needed to help develop intervention strategies for successful written language. Strategies for Monolingual Clinicians: A Postscript Although I did have formal training in Spanish in high school and college and had worked in a predominately Spanish-speaking community in Southern California, I still lacked confidence to converse with Yvonne in Spanish when I first met her. I knew that there were many dialects of Spanish, and I initially did not know enough about the Spanish that she and her family used. Clinicians who are monolingual or who lack information about a second-language-speaking student must do the research to find linguistic information particular to that student. Such knowledge is also helpful in understanding the contexualized uses of literacy in the home that will complement those used in the classroom. General professional development in the area of bilingual literacy learning is highly recommended, as is professional development in AAC. Understanding policies in educating bilingual students that are implemented in your school district is important. Clinicians should understand how policy affects access to instruction for bilingual students. Social, cultural, and economic issues that affect student learning and instruction also should be well understood. It is helpful to gain information from parents, other teachers, and community members about ways that they find helpful in instructing bilingual AAC users. Ovetta Harrison-Harris is chair of the department of communication sciences and disorders at Howard University. She is project director for a U.S. Department of Education Office of Special Education and Rehabilitative Services-funded graduate training program in AAC with an emphasis in multiculturalism and literacy development. For More Information Light J., Binger C., & Smith A.K. (1994). Story reading interactions between pre-schoolers who use AAC and their mothers. Augmentative and Alternative Communication, 10, 225–268. CrossrefGoogle Scholar Light J., & McNaughton D. (1993). Literacy and Augmentative and Alternative Communication (AAC): Expectations and Priorities of Parent and Teachers. Topics and Language Disorders, 13(2), 33–46. CrossrefGoogle Scholar Light J., & Smith A.K. (1993). Home literacy experience of pre-schoolers who use augmentative communication systems and their non disabled peers. Augmentative and Alternative Communication, 9, 10–25. CrossrefGoogle Scholar Pearson B.Z., Fernandez S., & Oller D.K. (1993a). Lexical developmental in simultaneous bilingual infants: Comparison to monolinguals. Language Learning, 43, 93–120. CrossrefGoogle Scholar Pearson B.Z., Fernandez S.C., & Oller D.K. (1993b). Lexical development in bilingual infants and toddlers: Comparison to monolingual norms. Language Learning, 43(1), 93–120. CrossrefGoogle Scholar Pearson B.Z., Fernandez S., & Oller D.K. (1995). Cross-language synonyms in the lexicons of bilingual infants: One language or two?, Journal of Child Language, 22, 345–68. CrossrefGoogle Scholar Pearson B.Z., Oller D.K., Umbel V.M., & Fernandez M.C. (1996, October). The Relationship of Lexical Knowledge to Measures of Literacy and Narrative Discourse in Monolingual and Bilingual Children. Paper presented at the Second Language Research Forum, Tucson, Google Scholar Tucker A perspective on and bilingual education Google Scholar Ovetta L. is chair of the department of communication sciences and disorders at Howard University. She is project director for a U.S. Department of Education Office of Special Education and Rehabilitative Services-funded graduate training program in AAC with an emphasis in multiculturalism and literacy development. of the ASHA Special Augmentative and Alternative Communication, and a for With Communication to your in Nov &
Introduction to WordNet: an on-line lexical database. International Journal of Lexicography, 3(4):235--44. Miller, G. and Charles, W. (1991). Contextual correlates of semantic similarity. Language and Cognitive Processes, 6(1):1--28. Moon, R., editor (2000). Collins Cobuild Dictionary of Phrasal Verbs. Harper Collins. M.P. Marcus, B. S. and Marcinkiewicz, M. (1993). Building a large annotated corpus of english: the penn treebank. Computational Linguistics, 19(2):313--30. Pearce, D. (2002). A comparative evaluation of collocation extraction techniques. In Proceedings of Third International Conference on Language Resources and Evaluation. Pedersen, T. (2002). Distance 0.1. Pulman, S. G. (1993). The recognition and interpretation of idioms. In Cacciari, C. and Tabossi, P., editors, Idioms: Processing, Structure and Interpretation, chapter 11. Lawrence Erlbaum Associates, Hillsdale, NJ. Sag, I., Baldwin, T., Bond, F., Copestake, A., and Flickinger, D. (2002). Multiword expressions: A
This paper proposes a novel class of PCFG parameterizations that support linguistically reasonable priors over PCFGs. To estimate the parameters is to discover a notion of relatedness among context-free rules such that related rules tend to have related probabilities. The prior favors grammars in which the relationships are simple to describe and have few major exceptions. A basic version that bases relatedness on weighted edit distance yields superior smoothing of grammars learned from the Penn Treebank (20 % reduction of rule perplexity over the best previous method). 1 A Sketch of the Concrete Problem This paper uses a new kind of statistical model to smooth the probabilities of PCFG rules. It focuses on “flat ” or “dependency-style ” rules. These resemble subcategorization
Lexical-Functional Grammar f-structures are abstract syntactic representations approximating basic predicate-argument structure. Treebanks annotated with f-structure information are required as training resources for stochastic versions of unification and constraint-based grammars and for the automatic extraction of such resources. In a number of papers (Frank, 2000; Sadler, van Genabith and Way, 2000) have developed methods for automatically annotating treebank resources with f-structure information. However, to date, these methods have only been applied to treebank fragments of the order of a few hundred trees. In the present paper we present a new method that scales and has been applied to a complete treebank, in our case the WSJ section of Penn-II (Marcus et al, 1994), with more than 1,000,000 words in about 50,000 sentences.
Among the resources developed in SI-TAL (Integrated Systems for the Automatic Treatment of Language), ItalWordNet (IWN) was built as reference semantic database, enlarging the Italian WordNet developed in the framework of the European project EuroWordNet (EWN). The Italian lexical database was increased, by introducing and codifying, besides the new grammatical categories of the adjectives and adverbs, a subset of proper names. In the IWN context, the subset of proper names represents a quantitatively limited portion, about 3600 synsets, but it may become a qualitatively important extension. The ever growing amount of non-structured information, stored in natural language, requires the availability of computational instruments able to manage this kind of information where proper names show a remarkable incidence in any types of texts. The work here presented falls in this context, taking into account the proper names, and is focussed on: i) encoding in the IWN database; ii) more typical uses in either proper or metaphorical and metonymic ways such as textual corpora evidence; iii) possibility of a well reasoned and structured enlarging of this data on the basis of the recent experience carried out in IWN. 1. Building the set of proper names in the IWN database IWN was first developed within the EWN1 project (Vossen, 1999) and then extended in the framework of an
DOAJ is a unique and extensive index of diverse open access journals from around the world, driven by a growing community, committed to ensuring quality content is freely available online for everyone.
'I~-eebanks, such as the Penn Treebank (PTB), offer a simple approach to obtaining a broad (:overage grammar: one can simply read the g rammar off the parse trees in the treebank. While such a g rammar is easy to obtain, a square-root rate of growth of the rule set with corpus size suggests that the derived grammar is far fi'om complete and that much more treebanked text would be required to obtain a complete grammar, if one exists at some limit. However, we offer an alternative explanation in terms of the underspecification of structures within the treebank. This hypothesis is explored by applying an algorithm to compact the derived grammar by eliminating redundant rules rules whose right hand sides can be parsed by other rules. The size of the resulting compacted grammar, which is significantly less than that of the full t reebank grammar, is shown to approach a limit. However, such a compacted grammar does not yield very good performance figures. A version of the compaction algorithm taking rule probabilities into account is proposed, which is argued to be more linguistically motivated. Combined with simple thresholding, this method can be used to give a 58% reduction in g rammar size without significant change in parsing performance, and can produce a 69% reduction with some gain in recall, but a loss in precision. 1 I n t r o d u c t i o n The Penn Treebank (PTB) (Marcus et al., 1994) has been used for a ra ther simple approach to deriving large grammars automatically: one where the g rammar rules are simply 'read off' the parse trees in the corpus, with each local subtree providing the left and right hand sides of a rule. Charniak (Charniak, 1996) reports precision and recall figures of around 80% for a parser employing such a grammar. In this paper we show that the huge size of such a treebank grammar (see below) can be reduced in size without appreciable loss in performance, and, in fact, an improvement in recall can be achieved. Our approach can be generalised in terms of Data-Oriented Parsing (DOP) methods (see (Bonnema et al., 1997)) with the tree depth of 1. However, the number of trees produced with a general DOP method is so large that Bonnema (Bonnema et al., 1997) has to resort to restricting the tree depth, using a very domain-specific corpus such as ATIS or OVIS, and parsing very short sentences of average length 4.74 words. Our compaction algorithm can be easily extended for the use within the DOP framework but, because of the huge size of the derived grammar (see below), we chose to use the simplest PCFG framework for our experiments. We are concerned with the nature of the rule set extracted, and how it can be improved, with regard both to linguistic criteria and processing efficiency. In what tbllows, we report the worrying observation that the growth of the rule set continues at a square root rate throughout processing of the entire t reebank (suggesting, perhaps tha t the rule set is far from complete). Our results are similar to those reported in (Krotov et al., 1994). 1 We discuss an alternative possible source of thi,~ rule growth phenomenon, partial bracketting, and suggest that it can be alleviated by compaction, where rules that are redundant (in a sense to be defined) are eliminated from the grammar. Our experiments on compacting a PTB tree1For the complete investigation of the grammar extracted from the Penn Treebank II see (Gaizauskas, 1995)
To test the hypothesis that lactate plays a central role in the distribution of carbohydrate (CHO) potential energy for oxidation and glucose production (GP), we performed a lactate clamp (LC) procedure during rest and moderate intensity exercise. Blood [lactate] was clamped at approximately 4 mM by exogenous lactate infusion. Subjects performed 90 min exercise trials at 65 % of the peak rate of oxygen consumption (V(O(2))(,peak); 65 %), 55 % V(O(2))(,peak) (55 %) and 55 % V(O(2))(,peak) with lactate clamped to the blood [lactate] that was measured at 65 % V(O(2))(,peak) (55 %-LC). Lactate and glucose rates of appearance (R(a)), disappearance (R(d)) and oxidation (R(ox)) were measured with a combination of [3-(13)C]lactate, H(13)CO(3)(-), and [6,6-(2)H(2)]glucose tracers. During rest and exercise, lactate R(a) and R(d) were increased at 55 %-LC compared to 55 %. Glucose R(a) and R(d) were decreased during 55 %-LC compared to 55 %. Lactate R(ox) was increased by LC during exercise (55 %: 6.52 +/- 0.65 and 55 %-LC: 10.01 +/- 0.68 mg kg(-1) min(-1)) which was concurrent with a decrease in glucose oxidation (55 %: 7.64 +/- 0.4 and 55 %-LC: 4.35 +/- 0.31 mg kg(-1) min(-1)). With LC, incorporation of (13)C from tracer lactate into blood glucose (L GNG) increased while both GP and calculated hepatic glycogenolysis (GLY) decreased. Therefore, increased blood [lactate] during moderate intensity exercise increased lactate oxidation, spared blood glucose and decreased glucose production. Further, exogenous lactate infusion did not affect rating of perceived exertion (RPE) during exercise. These results demonstrate that lactate is a useful carbohydrate in times of increased energy demand.
Is there a general model that can predict the perceived phrase structure in language and music? While it is usually assumed that humans have separate faculties for language and music, this work focuses on the commonalities rather than on the differences between these modalities, aiming at finding a deeper 'faculty'. Our key idea is that the perceptual system strives for the simplest structure (the 'simplicity principle'), but in doing so it is biased by the likelihood of previous structures (the 'likelihood principle'). We present a series of data-oriented parsing (DOP) models that combine these two principles and that are tested on the Penn Treebank and the Essen Folksong Collection. Our experiments show that (1) a combination of the two principles outperforms the use of either of them, and (2) exactly the same model with the same parameter setting achieves maximum accuracy for both language and music. We argue that our results suggest an interesting parallel between linguistic and musical structuring.
This paper describes an approach to treebank development which relies on the manual development of annotation tools. The overall process of tree annotation is described, and a special emphasis is put on the description of the last tool which has been built, i.e. a dependency-based robust chunk parser. The modularization of the parser and the central role of verbal subcategorization is presented. The first experimental results, carried out on a corpus of 645 sentences are reported and discussed. 1.
In this paper, we describe experiments on HPSG parse disambiguation using the Redwoods HPSG treebank (Oepen et al. 2002a,b,c). HPSG is a constraint-based lexicalist ("unification") grammar formalism The
REVIEWS 529 Zubova, L. V. Sovremennaia russkaia poeziia v kontekste istorni iazyka.Novoe literaturnoe obozrenie, Moscow, 2000. 432 pp. Notes. Bibliography. Index. Priceunknown. ZUBOVA'S study offers a detailed and provocative analysis of Russian poetry of the I96o-9os. It discusses almost three hundred authors, including leading postmodernist figures such as Joseph Brodsky, Viktor Krivulin, Genrikh Sapgir, Sergei Stratanovskii, Dmitrii Prigov, Elena Shvarts and Viktor Sosnora. Zubova highlights playful and innovative aspects of Russian postmodernistpoetry, arguingthat many linguisticexperimentsembedded in the texts under scrutiny in the present study explore the shortcomings and inadequacyof contemporaryRussianlanguage. Zubova arguesthat linguistic games of Russian postmodernistpoets offeralternativeways of development for phonetic, semantic and grammar structures of the Russian language. Zubova holds an optimisticview that deviations from the standardlanguage, as observed in the texts she studies, do not destroy the language, but help to preserve it, especially because they resurrect from oblivion some forgotten linguistic norms from the past (p. 399). Zubova's main thesis is based on the belief that language 'is a self-correctingsystem as well as a combination of options to express various meanings' (p. 399). The book will be of great interestto linguistsand to studentsof Russianpoetry, since it offersimportant insightsinto today'sstateof the Russianlanguage. The book comprises seven chapters, including an introduction and conclusion. Chapter one outlines the main theoretical frameworkwhich is applied throughout the book; chapter two discusses phonetic aspects of contemporary poetic experiments; chapter three investigates etymological innovations; chapter four is devoted to lexical changes; chapter five talks about archaic aspects of various grammaticaldeviations;chapter six offersa detailed analysis of the linguistic games based around gender; and chapter seven analysessyntacticalstructuresof the texts. In addition, the book offersa briefsummaryof the main conclusions (pp. 398-99), bibliographyand index. To some extent, all the texts Zubova discussesmight be viewed as hypertext, with no significantdifferentiationbetween the language of the i 960S and of the I990s. Furthermore,Zubova revealspostmodernisttendencies in Russian poetry of this period by demonstratinghow the linguisticexpressionis linked to the postmodernist worldview of the authors she discusses. Zubova encouragesher readersto considersome seeminglybad poems and treatthem with a sensitivityto the irony and parodic intentions they display.As Zubova states, 'the most constructivedevices of postmodernisttext include irony, selfirony and linguistic game' (p. i i). Zubova also highlights the authors' balancing acts between high and low cultures. In this respect, such poets as Prigov, Sosnora, Shvartsand Krivulin appear to be particularlyimaginative in their use of language, exploring precarious borders between modern and archaic forms of communication, and between elitist and popular forms of expression. Zubova's findings illustratewell the intrinsicbond between postSoviet poetry and Soviet undergroundliterature.Zubova'sstudyis a welcome additionto the extensiveanalysisof postmodernistfictionundertakenby Mark Lipovetsky (RussianPostmodernist Fiction.Dialoguewith Chaos,Armonk, NY, 530 SEER, 8o, 3, 2002 I999). Zubova is well aware of the metatextual qualities of Russian postmodernism, and points to the intertextualgames of varioustexts. In addition to the more familiarnames, Zubova introducesher audience to lesser known authors, such as Vladimir Strochkov, Ian Satunovskii, and Vladimir Erl'.Unfortunately, Zubova's studydoes not contain any biographical detailsof the authorsshe quotes. Such an appendixwould be a usefultool for assessingthe spreadof linguisticdeviations, from the point of view of age groups, regional variations and the aesthetic preferences of the poets. It is difficultto assess whether some of the deviations from established linguistic norms were intentional, or derive from the contemporary sloppy usage of Russianlanguage that isparticularlynoticeable in post-Soviet Russianmedia. Krivulinand Shvarts,for example, are philologistsby training,and therefore are more inclined to have playful appropriationof some idioms, or absurd examplesof Soviet newspeak. As Zubova's study demonstrates, numerous poetic experiments reflect on the fluid state of the Russian language itself. In this respect, Zubova's discussion of the satirical elements relating to the concept of gender in contemporary Russian poetry is particularlyrewarding. Zubova's examples from Russian poetry reveal, for example, the uncertaintiesrelatingto gender of animals. Thus, some poets use the feminine form of the noun koshka (cat) with the additional note that it is used in their poem as a noun of masculine gender. Such examples are both amusing and obscure. Zubova suggeststhat contemporary Russian poets are struggling to revive the use of the neuter gender that otherwise has been steadily disappearing from the standard language (p. 301...
While initial treebanks and treebank parsers primarily involved surface analysis, recent work focuses on predicate argument (PA) structure. PA structure provides means to regularize variants (e.g., actives/passives) of sentences so that individual patterns may have better
his contribution evaluates some aspects ofthe reversing ofthe Dutch-Estonian electronic bilingual dictionary database to Estonian-Dutch. The project has linked two monolingual lexical databases and added new lexical and example units with the editor tool OMBI. The links are provided with information about the status of equivalence. The two sources are the Dutch Reference File and an Estonian database ofpolysemous words. The strategies ofderiving correct polysemy representations ofthe Estonian items in the course of editing the Dutch-Estonian dictionary are be evaluated. Prior to dictionary editing, an Estonian reference file for polysemous words was created. In the course of editing, many missing entries and senses were added. The Estonian reference file consists of three structurally different parts: first, the left side of another bilingual dictionary, second, a database ofamonolingual dictionary, third, a part created specially for the database. It is argued that the high quality ofthe target language database and a correct specification ofthe equivalence information are crucial for successful reversing. Verbal polysemy and its relation to the Estonian object case have posed a major challenge for the project.
OBJECTIVE: Since Freud's "Interpretation of Dreams," sleep has been related to emotional functions, where dreams were assumed to play a cathartic role. In psychophysiological research, this role was attributed mainly to rapid eye movement (REM) sleep. The present study compared processing pictures with negative emotional impact over intervals covering either early sleep dominated by slow-wave sleep (SWS) or late REM sleep-dominated sleep. METHOD: Emotional reactions were assessed by a nonverbal rating procedure along the two emotional dimensions valence (positive vs. negative) and arousal (low vs. high). Two groups of healthy men were tested across 3-hour periods of early and late nocturnal sleep (sleep group) or corresponding intervals filled with wakefulness (wake group). After the intervals, subjects rated new pictures together with old pictures already presented before the interval. Sleep was recorded polysomnographically. RESULTS: As expected, the amount of REM sleep was about three times greater during late than early nocturnal sleep, whereas a reversed distribution was observed for SWS (p<.001). Valence ratings indicated a shift toward enhanced negative ratings after late sleep (p<.05), contrasting with a trend toward more positive ratings after early sleep (p<.10). Arousal habituated slightly to repeated presentation of the same stimuli, but sleep generally enhanced subsequent arousal ratings (p<.05). Effects of sleep did not depend on whether pictures had low or high emotional impact. CONCLUSIONS: Indicating a priming-like enhancement of emotional reactivity after periods rich in REM sleep, results do not confirm a cathartic function of REM sleep or sleep in general.
In this paper we present the Alpino Dependency Treebank and the tools that we have developed to facilitate the annotation process. Annotation typically starts with parsing a sentence with the Alpino parser, a wide coverage parser of Dutch text. The number of parses that is generated is reduced through interactive lexical analysis and constituent marking. A tool for on line addition of lexical information facilitates the parsing of sentences with unknown words. The selection of the best parse is done efficiently with the parse selection tool. At this moment, the Alpino Dependency Treebank consists of about 6,000 sentences of newspaper text that are annotated with dependency trees. The corpus can be used for linguistic exploration as well as for training and evaluation purposes.
We presented the development process and the technical specifications of K-CDA IG. We explored how the results can be used as interoperability criteria in the national EHR systems certification program. Finally, we provided recommendations that could guide other entities planning their HIE programs.
"Things seem to decline..." - Language, ethnicity and identity illustrated by material from a former Swedish colony in Misiones, Argentina and dialect material from Bjurholm, Sweden Nowadays there is a universal tendency towards convergence and simplification in several official European languages, the existence of non-codified languages is threatened and dialectal varieties are subject to levelling. If minority languages and intralinguistic varieties are to survive, this depends to a great extent on the identity of the individual speaker and the values and attitudes attached to his/her variety and speech behaviour. Besides, it is also due to the size of the language group, sharing the same values, and its cultural activities, forming part of its tradition. For that reason I have selected material from two threatened speech communities: one from a former Swedish colony in Misiones, Argentina, the other from Bjurholm, a small dialect-speaking community in the interior of Västerbotten, Sweden, in order to study the mechanisms causing the preservation or loss of the linguistic varieties as well as the cultural boundaries. This study consists of three parts and is divided into eleven chapters. The first part consists of chapter 1-4. Initially concepts of ethnicity, identity and culture are discussed in the light of different sciences, i.e. social anthropology (Barth 1969; Hylland Eriksen 1998), ethnology, the sociology of language (Fishman 1989) and sociolinguistics (Edwards 1985). In the following chapter emigrant material (narratives, letters, local history) forms part of the historical dimension: from the cultural contacts of individuals arriving in Brazil to the later Swedish settlement in Misiones, where it is appropriate to talk about an ethnic group, its collective history and Swedishness. In chapter 3 the continuity of the cultural heritage is illustrated by onomastic material (personal names) from three generations of Swedish descendants. Chapter 4 is a report of investigations carried out in the 1990's. In 1999, the Swedish language had been maintained among 20 of the 32 informants of Swedish descent, each one representing one family network. Their identity is hyphenated: they are all Argentine but of Swedish descent, and constitute an ethnic group. 24 of them had grown up with Swedish as their first language. Language attitudes had become more positive since 1988, when a similar investigation took place. Lately a new denomination has appeared: Los Nordicos, and it is discussed whether it is ethnic or not. Part two consists of chapter 5-10 and is the result of the project Dialects in change. The principal questions in this pilot study concern dialect boundaries: are they still maintained or subject to levelling? Which dialectal items are used as boundary markers and which are substituted by standard forms? A dialect boundary implies that dialects are still used in fairly genuine forms and that a local norm is prevailing, which seems to be the case in Bjurholm. Via three types of data, including inquiries, two tests: on dialectal vocabulary (50 lexical items) and translation of twelve standard sentences into dialect versions, besides type recordings of authentic speech, the author has tried to describe the local norm, based on individual micro data. In chapter 9 criteria for dialect variables used in quantitative studies are discussed. Based on this material there is evidence for dialect boundaries towards Lappland as well as towards the neighbour parish of Vindeln (former Degerfors). This material has to be extended to serve as a base for general conclusions and methods must be refined for further investigations. This will be possible, as the old dialects are still in use among the two oldest generations of adults. If they are to be used in the future depends on the younger generation (25-^4-0 years). In the final chapter data from the bilingual Swedish speakers in Misiones and the dialect speakers of Bjurholm are summarized, discussed and compared. For linguistic survival three key concepts are important: contact, prestige and identification, which can be related to ethnicity on a group level and identity on an individual level. There are striking similarities between them. In both cases there are boundaries between "us" and "them", but these are more subtle in an intralinguistic perspective. Both categories are using spoken varieties in a transitional stage, as Misiones Swedish soon will be extinct and the genuine Bjurholm dialect subject to levelling. Both varieties are also informal and diglossie in function, although codes are not always strictly kept apart. While the use of Misiones Swedish is reduced to the family sphere, the Bjurholm dialect can be extended to a wider range of domains. The great difference seems to concern the history of the varieties and the size of the group of speakers: the Bjurholm dialect can be traced back to at least 1750, maybe even to dialect splitting in the medieval time, while Misiones Swedish has been used for about 100 years by three generations of speakers, nowadays reduced to a number of approximately 150 persons.
Quantitative evaluation of parsers has traditionally centered around the PARSEVAL measures of crossing brackets, (labeled) precision, and (labeled) recall. However, it is well known that these measures do not give an accurate picture of the quality of the parsers output. Furthermore, we will show that they are especially unsuited for partial parsers. In recent years, research has concentrated on dependencybased evaluation measures. We will show in this paper that such a dependency-based evaluation scheme is particularly suitable for partial parsers. TüBa-D, the treebank used here for evaluation, contains all the necessary dependency information so that the conversion of trees into a dependency structure does not have to rely on heuristics. Therefore, the dependency representations are not only reliable, they are also linguistically motivated and can be used for linguistic purposes.
The traditional notion of word meaning used in natural language processing is literal or lexical meaning as used in dictionaries and lexicons. This relatively objective notion of lexical meaning is different from more subjective notions of emotive or affective meaning. Our aim is to come to grips with subjective aspects of meaning expressed in written texts, such as the attitude or value expressed in them. This paper explores how the structure of the WordNet lexical database might be used to assess affective or emotive meaning. In particular, we construct measures based on Osgood’s semantic differential technique. Suppose we can evaluate the subjective meaning expressed in a text. This would allow us to classify documents on subjective criteria, rather than on their factual content. There are several potential applications for such classifications, for example, providing summary statistics for search engines. Given the query “Crete travel review”, a search engine could report, “There are 1000 hits of which 3/4 is a positive review”. Another potential application is filtering “flames ” for newsgroups.
1.1 The importance of developing sociocultural competence If we were to meet an adult native speaker who had grown up in a place where there were no other people, but sufficient language input, through, for example, tapes, for that person to be in linguistic terms a fluent speaker, then it seems reasonable to say that this person would in all likelihood be regarded as socially dysfunctional. Our unfortunate would not know how to deal with the most simple situations, and unless he or she were protected and educated, a sorry end may well be just around the corner. While such a case is fantastical, the non-native speaker (NNS) who arrives in an alien culture which is markedly different from their own and who lacks sociocultural knowledge is in a position with certain parallels to that of a socially inadequate individual (Furnham 1993). While knowledge could be transferred from the native culture, there is no way of guessing correctly what the possible cultural differences or similarities are. Native speakers (NSs) would be unaware of the visitor’s lack of sociocultural knowledge (Blum-Kulka 1997), and both NNSs and NSs may even be unaware that cultures can vary as much as they do (Hinkel 2001). NSs are also likely to be find behaviour that runs counter to their society’s beliefs or norms unacceptable, and to react accordingly. After perusing Celce-Murcia et al’s (1995) list of sociocultural factors (see appendix 1), it is not difficult to see how inappropriacy in any of the listed areas could lead to problems. The acceptable length of a silence varies across cultures, and one possible reason for some students’ perceived reticence in ESL contexts could be caused by the fact that in certain cultures, people are comfortable with longer response times than is the case in English. Gestures vary across cultures, and are used to express abstract ideas (McCafferty and Ahmed 2000); potential for confusion is therefore plentiful and plain. In a liberal Western country such as England, men coming from a more patriarchal society could easily find themselves being rebuked or criticised, and might feel at a loss as to why. When and to whom the words ‘Thank you’ are required to be said in England is a notoriously confusing area, and a source of much resentment among the inhabitants of towns where there is a constant influx of language learners. While the above examples show the significance of sociocultural factors in communication, the key question is how this knowledge relates to and is formulated in language, and in particular a second language. Pavlenko and Lantolf (2000) argue that traditional models of second language acquisition account for the way we acquire lexical, phonological and grammatical units of knowledge, but that in order to understand language use in context, and therefore the pervasiveness of culture in communication, a model which accounts for learning as participation is necessary. In this model, the learner develops skills which enable him or her to engage with contextual and cultural factors of communication. Although the two models are not mutually exclusive but in fact complementary, the latter is far more appropriate for understanding language as socialisation, as an ongoing process of engagement
The increasing prestige of medicine as a science, accompanied by the social rise of the doctor, in eighteenth-century France is well documented. What I would like to argue here, however, is that there exists a correspondence between the establishment of medicine as an independent field of study in eighteenth-century France and the increasing use and influence of an autonomous form of medical discourse, namely, the aphorism, in this period. This is not so much a question of the language used by the more renowned doctors of the day but of a form of discourse deeply imbued and associated with medical practice. (It is nonetheless true that certain famous physicians combined medical and literary roles. For instance, Theophile Bordeu intervenes significantly in Diderot's Le Reve d'Alembert, and Vicq d'Azyr, Marie-Antoinette's doctor, was elected to the Académie Française in 1788 in a sort of social consecration or medical discourse, implicitly incorporating his medical figure and figures into the socio-linguistic norms of 'le bon usage’ promoted by the Académie itself.) Yet what interests me particularly here is the insinuation of the medical aphorism itself into other fields of late eighteenth-century discourse, notably those of literature and politics, the traditional domains of the maxim.
We present a flexible approach for extracting hierarchical classifications from linguistic data. To this end, the framework of observational logic is introduced, which extends the logic that underlies standard Formal Concept Analysis by allowing disjunctive rules and exclusions. We give a rigorous mathematical characterization of how the chosen rule type affects the structure of the induced hierarchy. The framework is applied to the induction of hierarchical classifications from linguistic databases. The pros and cons of several types of hierarchies are discussed in detail with respect to criteria such as compactness of representation, suitability for inference tasks, and intelligibility for the human user.
In pace with the success of corpus-based approaches to theoretical and computational linguistics, the collocation of corpora has evolved into a research activity in its own. As the currently available corpora either lack annotation depth or closure, more data will be annotated in the future, preferably with minimal human intervention. This paper tries to approach the problem of treebank development from a logic-based learning perspective, applying several alternative forms of inference in order to assess their potential for automatically generalizing from a seed corpus annotated by hand to a corpus of POS annotated sentences, in order to automatically produce syntactic annotations that are good enough to use as training material for a parser. We shall show that syntactic annotations can be created automatically in large quantity via deductive and abductive explanation-based learning (EBL). Although these automatically created structures are not statistically representative with respect to many quantitative aspects of the treebank, the annotations may provide useful qualitative and quantitative data which might be extracted and reinvested into a parser. We shall compare the benefits and investments of automatically created structures to that of human-annotated structures and suggest some possible strategies how EBL approaches can be combined with manual annotation.
The majority of humanities computingprojects within the discipline of literaturehave been conceived more as digital librariesthan monographs which utilise the medium as asite of interpretation. The impetus to conceiveelectronic research in this way comes from theunderlying philosophy of texts and textualityimplicit in SGML and its instantiation for thehumanities, the TEI, which was conceived as ``amarkup system intended for representing alreadyexisting literary texts''. This article exploresthe most common theories used to conceiveelectronic research in literature, such ashypertext theory, OCHO (Ordered Hierarchy ofContent Objects), and Jerome J. McGann's``noninformational'' forms of textuality. It alsoargues that as our understanding of electronictexts and textuality deepens, and as advancesin technology progresses, other theories, suchas Reception Theory and Versioning, may well beadapted to serve as a theoretical basis forconceiving research more akin to an electronicmonograph than a digital library.
Looking specifically at the genre ofadaptive narrative, this article explores thefuture of literature created for and withcomputer technology, focusing primarily on thetrope of mutability as it is played out withnew media. Some of the questions askedare: What can the medium of a work ofliterature, that is its material aspect, tellus about the text? About character? What canit possibly matter if narrative is recounted onpapyrus, retold on parchment and rag, and thenremediated in pixels? Isn't it the messagecarried by the medium we are most concernedwith, stable or unstable throughout the processof inscription, reinscription, encoding anddecoding, translation and remediation? Thispaper speculates about possibilities ratherthan attempts to answer these questions, butthe structuring and mean-making componentsconsidered here stand as examples of some wemay want to think about when developing futuretheories about literature – and all types ofwriting – generated by and for electronicenvironments.