Papers reviewed and determined not to be word norm studies. Use the flag icon to report errors or suggest re-inclusion.
18265 papers
This paper describes a lexicalized tree adjoining grammar (LTAG) based parsing system for Korean which combines corpus-based morphological analysis and tagging with a statistical parser. Part of the challenge of statistical parsing for Korean comes from the fact that Korean has free word order and a complex morphological system. The parser uses an LTAG grammar which is automatically extracted using LexTract (Xia et al., 2000) from the Penn Korean TreeBank (Han et al., 2002). The morphological tagger/analyzer is also trained on the TreeBank. The tagger/analyzer obtained the correctly disambiguated morphological analysis of words with 95.78/95.39% precision/recall when tested on a test set of 3,717 previously unseen words. The parser obtained an accuracy of 75.7% when tested on the same test set (of 425 sentences). These performance results are better than an existing off-the-shelf Korean morphological analyzer and parser run on the same data
This paper deals with up-translation - a process of lexical data transformation from any source format to the XML document. Relevant aspects of the XML format and many related technologies are surveyed first. Then, information content enhancement of existing lexical resources is discussed. The last part brings information about up-translation ofthe Dictionary of Literary Czech Language and the way of efficient storage and retrieval of data.
The Papillon project aims at building a multilingual lexical database for extracting dictionaries. This paper describes the Papillon monolingual lexie structure with an example and propose some changes for spotted problems.
We presented the development process and the technical specifications of K-CDA IG. We explored how the results can be used as interoperability criteria in the national EHR systems certification program. Finally, we provided recommendations that could guide other entities planning their HIE programs.
This paper presents a system to find automatically words from a definition or a paraphrase. The system uses a lexical database of French words that is comparable in its size to WordNet and an algorithm that evaluates distances in the semantic graph between hypernyms and hyponyms of the words in the definition. The paper first outlines the structure of the lexical network on which the method is based. It then describes the algorithm. Finally, it concludes with examples of results we have obtained.
We present a new approach to topological parsing of German which is corpus-based and built on a simple model of probabilistic CFG parsing. The topological field model of German provides a linguistically motivated, flat macro structure for complex sentences. Besides the practical aspect of developing a robust and accurate topological parser for hybrid shallow and deep NLP, we investigate to what extent topological structures can be handled by context-free probabilistic models. We discuss experiments with systematic variants of a topological treebank grammar, which yield competitive results.
his contribution evaluates some aspects ofthe reversing ofthe Dutch-Estonian electronic bilingual dictionary database to Estonian-Dutch. The project has linked two monolingual lexical databases and added new lexical and example units with the editor tool OMBI. The links are provided with information about the status of equivalence. The two sources are the Dutch Reference File and an Estonian database ofpolysemous words. The strategies ofderiving correct polysemy representations ofthe Estonian items in the course of editing the Dutch-Estonian dictionary are be evaluated. Prior to dictionary editing, an Estonian reference file for polysemous words was created. In the course of editing, many missing entries and senses were added. The Estonian reference file consists of three structurally different parts: first, the left side of another bilingual dictionary, second, a database ofamonolingual dictionary, third, a part created specially for the database. It is argued that the high quality ofthe target language database and a correct specification ofthe equivalence information are crucial for successful reversing. Verbal polysemy and its relation to the Estonian object case have posed a major challenge for the project.
We describe an LR parser of parts-of-speech (and punctuation labels) for Tree Adjoining Grammars (TAGs), that solves table conflicts in a greedy way, with limited amount of backtracking. We evaluate the parser using the Penn Treebank showing that the method yield very fast parsers with at least reasonable accuracy, confirming the intuition that LR parsing benefits from the use of rich grammars.
This contribution describes a new tool (named VisDic) for browsing and editing WordNet databases. It was developed in the Natural Language Processing Laboratory at the Faculty of Informatics, Masaryk University. In fact, it is not designed as a specialized tool for processing WordNet data only, generally, it has been developed as a tool for viewing and editing any lexical database as e.g. multilingual dictionaries, monolingual dictionaries, corpora, etc. From this point of view, WordNet can be also understood as a dictionary with special features.
We present an algorithm which translates the Penn Treebank into a corpus of Combinatory Categorial Grammar (CCG) derivations. To do this we have needed to make several systematic changes to the Treebank which have to effect of cleaning up a number of errors and inconsistencies. This process has yielded a cleaner treebank that can potentially be used in any framework. We also show how unary type-changing rules for certain types of modifiers can be introduced in a CCG grammar to ensure a compact lexicon without augmenting the generative power of the system. We demonstrate how the combination of preprocessing and type-changing rules minimizes the lexical coverage problem. 1.
The traditional notion of word meaning used in natural language processing is literal or lexical meaning as used in dictionaries and lexicons. This relatively objective notion of lexical meaning is different from more subjective notions of emotive or affective meaning. Our aim is to come to grips with subjective aspects of meaning expressed in written texts, such as the attitude or value expressed in them. This paper explores how the structure of the WordNet lexical database might be used to assess affective or emotive meaning. In particular, we construct measures based on Osgood’s semantic differential technique. Suppose we can evaluate the subjective meaning expressed in a text. This would allow us to classify documents on subjective criteria, rather than on their factual content. There are several potential applications for such classifications, for example, providing summary statistics for search engines. Given the query “Crete travel review”, a search engine could report, “There are 1000 hits of which 3/4 is a positive review”. Another potential application is filtering “flames ” for newsgroups.
This paper describes a general-purpose sentence generation system that can achieve both broad scale coverage and high quality while aiming to be suitable for a variety of generation tasks. We measure the coverage and correctness empirically using a section of the Penn Treebank corpus as a test set. We also describe novel features that help make the generator flexible and easier to use for a variety of tasks. To our knowledge, this is the first empirical measurement of coverage reported in the literature, and the highest reported measurements of correctness.
This study investigated the human eyeblink startle reflex as a measure of alcohol cue reactivity. Alcohol-dependent participants early (n = 36) and late (n = 34) in abstinence received presentations of alcohol and water cues. Consistent with previous research, greater salivation and higher ratings of urge to drink occurred in response to the alcohol cues. Differential salivary and urge responding to alcohol versus water cues did not vary as a function of abstinence duration. Of special interest was the finding that startle response magnitudes were relatively elevated to alcohol cues, but only in individuals early in abstinence. Affective ratings of alcohol cues suggested that alcohol cues were perceived as aversive. Methodological and theoretical implications of the findings are discussed.
Treebanks are widely recognised as a necessary source of information in NLP as well as in Linguistics studies. In this paper we present and justify methodological principles and syntactic criteria to build a Treebank for Spanish: annotating only explicit information, constituents and syntactic functions and being theory independent. Previous work is also presented in order to account for taken decisions. The annotation process will be done in different steps so that each one of them is the input of the next. We present the basic guidelines of syntactic annotation and the boundaries of the work to be done in a first step: annotation of low constituents and surface functions. Moreover, some semantic information (subject type) is likely to be included.
Lexical-Functional Grammar f-structures are abstract syntactic representations approximating basic predicate-argument structure. Treebanks annotated with f-structure information are required as training resources for stochastic versions of unification and constraint-based grammars and for the automatic extraction of such resources. In a number of papers (Frank, 2000; Sadler, van Genabith and Way, 2000) have developed methods for automatically annotating treebank resources with f-structure information. However, to date, these methods have only been applied to treebank fragments of the order of a few hundred trees. In the present paper we present a new method that scales and has been applied to a complete treebank, in our case the WSJ section of Penn-II (Marcus et al, 1994), with more than 1,000,000 words in about 50,000 sentences.
'I~-eebanks, such as the Penn Treebank (PTB), offer a simple approach to obtaining a broad (:overage grammar: one can simply read the g rammar off the parse trees in the treebank. While such a g rammar is easy to obtain, a square-root rate of growth of the rule set with corpus size suggests that the derived grammar is far fi'om complete and that much more treebanked text would be required to obtain a complete grammar, if one exists at some limit. However, we offer an alternative explanation in terms of the underspecification of structures within the treebank. This hypothesis is explored by applying an algorithm to compact the derived grammar by eliminating redundant rules rules whose right hand sides can be parsed by other rules. The size of the resulting compacted grammar, which is significantly less than that of the full t reebank grammar, is shown to approach a limit. However, such a compacted grammar does not yield very good performance figures. A version of the compaction algorithm taking rule probabilities into account is proposed, which is argued to be more linguistically motivated. Combined with simple thresholding, this method can be used to give a 58% reduction in g rammar size without significant change in parsing performance, and can produce a 69% reduction with some gain in recall, but a loss in precision. 1 I n t r o d u c t i o n The Penn Treebank (PTB) (Marcus et al., 1994) has been used for a ra ther simple approach to deriving large grammars automatically: one where the g rammar rules are simply 'read off' the parse trees in the corpus, with each local subtree providing the left and right hand sides of a rule. Charniak (Charniak, 1996) reports precision and recall figures of around 80% for a parser employing such a grammar. In this paper we show that the huge size of such a treebank grammar (see below) can be reduced in size without appreciable loss in performance, and, in fact, an improvement in recall can be achieved. Our approach can be generalised in terms of Data-Oriented Parsing (DOP) methods (see (Bonnema et al., 1997)) with the tree depth of 1. However, the number of trees produced with a general DOP method is so large that Bonnema (Bonnema et al., 1997) has to resort to restricting the tree depth, using a very domain-specific corpus such as ATIS or OVIS, and parsing very short sentences of average length 4.74 words. Our compaction algorithm can be easily extended for the use within the DOP framework but, because of the huge size of the derived grammar (see below), we chose to use the simplest PCFG framework for our experiments. We are concerned with the nature of the rule set extracted, and how it can be improved, with regard both to linguistic criteria and processing efficiency. In what tbllows, we report the worrying observation that the growth of the rule set continues at a square root rate throughout processing of the entire t reebank (suggesting, perhaps tha t the rule set is far from complete). Our results are similar to those reported in (Krotov et al., 1994). 1 We discuss an alternative possible source of thi,~ rule growth phenomenon, partial bracketting, and suggest that it can be alleviated by compaction, where rules that are redundant (in a sense to be defined) are eliminated from the grammar. Our experiments on compacting a PTB tree1For the complete investigation of the grammar extracted from the Penn Treebank II see (Gaizauskas, 1995)
The Swedish WordNet project aims at building a Swedish version of the EuroWordNet lexical database. The article accounts for some of the problems specific to the building of a Swedish net. As an i...
In this paper we present the Alpino Dependency Treebank and the tools that we have developed to facilitate the annotation process. Annotation typically starts with parsing a sentence with the Alpino parser, a wide coverage parser of Dutch text. The number of parses that is generated is reduced through interactive lexical analysis and constituent marking. A tool for on line addition of lexical information facilitates the parsing of sentences with unknown words. The selection of the best parse is done efficiently with the parse selection tool. At this moment, the Alpino Dependency Treebank consists of about 6,000 sentences of newspaper text that are annotated with dependency trees. The corpus can be used for linguistic exploration as well as for training and evaluation purposes.
We present a flexible approach for extracting hierarchical classifications from linguistic data. To this end, the framework of observational logic is introduced, which extends the logic that underlies standard Formal Concept Analysis by allowing disjunctive rules and exclusions. We give a rigorous mathematical characterization of how the chosen rule type affects the structure of the induced hierarchy. The framework is applied to the induction of hierarchical classifications from linguistic databases. The pros and cons of several types of hierarchies are discussed in detail with respect to criteria such as compactness of representation, suitability for inference tasks, and intelligibility for the human user.
We present a program for Matlab that quickly generates Attneave-style random polygons and families of similar polygons. The function allows a great deal of user control over various aspects of the shape generation process. It also has the ability to detect and eliminate shapes that do not match a variety of user-entered parameters regarding the lengths of the shapes’ sides, vertex angles, and topological form. The function eliminates the time-consuming task of generating such shapes by hand and should allow their broader use in behavioral research. The Matlab script function can be downloaded at www.dal.ca/ ~mcmullen/downloads.html.
1.1 The importance of developing sociocultural competence If we were to meet an adult native speaker who had grown up in a place where there were no other people, but sufficient language input, through, for example, tapes, for that person to be in linguistic terms a fluent speaker, then it seems reasonable to say that this person would in all likelihood be regarded as socially dysfunctional. Our unfortunate would not know how to deal with the most simple situations, and unless he or she were protected and educated, a sorry end may well be just around the corner. While such a case is fantastical, the non-native speaker (NNS) who arrives in an alien culture which is markedly different from their own and who lacks sociocultural knowledge is in a position with certain parallels to that of a socially inadequate individual (Furnham 1993). While knowledge could be transferred from the native culture, there is no way of guessing correctly what the possible cultural differences or similarities are. Native speakers (NSs) would be unaware of the visitor’s lack of sociocultural knowledge (Blum-Kulka 1997), and both NNSs and NSs may even be unaware that cultures can vary as much as they do (Hinkel 2001). NSs are also likely to be find behaviour that runs counter to their society’s beliefs or norms unacceptable, and to react accordingly. After perusing Celce-Murcia et al’s (1995) list of sociocultural factors (see appendix 1), it is not difficult to see how inappropriacy in any of the listed areas could lead to problems. The acceptable length of a silence varies across cultures, and one possible reason for some students’ perceived reticence in ESL contexts could be caused by the fact that in certain cultures, people are comfortable with longer response times than is the case in English. Gestures vary across cultures, and are used to express abstract ideas (McCafferty and Ahmed 2000); potential for confusion is therefore plentiful and plain. In a liberal Western country such as England, men coming from a more patriarchal society could easily find themselves being rebuked or criticised, and might feel at a loss as to why. When and to whom the words ‘Thank you’ are required to be said in England is a notoriously confusing area, and a source of much resentment among the inhabitants of towns where there is a constant influx of language learners. While the above examples show the significance of sociocultural factors in communication, the key question is how this knowledge relates to and is formulated in language, and in particular a second language. Pavlenko and Lantolf (2000) argue that traditional models of second language acquisition account for the way we acquire lexical, phonological and grammatical units of knowledge, but that in order to understand language use in context, and therefore the pervasiveness of culture in communication, a model which accounts for learning as participation is necessary. In this model, the learner develops skills which enable him or her to engage with contextual and cultural factors of communication. Although the two models are not mutually exclusive but in fact complementary, the latter is far more appropriate for understanding language as socialisation, as an ongoing process of engagement
The majority of humanities computingprojects within the discipline of literaturehave been conceived more as digital librariesthan monographs which utilise the medium as asite of interpretation. The impetus to conceiveelectronic research in this way comes from theunderlying philosophy of texts and textualityimplicit in SGML and its instantiation for thehumanities, the TEI, which was conceived as ``amarkup system intended for representing alreadyexisting literary texts''. This article exploresthe most common theories used to conceiveelectronic research in literature, such ashypertext theory, OCHO (Ordered Hierarchy ofContent Objects), and Jerome J. McGann's``noninformational'' forms of textuality. It alsoargues that as our understanding of electronictexts and textuality deepens, and as advancesin technology progresses, other theories, suchas Reception Theory and Versioning, may well beadapted to serve as a theoretical basis forconceiving research more akin to an electronicmonograph than a digital library.
"Things seem to decline..." - Language, ethnicity and identity illustrated by material from a former Swedish colony in Misiones, Argentina and dialect material from Bjurholm, Sweden Nowadays there is a universal tendency towards convergence and simplification in several official European languages, the existence of non-codified languages is threatened and dialectal varieties are subject to levelling. If minority languages and intralinguistic varieties are to survive, this depends to a great extent on the identity of the individual speaker and the values and attitudes attached to his/her variety and speech behaviour. Besides, it is also due to the size of the language group, sharing the same values, and its cultural activities, forming part of its tradition. For that reason I have selected material from two threatened speech communities: one from a former Swedish colony in Misiones, Argentina, the other from Bjurholm, a small dialect-speaking community in the interior of Västerbotten, Sweden, in order to study the mechanisms causing the preservation or loss of the linguistic varieties as well as the cultural boundaries. This study consists of three parts and is divided into eleven chapters. The first part consists of chapter 1-4. Initially concepts of ethnicity, identity and culture are discussed in the light of different sciences, i.e. social anthropology (Barth 1969; Hylland Eriksen 1998), ethnology, the sociology of language (Fishman 1989) and sociolinguistics (Edwards 1985). In the following chapter emigrant material (narratives, letters, local history) forms part of the historical dimension: from the cultural contacts of individuals arriving in Brazil to the later Swedish settlement in Misiones, where it is appropriate to talk about an ethnic group, its collective history and Swedishness. In chapter 3 the continuity of the cultural heritage is illustrated by onomastic material (personal names) from three generations of Swedish descendants. Chapter 4 is a report of investigations carried out in the 1990's. In 1999, the Swedish language had been maintained among 20 of the 32 informants of Swedish descent, each one representing one family network. Their identity is hyphenated: they are all Argentine but of Swedish descent, and constitute an ethnic group. 24 of them had grown up with Swedish as their first language. Language attitudes had become more positive since 1988, when a similar investigation took place. Lately a new denomination has appeared: Los Nordicos, and it is discussed whether it is ethnic or not. Part two consists of chapter 5-10 and is the result of the project Dialects in change. The principal questions in this pilot study concern dialect boundaries: are they still maintained or subject to levelling? Which dialectal items are used as boundary markers and which are substituted by standard forms? A dialect boundary implies that dialects are still used in fairly genuine forms and that a local norm is prevailing, which seems to be the case in Bjurholm. Via three types of data, including inquiries, two tests: on dialectal vocabulary (50 lexical items) and translation of twelve standard sentences into dialect versions, besides type recordings of authentic speech, the author has tried to describe the local norm, based on individual micro data. In chapter 9 criteria for dialect variables used in quantitative studies are discussed. Based on this material there is evidence for dialect boundaries towards Lappland as well as towards the neighbour parish of Vindeln (former Degerfors). This material has to be extended to serve as a base for general conclusions and methods must be refined for further investigations. This will be possible, as the old dialects are still in use among the two oldest generations of adults. If they are to be used in the future depends on the younger generation (25-^4-0 years). In the final chapter data from the bilingual Swedish speakers in Misiones and the dialect speakers of Bjurholm are summarized, discussed and compared. For linguistic survival three key concepts are important: contact, prestige and identification, which can be related to ethnicity on a group level and identity on an individual level. There are striking similarities between them. In both cases there are boundaries between "us" and "them", but these are more subtle in an intralinguistic perspective. Both categories are using spoken varieties in a transitional stage, as Misiones Swedish soon will be extinct and the genuine Bjurholm dialect subject to levelling. Both varieties are also informal and diglossie in function, although codes are not always strictly kept apart. While the use of Misiones Swedish is reduced to the family sphere, the Bjurholm dialect can be extended to a wider range of domains. The great difference seems to concern the history of the varieties and the size of the group of speakers: the Bjurholm dialect can be traced back to at least 1750, maybe even to dialect splitting in the medieval time, while Misiones Swedish has been used for about 100 years by three generations of speakers, nowadays reduced to a number of approximately 150 persons.
Looking specifically at the genre ofadaptive narrative, this article explores thefuture of literature created for and withcomputer technology, focusing primarily on thetrope of mutability as it is played out withnew media. Some of the questions askedare: What can the medium of a work ofliterature, that is its material aspect, tellus about the text? About character? What canit possibly matter if narrative is recounted onpapyrus, retold on parchment and rag, and thenremediated in pixels? Isn't it the messagecarried by the medium we are most concernedwith, stable or unstable throughout the processof inscription, reinscription, encoding anddecoding, translation and remediation? Thispaper speculates about possibilities ratherthan attempts to answer these questions, butthe structuring and mean-making componentsconsidered here stand as examples of some wemay want to think about when developing futuretheories about literature – and all types ofwriting – generated by and for electronicenvironments.
This paper describes a combined instrument (eye tracker and target generator, both head mounted, with integrated data analysis) that tests parameters of saccadic eye movement and fixation control to give insight into the status of functional brain systems. Using three minilasers, the target generator projects three visual stimuli, a fixation point and two lateral stimuli, with programmable timing. The controller allows the selection of overlap, 200-msec gap, or remembered saccade trials. Size, maximal velocity, and reaction time are determined for each primary saccade. The number of prosaccades and antisaccades are counted. More saccades—for example, the occurrence and latency of corrective saccades—may be evaluated off line by an interactive PC analysis program. The eye position data can be transferred to a PC. Off-line analysis compares each observed variable relative to an age-matched control group (300 healthy control subjects 7–70 years of age, tested in the overlap condition with prosaccade instructions and in the gap condition with antisaccades). The diagnostic results can be used to elaborate an individual optomotor training program.
Korean Combinatory Categorial Grammar (KCCG) is an extendedcombinatory categorial grammar formalism to capture thesyntax and interpretation of a relative freess word order, longdistance scrambling, and other specific characteristics of Korean.KCCG formalism can uniformly handle word order variations amongarguments and adjuncts within a clause, as well as in complexclauses and across clause boundaries, i.e. long distancescrambling. The approach we develop takes advantage of the ability of CCGfor type raising and composition along with the ability of variablecategories and unordered argument modeling for relatively freeword order treatment (Lee et al., 1994; Lee et al., 1997).We apply a probability model and heuristics using Koreancharacteristics to our KCCG parser.Results of the experiments on varioustext genre show that the KCCG parser performsat 87.67/87.03% constituent precision/recall.
REVIEWS 529 Zubova, L. V. Sovremennaia russkaia poeziia v kontekste istorni iazyka.Novoe literaturnoe obozrenie, Moscow, 2000. 432 pp. Notes. Bibliography. Index. Priceunknown. ZUBOVA'S study offers a detailed and provocative analysis of Russian poetry of the I96o-9os. It discusses almost three hundred authors, including leading postmodernist figures such as Joseph Brodsky, Viktor Krivulin, Genrikh Sapgir, Sergei Stratanovskii, Dmitrii Prigov, Elena Shvarts and Viktor Sosnora. Zubova highlights playful and innovative aspects of Russian postmodernistpoetry, arguingthat many linguisticexperimentsembedded in the texts under scrutiny in the present study explore the shortcomings and inadequacyof contemporaryRussianlanguage. Zubova arguesthat linguistic games of Russian postmodernistpoets offeralternativeways of development for phonetic, semantic and grammar structures of the Russian language. Zubova holds an optimisticview that deviations from the standardlanguage, as observed in the texts she studies, do not destroy the language, but help to preserve it, especially because they resurrect from oblivion some forgotten linguistic norms from the past (p. 399). Zubova's main thesis is based on the belief that language 'is a self-correctingsystem as well as a combination of options to express various meanings' (p. 399). The book will be of great interestto linguistsand to studentsof Russianpoetry, since it offersimportant insightsinto today'sstateof the Russianlanguage. The book comprises seven chapters, including an introduction and conclusion. Chapter one outlines the main theoretical frameworkwhich is applied throughout the book; chapter two discusses phonetic aspects of contemporary poetic experiments; chapter three investigates etymological innovations; chapter four is devoted to lexical changes; chapter five talks about archaic aspects of various grammaticaldeviations;chapter six offersa detailed analysis of the linguistic games based around gender; and chapter seven analysessyntacticalstructuresof the texts. In addition, the book offersa briefsummaryof the main conclusions (pp. 398-99), bibliographyand index. To some extent, all the texts Zubova discussesmight be viewed as hypertext, with no significantdifferentiationbetween the language of the i 960S and of the I990s. Furthermore,Zubova revealspostmodernisttendencies in Russian poetry of this period by demonstratinghow the linguisticexpressionis linked to the postmodernist worldview of the authors she discusses. Zubova encouragesher readersto considersome seeminglybad poems and treatthem with a sensitivityto the irony and parodic intentions they display.As Zubova states, 'the most constructivedevices of postmodernisttext include irony, selfirony and linguistic game' (p. i i). Zubova also highlights the authors' balancing acts between high and low cultures. In this respect, such poets as Prigov, Sosnora, Shvartsand Krivulin appear to be particularlyimaginative in their use of language, exploring precarious borders between modern and archaic forms of communication, and between elitist and popular forms of expression. Zubova's findings illustratewell the intrinsicbond between postSoviet poetry and Soviet undergroundliterature.Zubova'sstudyis a welcome additionto the extensiveanalysisof postmodernistfictionundertakenby Mark Lipovetsky (RussianPostmodernist Fiction.Dialoguewith Chaos,Armonk, NY, 530 SEER, 8o, 3, 2002 I999). Zubova is well aware of the metatextual qualities of Russian postmodernism, and points to the intertextualgames of varioustexts. In addition to the more familiarnames, Zubova introducesher audience to lesser known authors, such as Vladimir Strochkov, Ian Satunovskii, and Vladimir Erl'.Unfortunately, Zubova's studydoes not contain any biographical detailsof the authorsshe quotes. Such an appendixwould be a usefultool for assessingthe spreadof linguisticdeviations, from the point of view of age groups, regional variations and the aesthetic preferences of the poets. It is difficultto assess whether some of the deviations from established linguistic norms were intentional, or derive from the contemporary sloppy usage of Russianlanguage that isparticularlynoticeable in post-Soviet Russianmedia. Krivulinand Shvarts,for example, are philologistsby training,and therefore are more inclined to have playful appropriationof some idioms, or absurd examplesof Soviet newspeak. As Zubova's study demonstrates, numerous poetic experiments reflect on the fluid state of the Russian language itself. In this respect, Zubova's discussion of the satirical elements relating to the concept of gender in contemporary Russian poetry is particularlyrewarding. Zubova's examples from Russian poetry reveal, for example, the uncertaintiesrelatingto gender of animals. Thus, some poets use the feminine form of the noun koshka (cat) with the additional note that it is used in their poem as a noun of masculine gender. Such examples are both amusing and obscure. Zubova suggeststhat contemporary Russian poets are struggling to revive the use of the neuter gender that otherwise has been steadily disappearing from the standard language (p. 301...
^j;=5!? ITHIN the rich corpus of metrical psalms comtm itt^ta * posed during Spain's Golden Age, Fray Luis de Aivi bl T 0_Le6n's versions are universally accorded the high;@ ^ VV |@ est praise. As heir to a literary tradition that ex,.A s iGne * tended back to the late Middle Ages and early.f4Li J Renaissance, the Salamancan scholar and poet revolutionized Spain's engagement with the Psalter, establishing the lira or estrofa alirada as the dominant verse form for vernacular psalm translations (Rivers 112; Nufiez 357), and making close lexical parallelism and philological accuracy, rather than interpretive digression, the norm for most of his followers. As may be expected from the great Augustinian's role as el primer poeta humanista espaniol en lengua vulgar (A. Blecua 97), Fray Luis was widely imitated, especially among disciples of his own order, and questions of authorship and dating of the many psalm versions attributed to him continue to trouble literary historians (Nufiez 357-8; J. M. Blecua Poesia completa, 41-2). Jose Manuel Blecua, in his 1990 edition of the Poesia completa, based on all extant manuscripts, includes as genuine the following poems: Psalm 1 Beatus vir, 4 Cum invocarem, 6 ne in furore, 9 Salvum me fac, 12 Usquequo, Domine (2 versions), 17 Diligam te, 18 Coeli enarrant, 24 Ad te, Domine, levavi,
This paper presents the syntactic annotation level of a project aimed at providing a small dialog corpus with multiple levels of annotation. The syntactic annotation is based on dependency syntax. We outline the reasons for choosing dependency, and show the syntactic annotation for some constructions. We finish by describing the current state of the project. 1.
DOAJ is a unique and extensive index of diverse open access journals from around the world, driven by a growing community, committed to ensuring quality content is freely available online for everyone.
A novel eye-movement-contingent method is presented. It builds on and extends established eye-movement-contingent visual display change methods in that it uses movements of the eyes to control the presentation of acoustic information during sentence reading. In one implementation, an irrelevant spoken word is presented when the eyes cross a predetermined spatial boundary before they move on to a selected visual target word. The relationship between the spoken word and the visual target is manipulated, and the pattern of interference, caused by the presentation of the spoken word, is used to determine the nature and time course of activated representations. Results from three recently completed experiments in which the technique was used show that a word’s phonological code remains active after it has been read and that the activated code has speech-like properties.
To assess the effects of discrepancy between two independent variables, investigators sometimes compute difference scores and correlate such scores with a criterion variable. However, the correlation of the difference with the criterion is accounted for by the correlations of the difference constituents with the criterion and the constituents’ variances. It follows that when investigators are testing a prediction that is not captured by the difference constituents’ main effects, using the difference correlation analysis may be misleading. Under these circumstances, the effects of a discrepancy between two independent variables can be assessed by a test of their interaction. The problems inherent in using difference scores and the advantage of testing the interaction are illustrated in relation to research programs on two separate topics in social psychology.
You have accessThe ASHA LeaderFeature1 Nov 2002AAC, Literacy and Bilingualism Ovetta L. Harrison-Harris Ovetta L. Harrison-Harris Google Scholar https://doi.org/10.1044/leader.FTR2.07202002.4 SectionsAbout ToolsAdd to favorites ShareFacebookTwitterLinked In Children who use augmentative and alternative communication (AAC) have historically been challenged in their attainment of literacy skills. These challenges are even greater for AAC users who are bilingual. AAC users in the United States comprise large numbers of individuals from culturally and linguistically diverse backgrounds. Current demographic trends indicate that linguistic diversity will continue to intensify. During the 12 years between 1986 and 1998, the number of U. S. children who were identified as limited English proficient increased from 1.6 million to 9.9 million (see Tucker 1999). It is estimated that, by the year 2050, 40% of school-aged children in the United States will come from homes where English is not the first language. Individuals who use AAC systems surely will be represented in this group. The fact that many children in the United States, including those who use AAC systems, live amidst a sea of languages has captured national attention and has influenced our educational system. The new thrust to achieve educational equality represents a historic change. Many bilingual or monolingual schools that taught in languages other than English existed before World War II. For example, many German-only schools could be found in the northern Midwest. Afterward, a pattern of English-only instruction dominated our education system. As recognition of the cultural and linguistic diversity of the United States grew, a need to provide effective and appropriate education for bilingual children arose. Educators, parents, and researchers have challenged the notion of an English-only education for children from linguistically diverse backgrounds. Research supports the notion of education for limited-English-proficient children, including those relying on AAC systems, to be introduced in their first language, providing a transition to stronger second-language usage. This is logical given the fact that literacy attainment depends on language. Language learning, including reading and writing, is always culturally based. Reading and writing involve particular ways of using and thinking about written language that go beyond finding meaning in text and include the construction of sociocultural viewpoints or ways of understanding the world around us. It is important to realize the sociocultural and communicative nature of literacy, because of the possible therapeutic impact when working with bilingual AAC users. Writing, similarly, is a contextualized social event. It is a transactional, circular process created from a person's linguistic resources and interaction with past experiences. Viewing literacy learning as it is socially constructed through language provides a nice perspective of the need to educate linguistically and culturally diverse children from their first-language knowledge base. Research Challenges The challenges of literacy attainment for both monolingual and bilingual children who use AAC have become an area of focus for special educators, speech-language pathologists, parents, and researchers. Although research in reading and writing development of AAC users has increased steadily over the past 10 years, only recently have researchers turned their attention to reading and writing development of bilingual AAC users. Many of these children are unsuccessful in developing literacy, yet there is increasing recognition that this group is capable of developing sophisticated reading and writing skills. Bilingual AAC users who are highly successful in developing these skills make tremendous gains in overall language development and in use of their AAC systems. Acquisition of more vocabulary and the ability to compose text are just two advantages that literacy attainment brings to their receptive and expressive language development. Major focus has been brought to the topic of literacy attainment for bilingual AAC users because of its particular importance for this population. Attainment of literacy allows bilingual and monolingual AAC users, like all students, to be able to prepare messages to be used at a later time, produce exact messages, and learn vocabulary with which they can spontaneously spell out messages. But Light and McNaughton (l993) give three reasons why literacy development holds additional importance for AAC communicators. First, their face-to-face communication skills are often severely limited. Communication can be quite slow. Often the able-bodied message receiver doesn't have time to participate in communication interaction with an AAC communicator. Research shows that, in interactions between a person who is using an AAC system and a speaking person, the speaking person often dominates the interaction, and the person using the AAC system may not have opportunities to initiate topics or converse fully. Literacy gives an AAC communicator the opportunity to overcome many of the restrictions of face-to-face interaction, especially those imposed by slow AAC systems. Through writing it is possible for individuals to communicate more fully, to express themselves in more detail, and to circumvent some of the time limitations that they would normally experience in face-to-face interactions. The second aspect of school literacy importance for individuals who use AAC systems is that those who are preliterate are often limited to an ideographic literacy system. Some of these graphic systems force AAC communicators to use a closed vocabulary set and do not allow them to generate words to communicate new ideas. For example, an AAC communicator may operate a system composed of just 50 pictures or 100–200 ideographic symbols. They do not have access to the many thousands of concepts and ideas that they need in order to communicate fully and effectively. The use of orthographic literacy skills can be one way to open up access to a full range of concepts and vocabulary to students who use AAC. The literate AAC communicator, using traditional orthography, may spell words that are not printed on their communication boards or indicate first letters of words to which they don't have access on their communication system. In this way, they can use literacy skills to communicate in face-to-face interactions. The literacy development of augmentative communicators also may provide them with a means to participate in society by using written communication (as others also use written communication) to express opinions and give information. Using literacy as others do may help the bilingual AAC communicator advocate for bilingual education and acquire a sense of belonging to society as well as a stronger sense of value. The third way that literacy development carries added importance for bilingual AAC users involves vocational opportunities. In North America, there are very few individuals who use AAC systems who are competitively employed. The number holding white-collar jobs is few. The range of job opportunities available to individuals who have physical disabilities in general is restricted. AAC communicators are not usually employed in jobs requiring manual labor. Thus, they may need highly developed literacy skills for jobs involving, for example, data entry or word processing. Given limited vocational opportunities, the role of literacy in job preparation for bilingual AAC communicators is critical. Yvonne's Story AAC users must rely on innovative and sometimes creative strategies to learn to read, write, and monitor their understanding of what they are reading. Literacy-learning strategies for bilingual AAC users have not received as much attention as those of monolingual users. Some of the unique struggles and successes of literacy attainment can be seen in the story of Yvonne, a young Puerto Rican AAC user. Yvonne provides a wonderful example of the importance of first-language support and the use of specific literacy-learning strategies for bilingual AAC users. Yvonne is a 10-year-old girl with cerebral palsy of the spastic quadriplegic variety. She is nonambulatory and limited-speaking secondary to cerebral palsy. Her hearing and vision are within normal limits. During my initial contact with Yvonne, her intellectual functioning had not been formally determined. Yvonne's family immigrated to the United States one year before my initial contact with them. She is an only child. The primary language of the home is Spanish. Her father had limited English proficiency and her mother spoke no English at the time of my initial contact, although over the course of the school year they gained more proficiency. Another important characteristic of this family was the fact that the parents decided not to have any other children in order to devote total attention to Yvonne's education and health needs. Although no extended family lived in the area, they resided in a supportive neighborhood with other Puerto Ricans. Yvonne communicated primarily through use of an eye-gaze communication board. She used Mayer-Johnson Symbols and usually had a maximum of six symbols on her board. Other methods of communication included a smile/frown, yes/no response. A smile meant yes and a frown meant no. Yvonne also communicated by directing her eyes toward people or items that she wanted. Yvonne was not reading or writing very much in English when we first met. She may have recognized some English words that she encountered daily such as the names of her school, teacher, and classmates, and she had limited environmental vocabulary. I was not sure of her exact reading proficiency in Spanish; however, she did not demonstrate the ability to independently read upper-elementary-graded text w ritten in Spanish and answer basic content questions. Her listening comprehension for stories read to her in Spanish was good. We were not able to assess written language use because the classroom lacked the technology for text composition. Yvonne had a strong desire to learn to read more proficiently. Yvonne was a student in a general elementary school located in western Massachusetts. Her classroom was nongraded, but the students, all classified as special needs, were of comparable ages to those of fourth grade. The room was self-contained and designated by the school system as a special education classroom. The special need categories included physically and cognitively impaired. Half of the class comprised other Puerto Rican children. My role was that of AAC literacy consultant, but I also brought my expertise in the area of multiculturalism in speech-language pathology. My initial meeting with Yvonne occurred early in the school year, in her classroom with the classroom teacher and instructional aide. Yvonne immediately greeted me with a welcoming smile because she appeared to know that I was there especially to help her learn. During my initial meeting I was able to informally assess that Yvonne had good cognitive skills. She used her voice to initiate communication to bring attention to matters of need or interest. She laughed appropriately at jokes, her eyes followed speakers in a conversation, and she spontaneously used her eyes to appropriately answer yes/no questions. All of the conversations around her and directed to her by her teacher were in English. Yvonne obviously acquired some English proficiency, although she may not have understood everything. I had formal training in Spanish and worked some years earlier in a predominately Mexican-American school district in Southern California where I used the language daily. Although I lacked confidence in my use of Spanish, I greeted Yvonne and introduced myself in Spanish. Approaching her using Spanish set a tone for Yvonne that I was supportive of her background and language usage. She recognized that I needed help using the dialect of Spanish that she was familiar with as a primary way of communicating with her. We learned quickly to work together around the use of a language system. Honoring her first language was important to our working together. Another important factor was Yvonne's desire and willingness to learn English, which contributed significantly to her rapid acquisition of stronger English proficiency. On my second day of visiting the classroom, I was extremely pleased to meet the school SLP assigned to Yvonne. This wonderfully competent, energetic clinician just happened to be bilingual in English and Spanish. With a bilingual SLP and my knowledge of literacy-learning techniques for AAC users, Yvonne blossomed over the course of that academic year in her English proficiency and particularly in her ability to read and spell. A Successful Technique I first introduced a spelling/word-level reading technique to Yvonne that proved to be highly successful and allowed her to gain 10–12 new words in reading recognition and spelling each week. Upper-elementary-aged bilingual AAC users with profiles similar to Yvonne should start with whole-word-level reading aimed at teaching recognition of entire words such as swim, pool, the, or cap. Instruction of whole words leads to success in reading phrases and simple sentences quickly. Phonetic instruction should occur as well. The Words on the Wall technique, which can be used with monolingual as well as bilingual AAC users, begins by the teacher selecting approximately 3–5 new words that the student needs to learn. These should be words relevant to familiar situations and not spelling words from a spelling book. For example, Yvonne went swimming each week in school and thus, during her first week, she learned the words swimming, towel, pool, water, and splash. These words were initially introduced in Spanish only. The next step in this technique is to make the word accessible by writing it in large print on a sentence strip and attaching it to the wall. The word may initially be paired with a symbol, with the symbol being phased out over time leaving just the written word. The student and the teacher define the word and talk about events involving the target word. After all of the target words are discussed and displayed on the wall, the teacher asks the student to identify each word one at a time as in a spelling test. Yvonne used eye gaze to identify her target words. During the next day or week, depending on how well the student masters each set of words, introduce more words (1–3 a day). Leave all words on the wall for the school year, increasing the number of words each week. Review old and new words. After enough words are mastered, have the student begin to read simple sentences. Introduce words such as a and the to allow formation of sentences. The school SLP delivered all of the training to Yvonne in Spanish first and followed it with English only after she knew that Yvonne understood the word in Spanish. Because this literacy-learning technique is based primarily at the word level, it is easier to transition from the Spanish to the English word. The school SLP also kept in close contact with Yvonne's parents, phoning them and sending home each week the word that Yvonne was working on. Yvonne's communication reflected her increased vocabulary. A board in Spanish was sent home and used with her parents and an English board was used at school initially. As Yvonne's parents gained more English proficiency, they requested to have the English communication board as well. During the school year we piloted different types of high-tech AAC devices and switches with Yvonne. We also explored technology for writing purposes during this year. Assessment Words on the Wall lends itself to a Maze Reading Assessment technique once a student has acquired reading of simple sentences. This technique involves the deletion of target words in a sentence leaving a blank space. The student should be provided with three alternative words in random order at each blank (correct choice, incorrect choice of the same part of speech, incorrect choice of a different part of speech). For example: The boy ate a ______ (truck, this, banana). This technique can be used easily with many AAC users. Yvonne's eye gazed to her chosen word using this technique. The scale of reading proficiency most often used for informal reading assessments such as this is 90% accuracy indicating that the student is reading at an independent level, 60%–80% accuracy relating to a level where more instruction is needed, and below 60% is equivalent to a frustration level. For Yvonne, the Maze technique was delivered in English because she already had mastered the words on the wall and read simple sentences in English. The Words on the Wall technique and a Maze Reading Assessment Technique are two techniques that can be culturally and linguistically sensitive and used well with AAC users. Voice output is not required for these techniques, and the words are derived from the students' existing linguistic bases or contextual experiences. Other techniques also can be used to facilitate literacy development with bilingual AAC users. Techniques that contextualize instruction in the experiences of the home and first language are desirable. For young bilingual literacy-language learners, it is important to use interactive learning techniques that involve the teacher, peers, and the AAC user. Techniques that allow students to demonstrate competence in using language and literacy throughout the school day in all instructional activities are greatly beneficial. Techniques that use narratives such as storytelling, listening to stories, or writing are good for content development. These narratives should be delivered in the language that will allow the child to gain academic skill while learning English. My first year with Yvonne was a successful one. She gained approximately 10 new words a week over the course of the school year. For AAC users similar to Yvonne in age and cognitive ability, this is an expected rate of growth. There is no typical rate of growth for all AAC users because this population is so diverse in skill and ability. The Next Year I returned to visit Yvonne the next year when she had been promoted to a new class and school. The successful learning environment that she had previously experienced had come to an abrupt end. There was a lack of continuity with her education from the previous year. Yvonne was in a new school with all new staff. There was no Spanish language support. The literacy-learning methods had been abandoned. Communication with Yvonne was a problem. There was limited communication between the school and home. I spent the first few days in Yvonne's classroom as a participant observer and quickly assessed the social and literacy-learning needs of everyone involved in Yvonne's schooling. The goals of my intervention with Yvonne during this second school year included elimination of the communication problem between the school and the family and establishment of better trust and communication, reestablishment of appropriate instructional methods, eliminating AAC barriers, and supporting cultural identity through literacy lessons/interactions and development of a more efficient communication system. The lack of Spanish support and having to demonstrate and convince the new teachers of Yvonne's literacy-learning capabilities resulted in lost time in her development. Strong first-language support and knowledge of specific literacy-learning techniques for bilingual AAC users led to a successful outcome for Yvonne. She enjoys reading and had a strong desire to continue reading and learning English. This was compatible to the wishes of her parents. Like Yvonne, not all bilingual AAC users have significant difficulties learning to read and write; however, many of them do. Therefore, it becomes important to communication disorders specialists to identify variables of language that are predictive of later reading difficulties. Researchers and other professionals from different fields of study are combining their interests to close the knowledge/information gap that exists between what is already known about bilingual AAC users' acquisition and development and the information needed to help develop intervention strategies for successful written language. Strategies for Monolingual Clinicians: A Postscript Although I did have formal training in Spanish in high school and college and had worked in a predominately Spanish-speaking community in Southern California, I still lacked confidence to converse with Yvonne in Spanish when I first met her. I knew that there were many dialects of Spanish, and I initially did not know enough about the Spanish that she and her family used. Clinicians who are monolingual or who lack information about a second-language-speaking student must do the research to find linguistic information particular to that student. Such knowledge is also helpful in understanding the contexualized uses of literacy in the home that will complement those used in the classroom. General professional development in the area of bilingual literacy learning is highly recommended, as is professional development in AAC. Understanding policies in educating bilingual students that are implemented in your school district is important. Clinicians should understand how policy affects access to instruction for bilingual students. Social, cultural, and economic issues that affect student learning and instruction also should be well understood. It is helpful to gain information from parents, other teachers, and community members about ways that they find helpful in instructing bilingual AAC users. Ovetta Harrison-Harris is chair of the department of communication sciences and disorders at Howard University. She is project director for a U.S. Department of Education Office of Special Education and Rehabilitative Services-funded graduate training program in AAC with an emphasis in multiculturalism and literacy development. For More Information Light J., Binger C., & Smith A.K. (1994). Story reading interactions between pre-schoolers who use AAC and their mothers. Augmentative and Alternative Communication, 10, 225–268. CrossrefGoogle Scholar Light J., & McNaughton D. (1993). Literacy and Augmentative and Alternative Communication (AAC): Expectations and Priorities of Parent and Teachers. Topics and Language Disorders, 13(2), 33–46. CrossrefGoogle Scholar Light J., & Smith A.K. (1993). Home literacy experience of pre-schoolers who use augmentative communication systems and their non disabled peers. Augmentative and Alternative Communication, 9, 10–25. CrossrefGoogle Scholar Pearson B.Z., Fernandez S., & Oller D.K. (1993a). Lexical developmental in simultaneous bilingual infants: Comparison to monolinguals. Language Learning, 43, 93–120. CrossrefGoogle Scholar Pearson B.Z., Fernandez S.C., & Oller D.K. (1993b). Lexical development in bilingual infants and toddlers: Comparison to monolingual norms. Language Learning, 43(1), 93–120. CrossrefGoogle Scholar Pearson B.Z., Fernandez S., & Oller D.K. (1995). Cross-language synonyms in the lexicons of bilingual infants: One language or two?, Journal of Child Language, 22, 345–68. CrossrefGoogle Scholar Pearson B.Z., Oller D.K., Umbel V.M., & Fernandez M.C. (1996, October). The Relationship of Lexical Knowledge to Measures of Literacy and Narrative Discourse in Monolingual and Bilingual Children. Paper presented at the Second Language Research Forum, Tucson, Google Scholar Tucker A perspective on and bilingual education Google Scholar Ovetta L. is chair of the department of communication sciences and disorders at Howard University. She is project director for a U.S. Department of Education Office of Special Education and Rehabilitative Services-funded graduate training program in AAC with an emphasis in multiculturalism and literacy development. of the ASHA Special Augmentative and Alternative Communication, and a for With Communication to your in Nov &
Recent studies have suggested a theoretical distinction between active elaboration and passive storage in visuospatial working memory, but research with older adults has failed to demonstrate a differential preservation of these two abilities. The results are controversial, and the investigation of the active component has been inhibited by the absence of any appropriate experimental procedures. A new task was developed involving the mental reconstruction of pictures of objects from fragmented pieces, and this provides a useful procedure for exploring active visuospatial processing. Significant differences in terms of both correctness and response latency were obtained between young and older adults and between younger old and older old adults. Performance also varied with visual complexity, mental rotation, and processing load. It is concluded that this ecologically relevant procedure constitutes a very powerful, sensitive, and reliable tool for identifying individual differences in visuospatial working memory.
Users need more sophisticatedtools to handle the growing numberof image-based documents availablein databases. In this paper, wepresent a system devoted to theediting and browsing of complexliterary hypermedia includingoriginal manuscript documents andother handwritten sources. Editingcapabilities allow the user totranscribe manuscript images in aninteractive way and to encode theresulting textual representationby means of a logical markuplanguage (based on the XML/TEIspecification). Bothrepresentations (image andstructured text) are tightlylinked to facilitate the readingand the interpretation ofdocuments. This text/imagecoupling scheme is an attempt tounify several layers ofinformation in order to providethe user with a global vision ofthe work. Our system also suppliestools capable of processing andrelating information stored bothin images and structured texts.Finally, application-specificvisualization techniques have beendeveloped in order to provideusers with a way to identifyrelationships between sourcedocuments and help them tonavigate.
This thesis investigates the perceptual categories associated with contrasting pitch accents in Tokyo Japanese. Various aspects of pitch movements throughout the words are systematically varied in order to determine what aspects of the pitch contour affect categorization of words on the basis of accent. In addition, this thesis also investigates the effect on categorization of loss of pitch due to lack of voicing in various parts of the word. The model determined by these studies reveals two more-or-less orthogonal perceptual dimensions; pitch alignment with speech segments determines accent location and the amount of pitch drop determines accent presence. This study also investigates how to quantify the distinctive function of pitch accent in a way which incorporates the frequency of the contrasting items, as well as the peculiar category structure of accents. This model was applied in the analysis of a large-scale lexical database, revealing many irregularities in the distribution and use of accents. Comparing this quantification of the lexical use of accent with the perceptual experiments shows that accent-location detection is functionally more fundamental than accent-presence detection in short, 2-mora words. (PsycINFO Database Record (c) 2016 APA, all rights reserved)
Dictionaries can be used as a basis for lexicon development for NLP applications. However, it often takes a lot of pre-processing before they are usable. In the last 5 years a product-independent database of formal word features has been developed on the basis of the Van Dale dictionaries for Dutch. The database has proven to be useful in various NLP applications. This paper describes the history, some advantages and the constraints in the development of this database.
In recent years translation has come to be investigated increasingly from critical perspectives, with various studies high-lighting the translator’s mediating involvement in the construction of particular discourses. This article describes an exploratory think-aloud protocol study which attempts to observe this mediating involvement in the translation process, using degree-level language students as subjects. The study suggests that the students may not be conscious of the fact that certain lexical choices in their translations conform to common stereotypes of Britain. The relatively low level of critical discourse awareness among the students may in some ways be related to translation practice in their educational contexts, which focuses on norm compliance and may do little in most cases to encourage reflection on the effects of translational choices on readers’ perceptions of what is communicated.
In the field of empirical natural language processing, researchers constantly deal with large amounts of marked-up data; whether the markup is done by the researcher or someone else, human nature dictates that it will have errors in it. This paper will more fully characterise the problem and discuss whether and when (and how) to correct the errors. The discussion is illustrated with specific examples involving function tagging in the Penn treebank.
Ambiguity resolution in the parsing of natural language requires a vast repository of knowledge to guide disambiguation. An effective approach to this problem is to use machine learning algorithms to acquire the needed knowledge and to extract generalizations about disambiguation decisions. Such parsing methods require a corpus-based approach with a collection of correct parses compiled by human experts. Current statistical parsing models suffer from sparse data problems, and experiments have indicated that more labeled data will improve performance. In this dissertation, we explore methods that attempt to combine human supervision with machine learning algorithms to try and extend accuracy beyond what is possible with the use of limited amounts of labeled data. In each case we do this by exposing a machine learning algorithm to unlabeled data in addition to the existing labeled data. Most recent research in parsing has shown the advantage of having a lexicalized model, where the word relationships mediate knowledge about disambiguation decisions. We use Lexicalized Tree Adjoining Grammars (TAGs) as the basis of our machine learning algorithm since they arise naturally from the lexicalization of Context Free Grammars (CFGs). We show in this dissertation that probability measures applied to TAGs retain the simplicity of probabilistic CFGs along with its elegant formal properties and that while PCFGs need additional independence assumptions to be useful in statistical parsing, no such changes need to be made to probabilistic TAGs. The main results presented in this dissertation are: (1) We extend the Co-Training algorithm (Yarowsky 1995; Blum and Mitchell 1998), a machine learning technique for combining labeled and unlabeled data previously used with classifiers with 2/3 labels to the more complex problem of statistical parsing. Using empirical results based on parsing the Wall Street Journal corpus we show that training a statistical parser on the combined labeled and unlabeled data strongly outperforms training only on the labeled data. (2) We present a machine learning algorithm that can be used to discover previously unknown subcategorization frames. The algorithm can then be used to label dependents of a verb in a treebank as either arguments or adjuncts. We use this algorithm to augment the Czech Dependency Treebank with argument/adjunct information. (3) We extend a supervised classifier for automatically identifying verb alternation classes for a set of verbs so that it can be used on minimally annotated data. Previous work (Merlo and Stevenson 2001) provided a classifier for this task that used automatically parsed text. With the use of learning of subcategorization frames we construct the same type of classifier which now requires text annotated with part-of-speech tags and phrasal chunks. In each of these results we use some existing linguistic resource that has been annotated by humans and add some further significant linguistic annotation by applying statistical machine learning algorithms.