Papers reviewed and determined not to be word norm studies. Use the flag icon to report errors or suggest re-inclusion.
18265 papers
The present paper outlines an ongoing project of annotation of the extended nominal coreference and the bridging anaphora in the Prague Dependency Treebank. We describe the annotation scheme with respect to the linguistic classification of coreferential and bridging relations and focus also on details of the annotation process from the technical point of view. We present methods of helping the annotators -- by a pre-annotation and by several useful features implemented in the annotation tool. Our method of the inter-annotator agreement is focused on the improvement of the annotation guidelines; we present results of three subsequent measurements of the agreement.
The aim of the project entitled "Computerized Historical Linguistic Database of Latin Inscriptions of the Imperial Age" (http://lldb.elte.hu/) is to develop and digitally publish a fundamental computerized historical linguistic database that incorporates and treats the Vulgar Latin material of the Latin inscriptions from a specific group of the European provinces of the Roman Empire in the first phase. This will, on the one hand, allow for a more thorough study of the regional changes and the diversity of the Latin language of the Imperial Age. On the other hand, it could also serve as a basis for subsequent international co-operation, in the course of which further work on the computerized historical linguistic database may be executed. This paper intends to present the past and the present, as well as the future possibilities of this Database.
Generative lexicalized parsing models, which are the mainstay for probabilistic parsing of English, do not perform as well when applied to languages with different language-specific properties such as free(r) word order or rich morphology. For German and other non-English languages, linguistically motivated complex treebank transformations have been shown to improve performance within the framework of PCFG parsing, while generative lexicalized models do not seem to be as easily adaptable to these languages.
This paper presents a basic analysis of syntactic annotation errors and inconsistencies in the Prague Dependency Treebank, the biggest corpus of Czech with manual syntactic annotation. The corpus is used for developing and testing of many syntactic analysers of Czech and the problems in the annotation have an essential impact on the evaluation of the quality of these parsers and the results of precision measurements. We identify some of the basic annotation problems and in some cases, we outline possible solutions.
The lexical development system OntoNet is introduced, which includes a browser and an editor for the WordNet 3.0 database. The aim of the OntoNet project is to provide a comfortable and up-to-date access to the lexical database for the modification of WordNet or the development of new wordnets.
We describe the Hindi Discourse Relation Bank project, aimed at developing a large corpus annotated with discourse relations. We adopt the lexically grounded approach of the Penn Discourse Treebank, and describe our classification of Hindi discourse connectives, our modifications to the sense classification of discourse relations, and some crosslinguistic comparisons based on some initial annotations carried out so far.
The paper describes an auditory experiment aimed at testing whether the intrinsic loudness of a stimulus with a given voice quality influences the way in which it signals affect. Synthesised voice quality stimuli in which intrinsic loudness \nwas systematically manipulated were presented to listeners to test the effect of this manipulation on the affective colouring \nof the stimuli. The results showed that even when devoid of intrinsic loudness variation, non-modal voice quality stimuli \nwere capable of communicating affect. However, changing the loudness of a non-modal voice quality stimulus towards its \nintrinsic loudness resulted in the increase of affective ratings.
This paper presents a simple and effective approach to improve dependency parsing by using subtrees from auto-parsed data. First, we use a baseline parser to parse large-scale unannotated data. Then we extract subtrees from dependency parse trees in the auto-parsed data. Finally, we construct new subtree-based features for parsing algorithms. To demonstrate the effectiveness of our proposed approach, we present the experimental results on the English Penn Treebank and the Chinese Penn Treebank. These results show that our approach significantly outperforms baseline systems. And, it achieves the best accuracy for the Chinese data and an accuracy which is competitive with the best known systems for the English data.
It is unlikely that Standard Afrikaans has been based on one relatively uniform vernacular. Ana Deumert has convincingly argued that what we recognise as Standard Afrikaans today is a construction to be attributed to language entrepreneurs who strove for a unique South African identity towards the end of the nineteenth and early 20th century. This led to deliberately discarding some of the then metropolitan Dutch linguistic norms. The Afrikaans negative and diminutive systems will be shown to be the linguistic outcome of these conceptions of identity and purity.
The basic concept of semantic Web,ontology and semantic annotation are described.Then the semantic annotation technology and tool today are introduced and analyzed,and a way of automatic semantic annotation based on HTML documents that contain rich semantic data on the Web is presented.This method couples structural analysis of documents with semantic analysis incorporating domain ontologies and lexical database Hownet,discovers the semantic partition tree corresponding to documents,and annotates HTML documents with semantic lables.The experiment is based on the HTML documents of electronic products,the result shows the method is feasible.
The present study examined the degree to which acceptance, reappraisal, or suppression based strategies are associated with changes in heart rate, eyeblink startle magnitude, Event-Related Potentials (ERPs), and self-reports of subjective experience in a sample of college undergraduates. Participants were randomly assigned to use one of these strategies during an associative learning task that contained stimuli that signaled either threat or safety from a noxious stimulus as well as during exposure to highly arousing pleasant and unpleasant images. Participants in the reappraisal and suppression groups displayed greater eyeblink startle magnitudes during the emotion induction procedures compared with participants in the acceptance and control groups. No group differences were found with respect to heart rate or ERPs in response to the emotion inductions. Compared with participants assigned to the acceptance and control conditions, participants assigned to the reappraisal and suppression conditions rated unpleasant images as being less unpleasant; however, the groups did not differ in arousal ratings. Participants did not differ in their ratings of discomfort during the associative learning task, nor did they differ in their valence and arousal ratings for pleasant pictures. Findings suggest a possible dissociation of cognition and physiological reactivity for participants using reappraisal and suppression strategies to regulate mood and affect.
The PADT project might be summarized as an open-ended activity of the Center for Computational Linguistics, the Institute of Formal and Applied Linguistics, and the Institute of Comparative Linguistics, Charles University in Prague, resting in multi-level annotation of Arabic language resources in the light of the theory of Functional Generative Description (Sgall et al., 1986; Hajičová and Sgall, 2003).
Riflessioni sulle implicazioni didattiche e di ricerca degli strumenti di analisi della linguistica dei corpora con particolare riguardo alla linguistica computazionale (Treebank) per l'insegnamento dell'italiano come L2.
This paper describes and compares two algorithms that take as input a shared PCFG parse forest and produce shared forests that contain exactly the n most likely trees of the initial forest. Such forests are suitable for subsequent processing, such as (some types of) reranking or LFG f-structure computation, that can be performed ontop of a shared forest, but that may have a high (e.g., exponential) complexity w.r.t. the number of trees contained in the forest. We evaluate the performances of both algorithms on real-scale NLP forests generated with a PCFG extracted from the Penn Treebank.
Max Planck Institute for Psycholinguistics, Nijmegen, The Netherlands We present a coding system combined with an annotation tool for the analysis of gestural behavior. The NEUROGES coding system consists of three modules that progress from gesture kinetics to gesture function. Grounded on empirical neuropsychological and psychological studies, the theoretical assumption behind NEUROGES is that its main kinetic and functional movement categories are differentially associated with specific cognitive, emotional, and interactive functions. ELAN is a free, multimodal annotation tool for digital audio and video media. It supports multileveled transcription and complies with such standards as XML and Unicode. ELAN allows gesture categories to be stored with associated vocabularies that are reusable by means of template files. The combination of the NEUROGES coding system and the annotation tool ELAN creates an effective tool for empirical research on gestural behavior.
OXlearn is a free, platform-independent MATLAB toolbox in which standard connectionist neural network models can be set up, run, and analyzed by means of a user-friendly graphical interface. Due to its seamless integration with the MATLAB programming environment, the inner workings of the simulation tool can be easily inspected and/or extended using native MATLAB commands or components. This combination of usability, transparency, and extendability makes OXlearn an efficient tool for the implementation of basic research projects or the prototyping of more complex research endeavors, as well as for teaching. Both the MATLAB toolbox and a compiled version that does not require access to MATLAB can be downloaded from http://psych.brookes.ac.uk/oxlearn/.
ABSTRACT. In poetic language, ordinary language is subject to poetic organization. This organization results in deviation from ordinary-language linguistic norms. A number of Optimality-Theoretic studies analyze poetically motivated linguistic deviation as the domination of linguistic constraints by prosodic constraints (Rice 1997, Golston 1998, Reindl and Franks 2001, Michael 2003, Fitzgerald 2003, 2007). Adding to this line of scholarship, this paper examines how metrical mapping, metrical grouping and rhyme patterning govern stress shift, syllabic variation, and syntactic inversion as exemplified in the lyrics of honky tonk country music singer Hank Williams, Sr. The violation of the norm of the standard, its systematic violation, is what makes possible the poetic utilization of language; without this possibility there would be no poetry. (Mukarovsky 1970: 43) 1. INTRODUCTION. Poetic language necessarily deviates from ordinary language, violating ordinary language norms in order to satisfy poetic patterning. Metrical organization has been shown to govern word order (Youmans 1983, 1989; Golston 1998; Fitzgerald 2003), allomorphy (Youmans 1989), reduplication (Fitzgerald 1998), lexical stress (Janda and Morgan 1988), and the deletion and insertion of syllables (Fitzgerald 1998, Reindl and Franks 2001, Michael 2003). A number of articles analyze such poetically motivated linguistic deviation as the domination of prosodic constraints over other areas of the grammar (Rice 1997, Golston 1998, Reindl and Franks 2001, Michael 2003, and Fitzgerald 2003, 2007). Adding to this line of scholarship, this paper examines how meter, metrical grouping, and rhyme govern linguistic deviation in the lyrics of Hank Williams, Sr. Specifically, I show how metrical mapping and grouping constraints drive stress shift and syllabic variation, and how constraints requiring systematic rhyme govern syntactic inversion. Following a discussion of the methods used in this study is an introduction to the poetic grammar of the Hank Williams song. Three major types of poetic organization are identified: meter, metrical grouping, and rhyme. In the second half of the article, an analysis of the Hank Williams Corpus highlights how three types of linguistic deviation reflect the interaction of poetic constraints and ordinary language constraints: stress shift, syllabic variation, and syntactic inversion. 2. METHODS. The data for the Hank Williams Corpus consist of songs that Williams performed or recorded which were collected on the ten-compact disc compilation album, The Complete Hank Williams. The album contains 224 tracks including songs, recitations, and speech. Of the 164 discrete songs among them, Williams wrote or co-wrote 98 himself. The remaining songs were written by a number of different artists, among them Fred Rose, Mel Foree, Ernest Tubb, and Leon Payne. Each song was coded for its metrical and rhyming structure, and this information was then entered into corresponding databases to facilitate analysis. 3. POETIC ORGANIZATION IN HW CORPUS. Musical rhythm has two major components: meter and grouping (Lerdahl and Jackendoff 1983). In Williams' lyrics, musical meter is realized in the linguistic text in the distribution of downbeat- and non-downbeat-stressed syllables, linguistically-empty downbeats, and extended syllables on the metrical grid. Patterns in the distribution of these units half-line- and line-finally reflect the metrical grouping structure of the song. Systematic end-rhyme reinforces metrical grouping in its patterning. 3.1 METER. The musical meter for the majority of songs in the corpus, i.e. 115 of 164, is duple, 2/2, with two half-note beats per measure. The remaining 49 songs are in triple meter, 3/4, with three quarter-note beats per measure. In each song, the downbeat, i.e. the first beat of each measure, is acoustically prominent, often realized instrumentally as the thumping bass guitar. …
Currently, there is no international standard for the assessment of fitness to drive for cognitively or physically impaired persons. A computerized battery of driving-related sensory-motor and cognitive tests (SMCTests) has been developed, comprising tests of visuoperception, visuomotor ability, complex attention, visual search, decision making, impulse control, planning, and divided attention. Construct validity analysis was conducted in 60 normal, healthy subjects and showed that, overall, the novel cognitive tests assessed cognitive functions similar to a set of standard neuropsychological tests. The novel tests were found to have greater perceived face validity for predicting on-road driving ability than was found in the equivalent standard tests. Test—retest stability and reliability of SMCTests measures, as well as correlations between SMCTests and on-road driving, were determined in a subset of 12 subjects. The majority of test measures were stable and reliable across two sessions, and significant correlations were found between on-road driving scores and measures from ballistic movement, footbrake reaction, hand-control reaction, and complex attention. The substantial face validity, construct validity, stability, and reliability of SMCTests, together with the battery’s level of correlation with on-road driving in normal subjects, strengthen our confidence in the ability of SMCTests to detect and identify sensory-motor and cognitive deficits related to unsafe driving and increased risk of accidents.
In the article we compare the role of the dictionary and the lexical database, and address the issue of language register and correctness in dictionaries. We then deal with various types of sense distribution in dictionaries, the history of the word, and the principles of selection of dictionary headwords. We cite the corpus as an essential source for the treatment of meaning, collocation and syntagmatics, and investigate ways of interpreting corpus data – corpus profiling of headwords. We conclude with the thought that a dictionary represents the central language standard, whereby all of the expressed linguistic opinions contained in it must be based on corpus evidence.
espanolComo es bien sabido, aunque para los hablantes de una lengua las variedades dialectales resulten mas evidentes en los planos lexico, fonetico o fonologico, ellas se advierten en todos los niveles del lenguaje, orbita de la que, por supuesto, no escapa la sintaxis. Asi, en el caso particular del espanol de Buenos Aires, el uso del Preterito Perfecto Compuesto del Modo Indicativo difiere sensiblemente de la norma castellana, a la vez que la conciencia de los hablantes de la lengua respecto de el es practicamente nula: o lo niegan por completo, alegando que prefieren siempre el Preterito Perfecto Simple, o bien aducen que lo emplean segun la norma de Madrid; lo cual, como se vera a lo largo de nuestro trabajo, no resulta de ese modo en ninguno de los dos casos. Asi pues, intentaremos problematizar las cuestiones de norma y uso, en relacion con la conciencia de los hablantes portenos respecto de su empleo de los tiempos pasados. Para ello, partiremos de un trabajo de campo que hemos realizado y que nos ha permitido esbozar algunos matices caracteristicos del uso del tiempo verbal que nos ocupa, es decir, el Preterito Perfecto Compuesto del Modo Indicativo del dialecto rioplatense. EnglishIt is well known that dialectal language variations appear at every level of language including syntax. However, speakers are usually aware of lexical, phonetics, and phonological variations only. In this particular case, as expected, the use of perfect tenses in Buenos Aires (Argentina) is very different from that of Madrid (Spain). The problem is that most Argentinean speakers know how to use the Present Perfect according to Spanish rules they have learned in school, but their speech do not matches their learning. Most Argentinean speakers would say (and they believe) that they do not use the Present Perfect in everyday life, when they actually do, albeit in a different way. That is why I conducted a survey among speakers of all kind of age, in order to distinguish some specific characteristics of the Present Perfect use in rioplatense dialect. Finally, I intend to discuss the concept of language norm and use related to speakers' awareness in Buenos Aires.
Business English is characterized by a specialized vocabulary,polysemy,stylistic norms of formal,nicety,preciseness,concision and emerging new words.It can improve the study effectiveness to paying attention to the accumulation of professional knowledge,to understand vocabulary through context clues,and to learn new words by chunk approach and concern about the latest business information.
This paper presents a comparative study of Judgment and Assessing frames in English and Portuguese. The aim is to verify the possibility of using the FrameNet frames to construct a lexical database for Brazilian Portuguese. The research corpus is composed by 50 legal documents, totalizing 1.055,535 tokens and 39,108 types. Through a contrastive method the Judgment and Assessing frames were selected and translation equivalents for the English lexical units were established. The points considered in this research were the polysemy and the semantic relations of words. The polysemy is the main difficulty in applying FrameNet frames for Portuguese description.
grammatical resources from treebanks for English Here, we extend the LFG grammar acquisition approach to Arabic and the Penn Arabic Treebank (ATB) (Maamouri and Bies, 2004), adapting and extending the methodology of Arabic is challenging because of its morphological richness and syntactic complexity. Currently 98% of ATB trees (without FRAG and X) produce a covering and connected f-structure. We conduct a qualitative evaluation of our annotation against a gold standard and achieve an f-score of 95%.
Abstract In the preceding chapter we have delineated the boundaries of our inquiry by defining a cross-linguistically applicable domain of predicative (alienable) possession. On the basis of this definition we can now proceed to build a data base, which comprises the relevant linguistic material from the languages in the sample. Once this task has been completed (and we will assume here that it has) our next step is to construct a typology of predicative possession, on the basis of observable similarities and differences among the constructions included in the cross-linguistic database.
Evaluation of quantitative parameters in the English language. The article touches upon different lexical and grammatical means of expression of quantitative parameters in the modern English language and means of representation of evaluative concepts 'much' and 'little'. Specific features of the processes of evaluative conceptualization and evaluative categorization are under study, as well as the notion of quantitative standard/norm as the reflection of collective shared knowledge.
You have accessThe ASHA LeaderFeature1 Oct 2009Bilingualism: Consequences for Language, Cognition, Development, and the Brain Viorica Marian, PhD Yasmeen Faroqi-Shah, PhD, CCC-SLP Margarita Kaushanskaya, PhD Henrike K. Blumenfeld andPhD Li ShengPhD Viorica Marian Google Scholar More articles by this author, PhD, Yasmeen Faroqi-Shah Google Scholar More articles by this author, PhD, CCC-SLP, Margarita Kaushanskaya Google Scholar More articles by this author, PhD, Henrike K. Blumenfeld Google Scholar More articles by this author, PhD and Li Sheng Google Scholar More articles by this author, PhD https://doi.org/10.1044/leader.FTR2.14132009.10 SectionsAbout ToolsAdd to favorites ShareFacebookTwitterLinked In Every year, thousands of middle- and upper-class American children study a foreign language for enrichment. These children, their parents, and their teachers are guided by the belief that knowing another language “is good for you.” At the same time (and sometimes in the same schools) thousands of other children—usually from immigrant and lower-class backgrounds—are discouraged from and sometimes forbidden to speak their native language. Their families are told that communication in their native languages will prevent them from mastering English and that raising children with more than one language will “confuse” them and have long-lasting, detrimental effects. Given these two contradictory perspectives, what does research say about the consequences of bilingualism? Cognitive Development Empirical evidence suggests that bilingualism in children is associated with increased meta-cognitive skills and superior divergent thinking ability (a type of cognitive flexibility), as well as with better performance on some perceptual tasks (such as recognizing a perceptual object “embedded” in a visual background) and classification tasks (for reviews, see Bialystok, 2001; Cummins, 1976; Diaz, 1983, 1985). Other studies report that bilingualism has a negative impact on language development and is associated with delays in lexical acquisition (e.g., Pearson, Fernandez, & Oller, 1993; Umbel & Oller, 1995) and a smaller vocabulary than that of monolingual children (Verhallen & Schoonen, 1993; Vermeer, 1992). Bilingual children score on par with their monolingual counterparts on tests of verbal ability by middle school, and well-controlled studies provide no evidence for lower intellectual abilities of bilingual children compared to monolinguals (Baker & Jones, 1998; Cook, 1997; Hakuta, 1986). The early differences in linguistic performance of bilingual children can be attributed to a somewhat different language development pattern. Bilingual children learn earlier than their monolingual counterparts that objects and their names are not the same and that one object can have more than one name. Understanding that language is a symbolic reference system is advantageous for metacognitive development; it does not, however, necessarily translate to improved performance on early vocabulary development tests. Those vocabulary test results are due, in part, to the way language assessment usually takes place. If a monolingual child has three lexical labels for three semantic items (“milk,” “grandma,” and “dog”), and a bilingual child has two lexical labels in English (“milk” and “grandma”) and two in Spanish (“leche” and “abuela,” the Spanish words for “milk” and “grandma”), the monolingual child’s vocabulary will be counted as three words and the bilingual child’s vocabulary will be counted as two words—because vocabulary size is counted not as the number of lexical items known, but as the number of conceptual representations that have lexical labels. Therefore, even though the bilingual child has four words, they map onto two conceptual representations, compared with the three conceptual representations of the monolingual child. This assessment technique frequently places bilingual children at a disadvantage. Bilinguals often are assessed in only one language, providing an inaccurate assessment of the child’s actual level of linguistic and cognitive development. A child assessed in only one language, typically that of the country in which he or she is being tested (i.e., English in the United States, often the second and less-proficient language), may be placed erroneously at a lower level of cognitive development than his or her true level. This placement can have adverse academic consequences, such as inappropriate lower-grade placement, being held back a year, enrollment in inappropriate remedial programs, and other placement decisions. (For more comprehensive discussions of first/second language knowledge and cognitive processing in bilingual children, see work by Cummins). Comparisons of children’s performance in the first and second language indicate that performance in one language, even the dominant language, is not an accurate reflection of the child’s level of development. Instead, assessment is most accurate with “best performance” measures that assess the highest level of development attained by a bilingual child across both languages. Therefore, whenever possible, “best performance” measures across the two languages should be the technique of choice during bilingual assessments. Most school districts and speech-language pathology clinics lack the bilingual staff and financial resources to test individuals in the dozens of native languages of their client populations. The result is both over-identification (the client does not have an impairment, but just needs more time to learn the language) and under-identification (the client is assessed only in English, and the assessor inaccurately concludes that the client’s difficulties are related to learning a new language) of bilinguals. This state of affairs can be improved only if changes are made both at the systemic level—by increasing funding for services to linguistically diverse populations—and at the individual level—by raising clinicians’ understanding of bilingualism and its consequences. With regard to the latter, clinicians should be aware of the most recent findings in four areas: lexical organization, word-learning, cognitive control, and neural organization. Lexical Organization In children learning a first language, a noticeable change takes place in the salience of various word-word relations during middle childhood. For example, a 6-year-old is quick to point out the thematic relationship between an iron and a shirt (“Because you can iron a shirt!”) but has difficulties attributing the relationship between planes and buses to their shared taxonomy. They might say that “Planes and buses both have fumes” instead of recognizing that both are vehicles. By 8 years of age, most children readily acknowledge both thematic and taxonomic relationships (Hashimoto, McGregor, & Graham, 2007). Children learning two languages simultaneously or sequentially must store and retrieve a larger number of words, because vocabularies are distributed across two linguistic systems. Does access to different semantic relations arise in the same timeline for these children? “chair,” a child may produce “table,” “sit,” and “legs”). The bilingual children produced a similar number of taxonomic associations (e.g., chair-table) to the prompts in their two languages and in comparison to monolingual English-speaking peers (who were matched on performance IQ). The similarities in overall performance suggest that the emergence of taxonomic relations is largely determined by general cognitive abilities. Nevertheless, we found subtle differences—the bilingual children more frequently responded taxonomically than the monolingual children when the first associations and associations to verbs (e.g., jump-walk) were compared. This subtle bilingual advantage is interesting given that the bilingual children had a significantly smaller English receptive vocabulary than the monolinguals. The bilingual children’s need to store and retrieve more words across two linguistic systems may have rendered taxonomic relations more salient. More recently, Sheng, Bedore, and Peña (2008) compared word associations generated by Spanish-English bilingual children in their first and second languages. These children were considered relatively balanced bilinguals based on their linguistic input and output. The children showed overall comparable performance in the two languages, but there also was a subtle Spanish advantage over English in generating taxonomic associations to adjectives and verbs. We hypothesize that features of the Spanish language, such as the use of salient derivational endings (e.g., -oso, -ado, -ivo) to mark the adjective class and the use of verbs in more salient positions in an utterance may have led to an earlier appreciation of taxonomic relations for Spanish adjectives and verbs. Sheng, Bedore, and Peña (2009) are extending this research to bilingual children who have language impairment. Monolingual English-speaking children with language impairment exhibit a significant deficit in the use of both taxonomic and thematic relations in comparison to typically developing peers (Sheng & McGregor, in press). Investigations of bilingual children with language impairment will provide insights regarding the interactions among bilingualism (an experiential factor), linguistic capacity (a learner-internal factor), and vocabulary organization. Word-learning Speech-language pathologists have long been aware that application of monolingual language norms to bilingual clients is inappropriate. What are the alternatives? One possibility is to use processing-based measures, such as word-learning, to index language ability in bilinguals (e.g., Peña, Iglesias, & Lidz, 2001) because these tasks reflect a child’s general ability to process linguistic information but do not rely on extant linguistic knowledge. Therefore, bilinguals with poor language knowledge due to low proficiency should perform just as well on word-learning tasks as monolinguals, and better than bilinguals who experience language deficits. However, little is known about the effects of bilingualism on word-learning. How exactly does bilingualism influence word-learning ability? Our recent research comparing bilingual and monolingual adults on their ability to learn new words consistently suggests that bilingual adults tested in their native language outperform monolingual adults on word-learning tasks. For example, Kaushanskaya and Marian (2009a) examined word-learning performance in monolingual speakers, English-Spanish bilinguals, and English-Mandarin bilinguals, and found that both bilingual groups outperformed the monolingual group. A related study (Kaushanskaya & Marian, 2009b) examined the effects of bilingualism on adults’ ability to resolve cross-linguistic inconsistencies during novel word-learning. English monolinguals and English-Spanish bilinguals learned novel words that overlapped with English orthographically, but diverged from English phonologically. Native-language orthographic information presented during learning interfered with encoding of novel words in monolinguals, but not in bilinguals. These findings indicate that knowledge of two languages may shield bilinguals from native-language interference during novel word-learning. Current work (Kaushanskaya, Yoo, Van Hecke, & Mirsberger, 2009) suggests that monolinguals’ ability to learn new words depends on whether they learn new words silently or out loud. Conversely, bilinguals’ performance does not depend on any particular learning strategy, and they can acquire new words efficiently under any learning conditions. Our findings indicate that bilingualism facilitates word-learning performance in adults, although the precise mechanisms of this advantage remain unknown. Whether similar word-learning advantages can be observed in children is still under investigation. It appears that word-learning performance in bilingual children may be less contingent on latent vocabulary knowledge than in monolingual children (e.g., Kan & Kohnert, 2008; Wilkinson & Mazzitelli, 2003). However, studies that contrast word learning in simultaneous bilingual children (exposed to two languages from birth), sequential bilingual children, and monolingual children are necessary to identify the timeline and the mechanisms that underlie the development of the bilingual advantage for word learning. The finding that bilingualism facilitates word-learning performance has implications for the use of word-learning tasks to index language function in bilingual clients. If typically developing bilinguals perform at higher rates than typically developing monolinguals, then the expectations for bilingual clients with a suspected language difficulty may also need to be adjusted. Cognitive Control The consequences of bilingualism on cognition have implications for understanding the nature of linguistic-cognitive deficits. Linguistic and cognitive processes interact across the lifespan, with linguistic function tied to development of cognitive control throughout childhood and to its decline during aging (Comalli et al., 1962). For example, aging adults may have difficulty with language tasks that require inhibitory control, such as ignoring irrelevant language input when multiple speakers are present (Schrauf, 2008). Research suggests that the very processes that decline with normal aging also may be honed by lifelong bilingualism (Kavé et al., 2008). For example, aging bilinguals outperform monolingual peers at suppressing task-irrelevant information (Bialystok et al., 2004). Further, Bialystok, Craik, and Freedman (2007) showed that the onset of Alzheimer’s dementia may be delayed by up to four years in bilinguals relative to monolinguals. How does bilingual experience shape the cognitive system? In general, bilinguals face greater ambiguity during language processing because they consider similar-sounding words from two languages (instead of one) during comprehension and must choose between languages during production. For instance, a German-English bilingual who sees pictures of a bike and a leg while hearing “bike” & Marian, 2007). In our research, we aimed to identify a mechanism through which bilingual language processing may influence inhibitory control (Blumenfeld and Marian, in preparation). We measured the extent to which monolinguals and bilinguals activated similar-sounding words (e.g., “hamper” and “hammer”), and the extent to which they inhibited similar-sounding competitors as they identified the correct targets (e.g., a picture of a hammer). We found a correlation between how bilinguals (but not monolinguals) inhibited irrelevant words during comprehension and how well they performed on a non-linguistic task that required inhibition of irrelevant information. Bilinguals also showed higher accuracy rates on the non-linguistic inhibition task compared to monolingual peers. These findings suggest that a central inhibition mechanism may be recruited and altered by bilingual language processing. Identifying a link between language experience and cognitive processes is important because it may provide insights into how treatment can generalize from the cognitive into the linguistic domain and vice versa. Moreover, because inhibitory control deficits are thought to underlie (at least in part) a number of disorders, including attention-deficit (hyperactivity) disorder and frontal lobe impairments, monolingual/bilingual differences in this domain may become clinically relevant, generating the need to create bilingual norms even on non-linguistic neuropsychological assessment tools. In general, monolingual/bilingual differences should be considered in populations with potentially weaker cognitive control, such as children or older adults. Aspects of cognitive development or aging may differ across monolingual and bilingual populations, with potential consequences for the nature and severity of cognitive/linguistic symptoms related to inhibitory control. Neural Organization Investigations into the neural manifestations of bilingualism have included functional comparisons of a variety of linguistic and non-linguistic domains and studies of cortical anatomy. The earliest studies of the cortical correlates of bilingualism used behavioral approaches to examine hemispheric dominance differences between monolinguals and bilinguals, early- and late-acquired bilinguals, and high- and low-proficiency bilinguals. Hull and Vaid’s (2007) meta-analyses of the data reveal that early bilinguals were the only group that showed consistent bilateral dominance for language. Late bilinguals and monolinguals showed left-hemisphere dominance. Second-language proficiency was found to be less relevant than age of acquisition in influencing language lateralization. The authors proposed that a period of early monolingual development establishes left-hemispheric dominance that is then preserved irrespective of future bilingual experience. Interestingly, this decreased hemispheric dominance in early bilinguals also is observed for non-linguistic tasks. For example, Hausmann and colleagues (2004) used visual hemifield presentation to investigate face discrimination, a right-hemisphere-dominant task. Turkish-German bilinguals were more bilaterally dominant than both Turkish and German monolinguals. However, neuroimaging studies have failed to find consistent laterality differences between monolingual and bilingual speakers (e.g., Hernandez et al., 2001; Kim et al., 1997). When neural activations for single words are meta-analyzed on the basis of the lexical processes involved (semantic access, phonological code retrieval, or articulation), bilinguals and monolinguals activate similar neural regions for individual lexical processes (Indefrey, 2006; Indefrey & Levelt, 2004). What is different, though, is that specific perisylvian regions may differentially activate for individual languages of the bilingual speaker. The left inferior frontal gyrus (LIFG) has been shown to respond differentially to L1 and L2, either with different foci for L1 versus L2 or with greater volume of activation for L2 (Kim et al., 1997). This differential activation is found only for late bilinguals and for specific linguistic tasks. Marian and colleagues (2007), for example, found that the foci of LIFG activations differed across L1 and L2 for lexical and phonological processing, but not for orthographic processing. Others found L1 and L2 to activate the LIFG differentially for syntactic processing (Saur et al., 2009). The LIFG appears to make distinctions between L1 and L2 for linguistic processes for which it serves a unique role; further research is needed to elucidate these patterns. Moreover, bilingualism may have ramifications on cortical morphology: Using high-resolution magnetic resonance imaging scans and an analysis procedure called voxel-based morphometry, Mechelli and colleagues (2004) found that individuals with higher proficiency in and/or earlier age of second-language acquisition had a higher gray matter density in the left inferior parietal cortex. What Clinicians Should Know Knowledge of bilingualism suggests the following linguistic, cognitive, and neurophysiological differences between bilingual and monolingual speakers: Linguistic differences Bilingual children develop an earlier understanding of taxonomic relationships than their monolingual peers (e.g., car and bus are vehicles). This understanding is not dependent on vocabulary size, but could be influenced by the structural features of the speaker’s language. Bilingual adults are better than monolingual adults at learning new words. Bilinguals use a variety of word-learning strategies with similar efficiency and are less susceptible to interference from conflicting orthographic information during word-learning. Linguistic input co-activates both languages in bilinguals; when bilinguals hear or read words in one language, partially overlapping linguistic structures in the other language also are activated. Cognitive differences Bilinguals may be able to inhibit irrelevant verbal and nonverbal information with greater ease than monolinguals. Inhibitory control ability is slower to decline with age in bilinguals than in monolinguals. The average age of dementia onset is later in bilinguals than in monolinguals. Bilingual children have been found to exhibit superior performance in divergent thinking, figure-ground discrimination, and other related meta-cognitive skills. Neural differences Bilateral processing of language (and other nonverbal tasks) is most likely to occur only in early bilinguals. Monolinguals and bilinguals use similar neural regions for language processing. However, late bilinguals are likely to activate the LIFG differentially for processes in which the LIFG plays a crucial role, such as phonological and syntactic processing. Bilinguals have greater gray matter density than monolinguals in certain left hemisphere regions. Did You Know? According to the 2000 U.S. Census, a language other than English was spoken in approximately 18% of all American households. According to the U.S. Census, the Hispanic population in 2007 was 45.5 million, a number expected to grow to 47.7 million by 2010 and 59.7 million in 2020. Approximately 7.5 million bilingual children were enrolled in U.S. schools in 2002. U.S. Census information on language use and the incidence of bilingualism (based on the 2000 Census) can be found on the U.S. Census Bureau Web site. ASHA and the National Institute on Deafness and Other Communication Disorders estimate that 10–15% of the U.S. population has a speech-language or hearing disorder. These estimates are higher among persons from socially and economically disadvantaged groups, including recent immigrants. 3.5% of ASHA have to be bilingual speech-language pathologists or 2009). can and about bilingualism on the Bilingual families can through groups such as Bilingual and in the Web site. A map of languages spoken across the U.S. can be found on the Web site. The of language and of these languages are about the National for Bilingual can be found on the Web site. The of a used among has a Spanish-English bilingual More and in A of the of on & of and Bilingual Google Scholar in Language, and Google Scholar & and cognitive from the and Scholar Blumenfeld & Marian on activation in bilingual spoken language proficiency and lexical and Cognitive Scholar Blumenfeld & Marian preparation). inhibitory control in Google Scholar Blumenfeld & Marian interactions during bilingual language development in & in Google Scholar Blumenfeld & Marian and in comprehension across the & of the of the Cognitive Cognitive Google Scholar & effects of test in and of Google Scholar The consequences of bilingualism for cognitive & in Google Scholar The influence of bilingualism on cognitive A of research findings and on Google Scholar The impact of bilingualism on cognitive of Research in American Research Google Scholar Bilingual cognitive three in Development, Scholar K. of The on Google Scholar & comprehension and A and a new The of and Google Scholar & at and from the semantic of object of Language, and Scholar & for hemispheric in in of Google Scholar Hernandez & and language in Spanish-English Google Scholar Hull & Bilingual language A of two Google Scholar Indefrey A of studies on first and second language differences can we and what do they Google Scholar Indefrey & The and of word Google Scholar Kan & K. by developing bilinguals in L1 and of Language, Google Scholar Kaushanskaya & Marian The bilingual advantage in novel word and Google Scholar Kaushanskaya & Marian native-language interference in novel word of & Cognition, Google Scholar Kaushanskaya Van & new words silently between monolinguals and bilinguals. presented at the and and Knowledge in and Google Scholar & and cognitive state in the and Scholar Kim & cortical associated with native and second Google Scholar Marian Blumenfeld Kaushanskaya Faroqi-Shah & activation during word processing in late and differences as by functional magnetic resonance of and Google Scholar Mechelli et in the bilingual Google Scholar & Lexical development in bilingual and to monolingual Scholar Peña & test through of Scholar et processing in the bilingual Google Scholar and & to and Google Scholar Sheng & press). in children with specific language of Language, and Google Scholar Sheng & Marian in bilingual from a word of Language, and Scholar Sheng & Peña in Spanish-English bilingual of the on Bilingual in Google Scholar Sheng & Peña of semantic knowledge in Spanish-English bilingual children with specific language of the on Research in Google Scholar & Inhibitory processes and spoken word in and older The of lexical and semantic and Scholar & Cognitive and age differences in following control and Google Scholar Umbel & changes in receptive vocabulary in Hispanic bilingual school Lexical in Google Scholar & Lexical knowledge of monolingual and bilingual Google Scholar and of vocabulary in to acquisition and of Google Scholar Wilkinson & K. The of information on children’s of of Language, Google Scholar et and differences in inhibitory from a bilingual and Cognition, Scholar Viorica PhD, is of communication and disorders, and cognitive at and of the and cognitive, and measures to study the linguistic capacity and the ability to multiple languages her at Yasmeen PhD, CCC-SLP, is in the of and of the in and Cognitive and of the Research at the of research on and her at Margarita Kaushanskaya, PhD, is in the of Disorders and of the and at the of research the nature and the of cognitive mechanisms that underlie language learning and on second-language acquisition and bilingualism in children and adults. her at Henrike K. PhD, is in the of Language, and and of the and at research on the relationship between linguistic and cognitive processes in and aging bilinguals. her at Li Sheng, PhD, is in the of Communication and Disorders at the of at research on child language development and disorders, processing and organization, and her at With With to in Oct & American
We introduce a formal framework that allows the calculation of new purely statistical confidence measures for parsing, which are estimated from posterior probability of constituents. These measures allow us to mark each constituent of a parse tree as correct or incorrect. Experimental assessment using the Penn Treebank shows favorable results for the classical confidence evaluation metrics: the CER and the ROC curve. We also present preliminar experiments on application of confidence measures to improve parse trees by automatic constituent relabeling.
The Postmodern culture today breaks down historically a solid barrier, produced in the modern age, between the language and its users, signs and realities and subject and his object. This phenomenon brings about some translation problems at the same time; interpretative diversity, linguistic derivation in the mass media, permanent reproduction of translated text and meaning`s continuity, definition of translator etc. I examine these problems through a theoretical approach to the mediative nature of the act of translation. I stress on two inevitable aspects of the postmodern linguistic tendencies: connotation generalized in our translation activities and its re-mediative culture. Focusing on a cycling perspective of our translating activities, I could reach the conclusion that the translated text have no relation with the denotative meanings which have been comprised closed and original with the realities from the modern age. A corrected model could be proposed. I call this cycle model of translating process which comprises ① Mediation or Creation ② Re-mediation or Translation ③ Re-re-mediation or Comprehension ④ Verification, Deduction or Correction. Each activity has not only its own translating process but its circulated role for the time in which everyone can participate equally as a reader, a sender, a translator and an individual who makes its contextual needs and desires. This model could show that the translating process is not for fixing a linguistic sign to some closed meanings but for expanding its pertinent meanings to diverse situations. Translating activity is not for making a linguistic norm by itself, but for making appropriate communication with as much of the population as possible.
A recent proposal (Pollock 1989) within the framework of Government and Binding (GB) grammatical theory has been that the members of INFL Agreement and Tense should be given full constituent status as maximal projections in their own right. This idea has been applied to the syntax of Modern Irish in order both to test the universality of the expanded INFL proposal and to investigate what new perspectives it might have to offer on some remaining problems of Irish syntax. The results are presented in the following paper along with discussions of the direction they suggest for further research. INTRODUCTION Using data from mostly English and French, J.Y. Pollock argues in a recent proposal (1989) that if the usual members of INFL, Agreement and Tense, are included in the syntax as full maximal projections, many of the phenomena surrounding auxiliaries, negation, and verb movement can receive straightforward explanations. The proposal seems readily adaptable for other SVO languages which are generally accepted as showing evidence of verb movement, notably the so-called Verb Second (V2) languages. In order to test the universality of the expanded-INFL proposal, an expandedINFL syntax has been applied to the model VSO language Modern Irish. The result has been a quite promising new syntactic structure for Irish which seems to confirm the universality of expanded-INFL. While it is fully compatible with existing analyses for Irish word order in which V S O is derived from SVO, the new expanded syntax is equally adaptable to an account deriving VSO from SOV. Such an account is suggested by the Irish infinitive clause, which is built around the verbal noun (VN), and which regularly shows surface SOV order. The new syntax provides an attractive solution for the placement of preverbal particles (interrogative, relative, negative, and copula), which are the only elements regularly allowed to precede the verb in Irish. It also suggests some interesting perspectives for the analysis of copula constructions, an area which remains an open question in Irish syntax. 58 SHEILA DOOLEY COLLBERG Expanded-INFL syntax I would like to begin by defining exactly what is meant here by an expanded-INFL syntax. This is my own terminology for the kind of structure proposed in Pollock 1989. It is probably easiest to see what is new about this structure if we compare it to earlier models of universal syntax. Through the years, the 'basic' syntactic tree structure assumed within the G B theoretical framework has steadily grown more complex and abstract. The first tree structure (a) above shows a pre Barriers (Chomsky 1986) type of syntax with really the bare essentials. The S portion of the tree is the area which undergoes the most change. In the second tree (b), after Barriers, we have a new level of constituent structure introduced: INFL (inflection). It corresponds roughly to the S level of the previous structure. We also see that there is an abstract element Agr (Agreement) which is assumed to be generated in INFL. The whole tree shows consistent 2-level expansion of X-bar syntax for each phrasal projection. The last tree above (c) is an example of the expanded-INFL syntax: The IP of (b) has grown into two fully expanded phrasal projections in their own right: AgrP and TP (Tense). This of course gives us a lot more 'room' in the syntax to propose analyses for grammatical phenomena involving the abstract (or AN EXPANDED-INFL SYNTAX FOR MODERN IRISH 59 overt) elements Agr and Tense, namely things like the behavior of auxiliaries, subject-verb inversion, negation, quantifiers, and verb movement. As Pollock demonstrates, this kind of structure can be used to explain many of the word order details of the SVO languages French and English — details which otherwise seem unexplainable except by recourse to ad hoc stipulations. B A S I C I R I S H S Y N T A C T I C S T R U C T U R E Can the kind of structure pictured in (lc) say anything new to us about Irish? Can we implement such a structure at all for a V S O language like Irish? The answer depends in part upon how one decides to analyze the surface V S O order of Irish. There are two possible analyses, both represented in the existing literature. V S O is base-generated Stenson 1981 and Chung 1983 are two studies which represent the view that the V S O order in Irish is base-generated. This implies that the syntactic structure is a flat, one-level tree with all constituent phrases placed as sisters to the initial verb and no verb movement involved. It accurately represents the observed surface word order of Irish and is thus descriptively adequate, but it offers little explanation for the verb-initial order. Chung attempts to give a possible theoretical defense of the flat structure by appealing to the observation that VSO languages seem to lack the subject-object asymmetries with regard to extraction properties that one usually finds in S V O languages. However, this is not quite correct. The subject NP in Irish is much more closely tied to the verb than the object NP. While nothing can ever intervene between the subject and the verb, there are times when the object is in fact forced to move away from its canonical position. This occurs when the object is pronomimal. It must, appear in absolute final position in its clause, and it apparently reaches this position by means of some sort of a rule of Pronoun Postposing (Chung & McCloskey 1987). These facts suggest that the relationship of the subject and object NP to the verb is not simply one of equal sisterhood. The S V O Analysis If the VSO order of Irish is not base-generated, then it must arise through some sort of derivational process from a different underlying word order. This view is implicitly supported in an article devoted to establishing the 60 SHEILA DOOLEY COLLBERG existence of a V P in Irish (McCloskey 1983). The existence of a V P entails at least two hierarchical levels of sentence structure, with the verb originating in a V O or OV constituent and obligatorily fronted to some other position. Sproat 1985 builds on the work of McCloskey to develop a full SVO Analysis for Welsh, arguing that the same analysis may be applied to Irish. The underlying structure for the two languages is argued to be SVO, and the obligatory fronting of the finite verb is made to follow from the requirements of case theory. Sproat maintains that while INFL in SVO or SOV languages may assign nominative case either to the left or the right, INFL in VSO languages is restricted to assigning case rightward. The verb lexicalizing INFL is thus forced to appear to the left of the subject NP in order to assign nominative case successfully. Sproat's SVO Analysis is a step in the right direction in that it gives a theoretically attractive explanation for the obligatory fronting of the verb, but it is incomplete in that Sproat does not specify any landing site for the conjoined verb and INFL. Without going into any more detail, it may be said that the arguments for the SVO Analysis are quite attractive, and the general consensus among Celtic syntacticians seems to be that Irish is SVO underlyingly. In general, a derivational account like this for verb-initial languages is pretty much the norm now, as can be seen in recent works of a typological, nature such as Koopman & Sportiche 1988. EXPANDED-INFL FOR IRISH Obviously, it should be possible to adapt the Pollock type of syntax for Irish if we accept that Irish VSO order is derived from SVO. So let us assume that for the moment. Then, of course, there are plenty of language-specific details to work out, and the following sections contain suggestions for handling these. My proposal for the full syntactic structure of Irish is given in (2) and wil l be referred to throughout the ensuing discussion. Principles and parameters according to Pollock Given in (3) is a very brief summary of the most important points that Pollock argues for in his article. These can be reduced to a pair of universal principles (I and II) and a set of parameters (III) which vary from language to language. AN EXPANDED-INFL SYNTAX FOR MODERN IRISH 61
This paper presents a novel application of incorporating Alternating Structure Optimization (ASO) to conduct the task of text chunking of Semantic Role Labeling (SRL) in Chinese texts. ASO is a competent linear algorithm based on the theory of multi-task learning. In this paper, by constructing several SRL tasks to constitute a multi-task, we are able to encode the inference obtained by ASO algorithm as additional feature to further boost the performance of the target task employing Conditional Random Fields (CRFs). To our knowledge, our method is the first that incorporates multi-task learning into a statistical model in SRL for Chinese texts. We evaluate our approach on Penn Treebank data sets and obtain encouraging result.
This paper presents a tool for extracting multi-word expressions from corpora in Modern Greek, which is used together with a parallel concordancer to augment the lexicon of a rule-based machine-translation system. The tool is part of a larger extraction system that relies, in turn, on a multilingual parser developed over the past decade in our laboratory. The paper reviews the various NLP modules and resources which enable the retrieval of Greek multi-word expressions and their translations: the Greek parser, its lexical database, the extraction and concordancing system.
In this paper, we propose a modular cascaded approach to data driven dependency parsing. Each module or layer leading to the complete parse produces a linguistically valid partial parse. We do this by introducing an artificial root node in the dependency structure of a sentence and by catering to distinct dependency label sets that reflect the function of the set internal labels vis-a¿-vis a distinct and identifiable linguistic unit, at different layers. The linguistic unit in our approach is a clause. Output (partial parse) from each layer can be accessed independently. We applied this approach to Hindi, a morphologically rich free word order language using MST parser. We did all our experiments on a part of Hyderabad Dependency Treebank. The final results show an increase of 1.35% in unlabeled attachment and 1.36% in labeled attachment accuracies over state-of-the-art data driven Hindi parser.
The Arabic language has a very rich/complex morphology. Each Arabic word is composed of zero or more prefixes, one stem and zero or more suffixes. Consequently, the Arabic data is sparse compared to other languages such as English, and it is necessary to conduct word segmentation before any natural language processing task. Therefore, the word-segmentation step is worth a deeper study since it is a preprocessing step which shall have a significant impact on all the steps coming afterward. In this article, we present an Arabic mention detection system that has very competitive results in the recent Automatic Content Extraction (ACE) evaluation campaign. We investigate the impact of different segmentation schemes on Arabic mention detection systems and we show how these systems may benefit from more than one segmentation scheme. We report the performance of several mention detection models using different kinds of possible and known segmentation schemes for Arabic text: punctuation separation, Arabic Treebank, and morphological and character-level segmentations. We show that the combination of competitive segmentation styles leads to a better performance. Results indicate a statistically significant improvement when Arabic Treebank and morphological segmentations are combined.
Spatial models are employed to represent conceptual data in a wide range of fields within psychological research. In order to generate spatial models, it is necessary to first obtain empirical similarity data. A number of methods are available for collecting these data, but little effort has been made to compare their relative utility. In this article, we compare directly rated and five feature-based similarity data types in regard to their ability to be adequately represented by a spatial model (representational goodness of fit), and the ability of the representations to predict three external empirical variables (predictive validity). The results indicate that the representational goodness of fit of the feature-based similarities is noticeably superior to the directly rated similarities, and that the predictions of representations derived from common feature similarity data are substantially more likely than the predictions of all of the alternative representations. It is suggested that these findings are highly relevant to researchers employing spatial models to represent conceptual data, given that direct pairwise ratings have generally been considered the “gold standard” means of obtaining empirical similarities.
OBJECTIVES: This study is the first in a series designed to develop and norm new theoretically motivated sentence tests for children. The purpose was to examine the independent contributions of word frequency (i.e., how often words occur in language) and lexical density (the number of similar sounding words or "neighbors" to a target word) to the perception of key words in the new sentence set. DESIGN: Twenty-four children with normal hearing aged 5 to 12 yrs served as participants; they were divided into four equal age-matched groups. The stimuli consisted of 100 semantically neutral sentences that were 5 to 7 words in length. Each sentence contained 3 key words that were controlled for word frequency and lexical density. Words with few neighbors come from sparse neighborhoods, whereas words with many neighbors come from dense neighborhoods. The key words within a sentence belonged to one of the four lexical categories: (1) high-frequency sparse, (2) low-frequency dense, (3) high-frequency dense, and (4) low-frequency sparse. Participants were administered the sentence list and the 300 key words in isolation at 65 dB SPL. Each participant group was tested in spectrally matched noise at one of the four signal-to-noise ratios (SNRs -2, 0, 2, and 4 dB). The percent of words correctly identified was calculated as a function of SNR, key word context (sentences vs. words), and key word lexical category. RESULTS: SNR had a significant effect on the recognition of key words in sentences and in isolation; performance improved at higher SNRs. There were significant main effects of word frequency and lexical density as well as a significant interaction between the two lexical factors. In isolation, high-frequency words were recognized more accurately than low-frequency words. In both word and sentence contexts, sparse words yielded greater accuracy than dense words, irrespective of word frequency. There was a modest but significant negative correlation between lexical density and the recognition of words in isolation and in sentences. CONCLUSIONS: Word frequency and lexical density seem to influence word recognition independently in children with normal hearing. This is similar to earlier results in adults with normal hearing. In addition, there seems to be an interaction between the two factors, with lexical density being more heavily weighted than word frequency. These results give us further insight into the way children organize and access words from long-term lexical memory in a relational way. Our results showed that lexical effects were most evident at poorer SNRs. This may have important implications for assessing spoken-word recognition performance in children with sensory aids because they typically receive a degraded auditory signal.
Semantic processing represents the new challenge for all applications that require text understanding, as for instance Q/A. In this paper we will highlight the need to couple statistical approaches with deep linguistic processing and will focus on ldquoimplicitrdquo or lexically unexpressed linguistic elements that are nonetheless necessary for a complete semantic interpretation of a text. We will address the following types of ldquoimplicitrdquo entities and events: - grammatical ones, as suggested by a linguistic theories like LFG or similar generative theories; - semantic ones suggested in the FrameNet project, i.e. CNI, DNI, INI; - pragmatic ones: here we will present a theory and an implementation for the recovery of implicit entities and events of (non-) standard implicatures. In particular we will show how the use of commonsense knowledge may fruitfully contribute in finding relevant implied meanings. We will also briefly explore the subject of point of view which is computed by semantic informational structure and contributes the intended entity from whose point of view is expressed a given subjective statement. We also present an evaluation based on section 24 of Penn Treebank as encoded by LFG people in the PARC-700 treebank where lexically unexpressed are adequately classified and diversified.
Abstract. This paper presents a method to incorporate statistical information into a rule-based parser to resolve syntactic ambiguities. We extract the statistical information from the Penn Treebank, and apply the information to the rule-based parser. For the extraction of the statistical information the tag conversion is needed because of the disagreement of the tags and the bracketing style. We will show the effect of the tag conversion with experiments. The final result shows about 7 % error rate reduction in the dependency evaluation. We will also show how much each type of statistical information affects the parsing performance.
The paper concentrates on obtaining hidden relationships among individual clauses of complex sentences from the Prague Dependency Treebank. The treebank contains only an information about mutual relationships among individual tokens (words, punctuation marks), not about more complex units (clauses). For the experiments with clauses and their parts (segments) it was therefore necessary to develop an automatic method transforming the original annotation into a scheme describing the syntactic relationships between clauses. The task was complicated by a certain degree of inconsistency in original annotation with regard to clauses and their structure. The paper describes the algorithm of deriving clause-related information from the existing annotation and its evaluation.
Voice acoustic analysis is typically a labor-intensive, time-consuming process that requires the application of idiosyncratic parameters tailored to individual aspects of the speech signal. Such processes limit the efficiency and utility of voice analysis in clinical practice as well as in applied research and development. In the present study, we analyzed 1,120 voice files, using standard techniques (case-by-case hand analysis), taking roughly 10 work weeks of personnel time to complete. The results were compared with the analytic output of several automated analysis scripts that made use of preset pitch-range parameters. After pitch windows were selected to appropriately account for sex differences, the automated analysis scripts reduced processing time of the 1,120 speech samples to less than 2.5 h and produced results comparable to those obtained with hand analysis. However, caution should be exercised when applying the suggested preset values to pathological voice populations.