Papers reviewed and determined not to be word norm studies. Use the flag icon to report errors or suggest re-inclusion.
16504 papers
OBJECTIVES: This study is the first in a series designed to develop and norm new theoretically motivated sentence tests for children. The purpose was to examine the independent contributions of word frequency (i.e., how often words occur in language) and lexical density (the number of similar sounding words or "neighbors" to a target word) to the perception of key words in the new sentence set. DESIGN: Twenty-four children with normal hearing aged 5 to 12 yrs served as participants; they were divided into four equal age-matched groups. The stimuli consisted of 100 semantically neutral sentences that were 5 to 7 words in length. Each sentence contained 3 key words that were controlled for word frequency and lexical density. Words with few neighbors come from sparse neighborhoods, whereas words with many neighbors come from dense neighborhoods. The key words within a sentence belonged to one of the four lexical categories: (1) high-frequency sparse, (2) low-frequency dense, (3) high-frequency dense, and (4) low-frequency sparse. Participants were administered the sentence list and the 300 key words in isolation at 65 dB SPL. Each participant group was tested in spectrally matched noise at one of the four signal-to-noise ratios (SNRs -2, 0, 2, and 4 dB). The percent of words correctly identified was calculated as a function of SNR, key word context (sentences vs. words), and key word lexical category. RESULTS: SNR had a significant effect on the recognition of key words in sentences and in isolation; performance improved at higher SNRs. There were significant main effects of word frequency and lexical density as well as a significant interaction between the two lexical factors. In isolation, high-frequency words were recognized more accurately than low-frequency words. In both word and sentence contexts, sparse words yielded greater accuracy than dense words, irrespective of word frequency. There was a modest but significant negative correlation between lexical density and the recognition of words in isolation and in sentences. CONCLUSIONS: Word frequency and lexical density seem to influence word recognition independently in children with normal hearing. This is similar to earlier results in adults with normal hearing. In addition, there seems to be an interaction between the two factors, with lexical density being more heavily weighted than word frequency. These results give us further insight into the way children organize and access words from long-term lexical memory in a relational way. Our results showed that lexical effects were most evident at poorer SNRs. This may have important implications for assessing spoken-word recognition performance in children with sensory aids because they typically receive a degraded auditory signal.
Dolgozatomban, amint arra a címből is lehet következtetni, az 1996 és 2005 között adatolható magyarországi börtönszlenget mutatom be, azt a csoportnyelvet, amelynek az átfogó tanulmányozása hazánkban eddig még nem történt meg. Munkámban büntetés-végrehajtási intézeteink fogvatartottjainak belső, informális nyelvhasználatának általános kérdéseivel foglalkozom, és a mai magyar börtönszleng szó- és kifejezéskészletének szótárba foglalásán túl kísérletet teszek a vizsgált csoportnyelv nyelvi-szociolingvisztikai leírására. \n \nCélkitűzésemet, a magyar börtönszleng átfogó tanulmányozását az indokolta, hogy a kutatás első éveiben olyan mennyiségű és minőségű, a nyelvtudományban eddig még nem tárgyalt adatokra bukkantam, amelyek érdemesnek mutatkoztak arra, hogy egy mélyebb, megtervezett szlengkutatás irányuljon erre a területre. \n \nElsősorban célom volt a magyar börtöszlenget feltárni, bemutatni keletkezését, funkcióját, működését, a szlenghasználó közösségben betöltött szerepét. Célom volt továbbá rámutatni nyelvi előzményeire, összevetni a már létező bűnözői nyelvi adatbázissal, vagyis a tolvajnyelv elemeivel, egyben definiálni helyét a magyar szlengkutatás területén. Kutatásom során mindvégig azt tartottam szem előtt, hogy hol van az ember a szlengben, így célom volt annak leírása is, mikor, milyen körülmények között motiváltak a vizsgált csoport tagjai szlenghasználatra. Ennek kiderítéséhez a nyelvi adatok feltárásán és rendszerezésén túl, a zárt közeg csoportjainak vizsgálatára is ki kellett terjeszteni a kutatást, megfigyelve a csoportszerveződés lehetőségeit és okait a börtöntársadalomban. \n \nAs already the title has suggested, my dissertation presents Hungarian prison slang as attested from 1996 to 2005. This group language has never seen an overall study in Hungary up to now. I deal with general questions of the internal and informal language use of the prisoners of penal institutions in my work and, together with rendering the words and expressions into a dictionary, I attempt at the linguistic and sociolinguistic description of the group language examined. \n \nMy objective, the overall study of Hungarian slang, is justified by the data being of such quantity and quality and having never been dealt with in linguistics that seemed worth to be examined by a deeper and planned slang research. \n \nMy main objective was to explore Hungarian prison slang, to present its origins, functions and operation as well as its role within slang user communities. My aim also was to show its linguistic predecessors, to compare it to an existing linguistic database of criminals, that is, to the elements of cant, and, together with it, to define its place in the area of Hungarian slang research. During my studies, I always kept man’s role in slang in mind so my objective was the description of the conditions among which the members of a researched group are motivated for slang usage. In order to learn about it, I had to extend my research to the examination of the groups of closed space, observing the possibilities and reasons within prison community.
In recent years, the specter of litigants turning to religious or customary sources of law as authoritative guides to regulate their behavior, alongside or in lieu of secular norms, has risen to the forefront of politics in many countries worldwide. In this essay, we draw upon citizenship theory and comparative constitutional jurisprudence to identify two different categories of judicial response to religious-based claims for recognition, accommodation, and exemption: 1) 'diversity as inclusion;' and 2) 'non-state law as competition.' As long as legal claims for accommodation are not seen by courts as challenging the lexical superiority of the constitutional religion itself ('diversity as inclusion'), they stand a fair chance of success. Contrast that with the unyielding reluctance of legislatures and judiciaries to accept as binding or even cognizable any potentially competing legal order that originates in sacred or customary sources of identity and authority. This pattern of clamping down and refusing to accept any alternative sources of regulation becomes particularly visible where the legal challenge at issue is interpreted as raising doubts regarding which set of norms and institutions, or what set of high priests, should have the final word in authoritatively resolving legal disputes within a given society ('non-state law as competition'). This is a challenge that no secular legal order, no matter how tolerant and otherwise open to providing exemptions and accommodations to religious believers, can accept with indifference. For what perceived to be at stake here is the very authority and source of legitimacy of the accepted civil religion. We demonstrate these claims by focusing on recent jurisprudence from Canada and South Africa, two polities that represent the most difficult cases for our argument; if there is any place we would expect to find recognition by secular countries of religious or customary sources of law and authority, it would be in these diverse societies that have made an explicit constitutional commitment to promote their citizens’ freedom to preserve and enhance their multitude of backgrounds and distinctive cultural, linguistic and religious heritages as part of their 'mosaic' (Canada) or 'rainbow nation' (South Africa) conceptions of citizenship. Although operating in different contexts, the South African Constitutional Court and the Supreme Court of Canada seem to have made every effort to subject traditional legal regimes to general principles of constitutional law. By so doing, they have erected a new wall of separation that places noncompliance with the values of the civil religion beyond the pale of accepted accommodation, offering to those who espouse them the potential to either bring these alternative legal domains under the general rule of constitutional law or encounter the wrath of state fiat.
Phlebography, has been compared with the results of venous capacity measurement taken from 150 patients, 76 women and 74 men. The measurement of the venous capacity is a appropriate screening method to determine any irregularities of the deep venous hemodynamic. It is possible to attain a more exact indication for the phlebography. Maximum expression of the values of venous capacity, is however, contained in the functional periodical controls of the pathologic processes of the deep venous system.
The paper reports on the main findings of the LANCHART language attitudes studies. These studies were designed to falsify (or modify) the picture of adolescent language ideology – and its role in language change – that had emerged from previous sociolinguistic studies in Denmark. This picture is formulated as three hypotheses: (1) There are two value systems at two levels of consciousness, (2) Language change is governed by subconscious values, (3) Copenhagen is Denmark’s only linguistic norm centre. Following strict guidelines for data collection among 9th graders (aged 15–16) in Copenhagen, Næstved, Vissenbjerg, Odder, and Vinderup we obtained subconsciously offered attitudes that could be compared with consciously offered attitudes. The results neither falsify nor modify the established picture but strongly confirm it.
ABSTRACT. We present an overview of the Index Thomisticus Treebank project (IT-TB). The IT-TB consists of around 60,000 tokens from the Index Thomisticus by Roberto Busa SJ, an 11million-token Latin corpus of the texts by Thomas Aquinas. We briefly describe the annotation guidelines, shared with the Latin Dependency Treebank (LDT). The application of data-driven dependency parsers on IT-TB and LDT data is reported on. We present training and parsing results on several datasets and provide evaluation of learning algorithms and techniques. Furthermore, we introduce the IT-TB valency lexicon extracted from the treebank. We report on quantitative data of the lexicon and provide some statistical measures on subcategorisation structures. RÉSUMÉ. Nous présentons une vue d’ensemble du projet de l’Index Thomisticus Treebank (IT-TB). L’IT-TB consiste d’environ 60,000 occurrences tirées de l’Index Thomisticus de Roberto Busa SJ, un corpus de onze millions de mots latins de Thomas d’Aquin. Nous décrivons brièvement les règles d’étiquetage, qui sont en commun avec la Latin Dependency Treebank (LDT). Nous décrivons l’application des parseurs probabilistes dépendanciels sur les données de l’IT-TB et de la LDT. Nous présentons les résultats de l’entraînement et de l’analyse syntactique sur plusieurs ensembles des données et nous fournissons une évaluation des algorithmes et des techniques d’apprentissage. En outre, nous introduisons le lexique de valence de l’IT-TB tiré de la treebank. Nous reportons les données quantitatives du lexique et nous fournissons quelques mesures statistiques sur les structures de sous-catégorisation.
The paper describes an auditory experiment aimed at testing whether the intrinsic loudness of a stimulus with a given voice quality influences the way in which it signals affect. Synthesised voice quality stimuli in which intrinsic loudness \nwas systematically manipulated were presented to listeners to test the effect of this manipulation on the affective colouring \nof the stimuli. The results showed that even when devoid of intrinsic loudness variation, non-modal voice quality stimuli \nwere capable of communicating affect. However, changing the loudness of a non-modal voice quality stimulus towards its \nintrinsic loudness resulted in the increase of affective ratings.
creativeness / a pleasing field / of bloom Word associations are an important element of linguistic creativity. Traditional lexical knowledge bases such as WordNet formalize a limited set of systematic relations among words, such as synonymy, polysemy and hypernymy. Such relations maintain their systematicity when composed into lexical chains. We claim that such relations cannot explain the type of lexical associations common in poetic text. We explore in this paper the usage of Word Association Norms (WANs) as an alternative lexical knowledge source to analyze linguistic computational creativity. We specifically investigate the Haiku poetic genre, which is characterized by heavy reliance on lexical associations. We first compare the density of WAN-based word associations in a corpus of English Haiku poems to that of WordNet-based associations as well as in other non-poetic genres. These experiments confirm our hypothesis that the non-systematic lexical associations captured in WANs play an important role in poetic text. We then present Gaiku, a system to automatically generate Haikus from a seed word and using WAN-associations. Human evaluation indicate that generated Haikus are of lesser quality than human Haikus, but a high proportion of generated Haikus can confuse human readers, and a few of them trigger intriguing reactions.
espanolComo es bien sabido, aunque para los hablantes de una lengua las variedades dialectales resulten mas evidentes en los planos lexico, fonetico o fonologico, ellas se advierten en todos los niveles del lenguaje, orbita de la que, por supuesto, no escapa la sintaxis. Asi, en el caso particular del espanol de Buenos Aires, el uso del Preterito Perfecto Compuesto del Modo Indicativo difiere sensiblemente de la norma castellana, a la vez que la conciencia de los hablantes de la lengua respecto de el es practicamente nula: o lo niegan por completo, alegando que prefieren siempre el Preterito Perfecto Simple, o bien aducen que lo emplean segun la norma de Madrid; lo cual, como se vera a lo largo de nuestro trabajo, no resulta de ese modo en ninguno de los dos casos. Asi pues, intentaremos problematizar las cuestiones de norma y uso, en relacion con la conciencia de los hablantes portenos respecto de su empleo de los tiempos pasados. Para ello, partiremos de un trabajo de campo que hemos realizado y que nos ha permitido esbozar algunos matices caracteristicos del uso del tiempo verbal que nos ocupa, es decir, el Preterito Perfecto Compuesto del Modo Indicativo del dialecto rioplatense. EnglishIt is well known that dialectal language variations appear at every level of language including syntax. However, speakers are usually aware of lexical, phonetics, and phonological variations only. In this particular case, as expected, the use of perfect tenses in Buenos Aires (Argentina) is very different from that of Madrid (Spain). The problem is that most Argentinean speakers know how to use the Present Perfect according to Spanish rules they have learned in school, but their speech do not matches their learning. Most Argentinean speakers would say (and they believe) that they do not use the Present Perfect in everyday life, when they actually do, albeit in a different way. That is why I conducted a survey among speakers of all kind of age, in order to distinguish some specific characteristics of the Present Perfect use in rioplatense dialect. Finally, I intend to discuss the concept of language norm and use related to speakers' awareness in Buenos Aires.
The paper presents a preliminary study on discourse connectives (DC) in Czech. Aiming to build a computerized language corpus capturing discourse relations in Czech, we base our observations on current foreign projects with the same purpose. In this study, first, the different methods of linguistic analysis of the discourse structure and discourse connec- tives are described, next, the nature and properties of the group of DCs are analyzed and, finally, the procedure of the annotation of discourse connectives in Prague is presented.
Résumé Dans cette étude, qui se base sur un petit corpus de textes français et suédois originaux et traduits, l’usage des formes lexicales et pronominales en fonction anaphorique est examiné. Comme prévu, les formes pronominales s’avèrent plus fréquentes dans les textes français que dans les textes suédois, où la répétition du SN lexical thématique prédomine. Cette différence est en général liée aux normes rhétoriques ou stylistiques des cultures respectives, telles que la plus haute fréquence de marqueurs explicites de cohérence textuelle et le besoin plus fort de variation lexicale dans les langues romanes comparées aux langues germaniques, mais l’absence systématique de pronoms anaphoriques dans les textes suédois demande une autre explication. Elle semble due à une répugnance générale des pronoms personnels suédois inanimés ( den, det ) d’assumer une fonction anaphorique.
A novel series of trifluoromethyl-containing quinazoline derivatives with a variety of functional groups was designed, synthesized, and tested for their antitumor activity by following a pharmacophore hybridization strategy. Most of the 20 compounds displayed moderate to excellent antiproliferative activity against five different cell lines (PC3, LNCaP, K562, HeLa, and A549). After three rounds of screening and structural optimization, compound 10 b was identified as the most potent one, with IC<sub>50</sub> values of 3.02, 3.45, and 3.98 μM against PC3, LNCaP, and K562 cells, respectively, which were comparable to the effect of the positive control gefitinib. To further explore the mechanism of action of 10 b against cancer, experiments focusing on apoptosis induction, cell cycle arrest, and cell migration assay were conducted. The results showed that 10 b was able to induce apoptosis and prevent tumor cell migration, but had no effect on the cell cycle of tumor cells.
Abstract The aim of the article is to test empirically predictions formulated in the Transitivity Hypothesis framework. Methodological problems of the original approach are discussed and some solutions are offered. For the testing of the hypotheses two corpora of Czech were used (Prague Spoken Corpus and Prague Dependency Treebank). The results question both the predicted impact of the language form on transitivity and, more importantly, the concept of the Transitivity Hypothesis in general.
The first task of statistical computational linguistics, or any other type of datadriven processing of language, is the extraction of counts and distributions of phenomena. This is much more difficult for the type of complex structured data found in treebanks and in corpora with sophisticated annotation than for tokenized texts. Recent developments in data mining, particularly in the extraction of frequent subtrees from treebanks, offer some solutions. We have applied a modified version of the TreeMiner algorithm to a small treebank and present some promising results.
We compare two processing methods for a single natural language processing task. One uses a treebank created with a full parser while the other restricts itself to lexical and part-of-speech information. We show that for the task under investigation, automatic extraction of hypernym-hyponym pairs from text, the former does not outperform the latter. We compare the output of the two approaches and look for an explanation for this unexpected result.
The paper deals with valency frames for selected group of Czech verbs belonging to the domain of Law. Starting with the lexical database VerbaLex we propose semantic roles for these verbs and formulate their Complex Valency Frames. The lexical database Verbalex has been developed recently at the NLP Centre FI MU and contains approx. 10 500 Czech verbs. We integrate the proposed 'law' valency frames into it.
It is a commonplace, by now, to refer to the recent explosive growth in the power and availability of computers as an information revolution. The most casual of computer users, linguists included, have at their fingertips an enormous amount of computing power. Tasks such as writing a document or playing
Abstract Recently, there has been a growing interest in regional variation within African American English. This study reviews a work done on local speech in Pittsburgh, Pennsylvania, discussing trends for both African American and White ethnic groups. Just as scholars have found in other geographic regions, in Pittsburgh, African Americans and Whites share a number of feature characteristics of the local dialect, but remain distinct in a number of other ways. Research in Pittsburgh, as elsewhere, highlights the complexity, rather than the homogeneity, of African American speech across the country, as speakers exhibit alignment to both regional and supraregional ethnic linguistic norms.
This paper describes the simultaneous development of dependency structure and phrase structure treebanks for Hindi and Urdu, as well as a PropBank. The dependency structure and the PropBank are manually annotated, and then the phrase structure treebank is produced automatically. To ensure successful conversion the development of the guidelines for all three representations are carefully coordinated.
We present a cost effective strategy for the creation of a mid-size fine-grained dependency treebank of surface- and deep-syntactic structures as defined in the Meaning-Text Theory for Spanish. The strategy starts from a small seed dependency corpus, the AnCora corpus, whose annotation is considerably more coarse-grained than our target annotation. We show that this discrepancy can be bridged largely by automatic means, relying upon contextual information and leaving thus minimal work to the annotators. This allows us to develop the resources with limited human effort within a limited period of time. We also propose a preliminary evaluation of the actual amount of work that the annotation process requires. 1
We consider linguistic database summaries in the sense of Yager (1982), in an implementable form proposed by Kacprzyk & Yager (2001) and Kacprzyk, Yager & Zadrozny (2000), exemplified by, for a personnel database, “most employees are young and well paid” (with some degree of truth) and their extensions as a very general tool for a human consistent summarization of large data sets. We advocate the use of the concept of a protoform (prototypical form), vividly advocated by Zadeh and shown by Kacprzyk & Zadrozny (2005) as a general form of a linguistic data summary. Then, we present an extension of our interactive approach to fuzzy linguistic summaries, based on fuzzy logic and fuzzy database queries with linguistic quantifiers. We show how fuzzy queries are related to linguistic summaries, and that one can introduce a hierarchy of protoforms, or abstract summaries in the sense of latest Zadeh’s (2002) ideas meant mainly for increasing deduction capabilities of search engines. We show an implementation for the summarization of Web server logs.
This article describes a method for calculating the 'dependency distance' between the words in a text – i.e. the number of words that separate each word from the word on which it depends syntactically – and reports the results of applying this method to a Chinese treebank. This study shows that Chinese dependencies tend strongly to be governor-final and that the mean dependency distance of words is much higher for Chinese than for other languages that have been studied including English, German and Japanese. It is unclear whether this difference means that Chinese is syntactically more difficult to process.
LTAG-spinal is a novel variant of traditional Lexicalized Tree Adjoining Grammar (LTAG) introduced by The LTAG-spinal Treebank (Shen et al., 2008) combines elementary trees extracted from the Penn Treebank with Propbank annotation. In this paper, we present a semantic role labeling (SRL) system based on this new resource and provide an experimental comparison with CCGBank and a state-of-the-art SRL system based on Treebank phrase-structure trees. Deep linguistic information such as predicateargument relationships that are either implicit or absent from the original Penn Treebank are made explicit and accessible in the LTAG-spinal Treebank, which we show to be a useful resource for semantic role labeling.
For Arabic, diacritizing written text is important for many NLP tasks. In the work presented here, we investigate the quality of a diacritization approach, with a high success rate for treebank data but with a more limited success on realworld data. One of the problems we encountered is the non-standard use of the hamza diacritic, which leads to a decrease in diacritization accuracy. If an automatic hamza restoration module precedes diacritization, the results improve from a word error rate of 9.20% to 7.38% in treebank data, and from 7.96% to 5.93% on selected real-world texts. This shows clearly that hamza restoration is a necessary step for improving diacritization quality for Arabic real-world texts.
This paper reports a sociolinguistic study of the state of Greek language in Australia as spoken by native-speaking Greek immigrants and their children. Emphasis is given to the analysis of the linguistic behaviour of these Greek Australians which are attributed to contact with English and to other environmental, social and linguistic influences. The paper discusses the non-standard phenomena in various types of inter-lingual transferences in terms of their incidence and causes and, in correlation with social, linguistic and psychological factors in order to determine the extent of language assimilation, attrition, and the content and context and medium of the language-event. The paper also discusses the transferences from English to Greek and vice- versa from a qualitative and quantitative perspective, of the phonemic, lexical, morphological, syntactic, semantic, pragmatic and prosodic deviations. During the last 170 years of settlement, Greek Australians know and use a new communicative norm with some degree of stability, the Ethnolect, (a non-standard variety of language used by an ethnic group in a static or dynamic bilingual situation) which serves their linguistic needs.
We describe a heuristics-based system for automatic measurement of syntactic complexity using the revised Developmental Level (D-Level) scale (Rosenberg & Abbeduto 1987; Covington et al. 2006). The system takes a raw sentence as input and assigns it to an appropriate developmental level on the scale. The system is designed with child language acquisition and psycholinguistic research in mind, and is therefore developed and evaluated using both written data from the Penn Treebank (Marcus et al. 1993) and spoken child language acquisition data from the CHILDES database (MacWhinney 2000). Experiment results show that the model achieves an accuracy of 94.0% and 93.2% on unseen test data from the Penn Treebank and the CHILDES database respectively. We illustrate how the system is used in an example application to investigate the correlation of average D-Level score and speaker age.
Older adults' relatively better memory for positive over negative material (positivity effect) has been widely observed in Western samples. This study examined whether a relative preference for positive over negative material is also observed in older Koreans. Younger and older Korean participants viewed images from the International Affective Picture System (IAPS), were tested for recall and recognition of the images, and rated the images for valence. Cultural differences in the valence ratings of images emerged. Once considered, the relative preference for positive over negative material in memory observed in older Koreans was indistinguishable from that observed previously in older Americans.
In this paper, we report a work in progress on transforming syntactic structures from the Syn-TagRus corpus into tectogrammatical trees in the Prague Dependency Treebank (PDT) style. SynTagRus (Russian) and PDT (Czech) are both dependency treebanks sharing lots of common features and facing similar linguistic challenges due to the close relatedness of the two languages. While in PDT the tectogrammatical representation exists, sentences in SynTagRus are annotated on syntactic level only. annotation layers: the morphological layer, the analytical layer (describing the surface syntax) and the tectogrammatical layer (describing the deep syntax – transition between syntax and semantics). A highly simplified example of the annotation layers is in Figure
In this paper, we present a discriminative word-character hybrid model for joint Chinese word segmentation and POS tagging. Our word-character hybrid model offers high performance since it can handle both known and unknown words. We describe our strategies that yield good balance for learning the characteristics of known and unknown words and propose an error-driven policy that delivers such balance by acquiring examples of unknown words from particular errors in a training corpus. We describe an efficient framework for training our model based on the Margin Infused Relaxed Algorithm (MIRA), evaluate our approach on the Penn Chinese Treebank, and show that it achieves superior performance compared to the state-of-the-art approaches reported in the literature.
Abstract Textual data is at the forefront of information management problems today. One response has been the development of visualizations of text data. These visualizations, commonly based on simple attributes such as relative word frequency, have become increasingly popular tools. We extend this direction, presenting the first visualization of document content which combines word frequency with the human‐created structure in lexical databases to create a visualization that also reflects semantic content. DocuBurst is a radial, space‐filling layout of hyponymy (the IS‐A relation), overlaid with occurrence counts of words in a document of interest to provide visual summaries at varying levels of granularity. Interactive document analysis is supported with geometric and semantic zoom, selectable focus on individual words, and linked access to source text.
Over the past 15 years, there has been increasing use of linguistically annotated sentence collections such as the LDC Penn Tree Bank (PTB) for constructing statistically based parsers. While these parsers have generally been built for engineering purposes, more recently such approaches have been advanced as potential cognitive solutions, e.g., for the problem of human language acquisition. Here we examine this possibility critically: we assess how well these Treebank parsers actually approach human/child language competence. We find that such systems fail to replicate many, perhaps most, empirically attested grammaticality judgments; seem overly sensitive, rather than robust, to training data idiosyncrasies; and easily acquire unnatural syntactic constructions never attested in human languages. Overall, we conclude that existing statistically based treebank parsers fail to incorporate much knowledge of language in these three senses. We discuss the implications of these results for the improvement of Treebank parsers and their cognitive relevance.
We present an implicit discourse relation classifier in the Penn Discourse Treebank (PDTB). Our classifier considers the context of the two arguments, word pair information, as well as the arguments' internal constituent and dependency parses. Our results on the PDTB yields a significant 14.1% improvement over the baseline. In our error analysis, we discuss four challenges in recognizing implicit relations in the PDTB.
In this paper we describe and evaluate a top-down transfer component of a hybrid example-based machine translation system with an architecture similar to that of transfer MT systems, but with automatically derived transfer-rules and dictionary entries based on a parallel treebank. The tests were applied on the translation pair Dutch to English. Evaluation and error analysis have shown that the top-down transfer process has a number of shortcomings on which we wish to report and which we will try to solve in future work by applying bottom-up transfer.
We consider linguistic database summaries in the sense of Yager (1982), in an implementable form proposed by Kacprzyk & Yager (2001) and Kacprzyk, Yager & Zadrozny (2000), exemplified by, for a personnel database, “most employees are young and well paid” (with some degree of truth) and their extensions as a very general tool for a human consistent summarization of large data sets. We advocate the use of the concept of a protoform (prototypical form), vividly advocated by Zadeh and shown by Kacprzyk & Zadrozny (2005) as a general form of a linguistic data summary. Then, we present an extension of our interactive approach to fuzzy linguistic summaries, based on fuzzy logic and fuzzy database queries with linguistic quantifiers. We show how fuzzy queries are related to linguistic summaries, and that one can introduce a hierarchy of protoforms, or abstract summaries in the sense of latest Zadeh’s (2002) ideas meant mainly for increasing deduction capabilities of search engines. We show an implementation for the summarization of Web server logs.
One of the benefits of incremental sentence production is reduction of the working memory capacity needed for advance planning: The planning units can be considerably smaller (measured in terms of word length) than in case of non-incremental production. The same advantage has been claimed for the various forms of ellipsis, which preempt the need to plan the detailed shape of one or more constituents and thereby reduce the size of planning units. Because working memory load tends to be higher in spoken than in written language, one expects that speakers, in comparison with writers, will more frequently resort to the use of elliptical constructions. However, in two corpus studies into the incidence of Clausal Coordinate Ellipsis (CCE) in spoken and written English, Meyer (1995) and Greenbaum &amp; Nelson (1999) obtained a data pattern opposite to this prediction: In written clausal coordinations, the proportion of CCE versions was about twice as high as in spoken coordinations. The pattern was explained in terms of audience design: Non-elliptical (unreduced) clauses include more repetition and thereby facilitate comprehension. Recent treebanks with large numbers of hand-parsed spoken (CGN2.0) and written (ALPINO) Dutch sentences, enabled us to verify the data pattern for another language: In written Dutch, the percentage of elliptical versions within the set of all clausal coordinations was even three times higher than in spoken Dutch: 34 % versus 11 % (Table 1).
This paper presents an on-going effort which aims to annotate the Wall Street Journal sections of the Penn Treebank with the help of a hand-written large-scale and wide-coverage grammar of English. In doing so, we are not only focusing on the various stages of the semi-automated annotation process we have adopted, but we are also showing that rich linguistic annotations, which can apart from syntax also incorporate semantics, ensure that the treebank is guaranteed to be a truly sharable, re-usable and multi-functional linguistic resource.
A polyadic dynamic logic is introduced in which a model-theoretic version of nonlocal multicomponent tree-adjoining grammar can be formulated.It is shown to have a low polynomial time model checking procedure.This means that treebanks for nonlocal MCTAG, incl.all weaker extensions of TAG, can be efficiently corrected and queried.Our result is extended to HPSG treebanks (with some qualifications).The model checking procedures can also be used in heuristics-based parsing.* The model checking procedure described in this paper uses constructs from a model checking procedure introduced in joint work with Martin Lange.Thanks also to Laura Kallmeyer, Timm Lichte and Wolfgang Maier for introducing me to various extensions of tree-adjoining grammar, incl.nonlocal MCTAG.
This paper proposes an approach to enhance dependency parsing in a language by using a translated treebank from another language. A simple statistical machine translation method, word-by-word decoding, where not a parallel corpus but a bilingual lexicon is necessary, is adopted for the treebank translation. Using an ensemble method, the key information extracted from word pairs with dependency relations in the translated text is effectively integrated into the parser for the target language. The proposed method is evaluated in English and Chinese treebanks. It is shown that a translated English treebank helps a Chinese parser obtain a state-of-the-art result.
Generative lexicalized parsing models, which are the mainstay for probabilistic parsing of English, do not perform as well when applied to languages with different language-specific properties such as free(r) word order or rich morphology. For German and other non-English languages, linguistically motivated complex treebank transformations have been shown to improve performance within the framework of PCFG parsing, while generative lexicalized models do not seem to be as easily adaptable to these languages.
English is spoken worldwide by both native (L1) and nonnative (L2) speakers. It is therefore imperative to establish how easily L1 and L2 speakers understand each other. We know that L1 listeners adapt to foreign-accented speech very rapidly (Clarke & Garrett, 2004), and L2 listeners find L2 speakers (from matched and mismatched L1 backgrounds) as intelligible as native speakers (Bent & Bradlow, 2003). But foreign-accented speech can deviate widely from L1 pronunciation norms, for example when adult L2 learners experience difficulties in producing L2 phonemes that are not part of their native repertoire (Strange, 1995). For instance, Italian L2 learners of English often lengthen the lax English vowel /I/, making it sound more like the tense vowel /i/ (Flege et al., 1999). This blurs the distinction between words such as bin and bean. Unless listeners are able to adapt to this kind of pronunciation variance, it would hinder word recognition by both L1 and L2 listeners (e.g., /bin/ could mean either bin or bean). In this study we investigate whether Italian-accented English interferes with on-line word recognition for native English listeners and for nonnative English listeners, both those where the L1 matches the speaker accent (i.e., Italian listeners) and those with an L1 mismatch (i.e., Dutch listeners). Second, we test whether there is perceptual adaptation to the Italian-accented speech during the experiment in each of the three listener groups. Participants in all groups took part in the same cross-modal priming experiment. They heard spoken primes and made lexical decisions to printed targets, presented at the acoustic offset of the prime. The primes, spoken by a native Italian, consisted of 80 English words, half with /I/ in their standard pronunciation but mispronounced with an /i/ (e.g., trick spoken as treek), and half with /i/ in their standard pronunciation and pronounced correctly (e.g., treat). These words also appeared as targets, following either a related prime (which was either identical, e.g., treat-treat, or mispronounced, e.g., treek-trick) or an unrelated prime. All three listener groups showed identity priming (i.e., faster decisions to treat after hearing treat than after an unrelated prime), both overall and in each of the two halves of the experiment. In addition, the Italian listeners showed mispronunciation priming (i.e., faster decisions to trick after hearing treek than after an unrelated prime) in both halves of the experiment, while the English and Dutch listeners showed mispronunciation priming only in the second half of the experiment. These results suggest that Italian listeners, prior to the experiment, have learned to deal with Italian-accented English, and that English and Dutch listeners, during the experiment, can rapidly adapt to Italian-accented English. For listeners already familiar with a particular accent (e.g., through their own pronunciation), it appears that they have already learned how to interpret words with mispronounced vowels. Listeners who are less familiar with a foreign accent can quickly adapt to the way a particular speaker with that accent talks, even if that speaker is not talking in the listeners’ native language.
The Turin University Treebank (TUT) is a treebank with dependency-based annotations of 2,400 Italian sentences. By converting TUT to binary constituency trees, it is possible to produce a treebank of derivations of Combinatory Categorial Grammar (CCG), with an algorithm that traverses a tree in a top-down manner, employing a stack to record argument structure, using Part of Speech tags to determine the lexical categories. This method reaches a coverage of 77%, resulting in a CCGbank for Italian comprising 1,837 sentences, with an average length of 22,9 tokens. The CCGbank for English has proven to be a useful tool for developing efficient wide-coverage parsers for semantic interpretation, and the Italian CCGbank is expected to be an equally useful linguistic resource for training statistical parsers.