Papers reviewed and determined not to be word norm studies. Use the flag icon to report errors or suggest re-inclusion.
18265 papers
This paper describes the mutually beneficial relationship between a cultural heritage digital library and a historical treebank: an established digital library can provide the resources and structure necessary for efficiently building a treebank, while a treebank, as a language resource, is a valuable tool for audiences traditionally served by such libraries. 1
In this paper, we describe the application of a bidirectional dependency parser trained on the Turin University Treebank.
The lexicon of the modern Norwegian bokmal standard needs a better description and documentation than what is the situation today. The article presents a plan for building a modern lexical database based on a balanced corpus of 40 million words of modern bokmal. This base should serve as a source for a traditional scientific dictionary as well as a dictionary for language technological applications.
782 SEER, 85, 4, OCTOBER 2007 television is largely Prague-based may have a more profound bearing on the use of language than has generally been appreciated. Not surprisingly, this study has many of the strengths and some of the weaknesses of a typical doctoral thesis. It offers a comprehensive summary and evaluation of existing research and provides very useful cross-references. It also highlights the complexity of language usage in a linguistic settingwhere stylisticallyand functionally divergent forms coexist, and where theprestigious 'standard' variant is not the spoken norm. Most importantly, it offers new statistical information to add to the existing body of data on morphological, phonological and lexical variation, and to substantiate claims that language choice always depends to a significant extent on the purpose of the dialogue and the formality of the situation.However, minor problems with editing and proof-reading detract from the overall quality of thework. Furthermore, the selection of television broadcasts inevitably contains a degree of subjectivity and is not indicative of the speech of the population as a whole. Finally, itwould appear that a lack of space may have prevented the author from developing some of hermore interesting ideas, such as the notion thatwomen may be treated differendy tomen in the television studio, and that this may be reflected in theiruse of language. In summary, despite some shortcomings, this is an original and stimulating study,which is of relevance to all scholars of language variation and change, and presents considerable scope for further research. School of Humanities, Languages and Social Sciences Tom Dickins Universityof Wolverhampton Pushkin, Alexander. 'TheGypsies' and Other NarrativePoems. Translated, with an introduction and notes, by Antony Wood. Engravings by Simon Brett. Angel Books, London, 2006. xl + 116pp. Notes.?14.95. Anyone who has ever attempted to translate nineteenth-century Russian verse into English should make a point of turning to the Afterword of Antony Wood's new book. Subtided 'Pushkin's Voice inEnglish', it is a pithy credo from one of the UK's leading translators of verse. One statement in particular should be writ large above any translator's desk: 'the rise of translation theory in recent decades has not been accompanied by the emergence of any sub stantial body of translation of Pushkin's verse that has impressed as verse in English' (p. 114). Wood makes clear how he intends to remedy thisdeficiency. To begin with he is uncontroversial. He will eschew alternating masculine and feminine rhymes as being too difficult to achieve inRussian. He will have recourse to half-rhymes, since rhymes are far easier to find in an inflected language than in an uninflected language. His other points, however, are more contentious. He is clearly no enthusiast for translations which reproduce exactly themetre and rhyme scheme of the original, considering that their effect 'tends to be self-conscious, self-satisfied,unengaged and disembodied' (p. in). Nor does he think that the number of lines of the original should necessarily be maintained. reviews 783 The Afterword is one of the items added toWood's earlier work 'The Bridegroom', with 'Count Nulin} and 'TheTale of the Golden CockereT,published by Angel Books in 2002 and reviewed in this journal (vol. 82, July 2004, no. 3). These three poems are reproduced here with slight amendments, one of which, fromThe Bridegroom,isparticularly felicitous.Whereas in 2002 we find in the tenth stanza of the poem 'and then a jet/Over Natasha's head', the revised translation reads 'then splash a/Dash of iton Natasha. The new book is some twice the length of the earlier book. The new translations are Pushkin's first 'problem' poema,The Gypsiesand the skazka,The Tale of the Dead Princess and theSevenChampions.True to his credo,Wood does not attempt to replicate Pushkin's iambic tetrameter throughout his transla tion of The Gypsies. His favoured departure from this involves removing the initial unstressed syllable and turning the line into trochaic tetrameter. There are numerous examples of the type 'Life resounds on every side' (p. 3 ). In addition there are variants of this variant, all scrupulously noted in the Afterword. These departures from Pushkin's metre are clearly no oversight and Wood shows considerable expertise in producing...
Traditional self-report measures suffer from weaknesses in either the quantitative or qualitative assessment of subjective experience. Researchers interested in the subjective intensity of oral sensation have attempted to reduce these scale limitations by developing rating scales with empirically determined placement of verbal descriptors along a continuous visual analogue scale continuum. In the present research, a similar empirical approach to scale construction was adopted to develop a rating scale of emotional valence. The potential benefits of using an empirically derived valence scale and techniques for validating the scale are discussed.
REVIEWS 781 Hedin, Tora. Changing Identities:Language Variation on Czech Television. Acta universitatis Stockholmeinsis. Stockholm Slavic Studies, 29. Stockholm University, Stockholm, 2005. xv + 219 pp. Tables. Figures. Illustrations. Notes. Bibliography. Appendix. Index. SEK 282.00 (paperback). This is a published version of the author's doctoral thesis,written inEnglish. The study examines indetail language variation inCzech television discourse, primarily between January 1997 and September 2003. It is a thoroughly researched and well informed work, with reference to all the important corpus-based studies of the spoken language? Kucera, Kravcisinova & Bednaf ova, and Bayer and Maglione, as well as toGammelgaard and Bermel (who both draw on dialogue in literary texts)? and to over seventy television broadcasts. The quantitative analysis is confined to a more manageable corpus of fifteen programmes, comprising 24,000 words, which are aimed at a range of different audiences and contain a mixture of prepared and unpre pared speech. Most of the examples cited have been well documented and analysed in other linguistic environments elsewhere, but comparatively little attention has hitherto been paid to themilieu of Czech television. The chapter on the historical background to theCzech language situation and the linguistic debate over the coexistence of Standard Czech and Common Czech (and their variant and intermediate forms) serves to contextualize the discussion of the empirical data. Chapter three provides a concise overview of language and the mass media, with some particularly helpful comments for the lesswell initiated on concepts such as the facade, the conversational framework and television as flow and image. It is to the author's advantage here that she is able to draw on a number of important studies in Scandinavian languages which are linguistically inaccessible tomost English- and Czech-speaking scholars. While the main part of the study employs a quantitative synchronic approach, the most engaging and thought provoking section of the book is arguably the qualitative diachronic analysis presented in chapter six. The explanation and exemplification of differences in language use in Czech television discourse before and after 1989 make for especially interesting reading and would provide the basis for a more detailed investigation in their own right.The final chapter on code-switching is similarly worthwhile, although this reviewer would have liked to see greater consideration given to the linguistic constraints on different types of code-mixing (both intrasentential and intersentential). The difficultyfor the author of thisbook is the sheer breadth of the subject matter and the range of the variables affecting television participants' choice of language variety.While a macrolinguistic study of this typemay provide a sound statistical basis for frequency-based analysis of usage, it can at best only offer tentative generalized suggestions as to the role played by the inter relationship between social, geographical, cultural and personal factors, variables such as age, professional status and gender, and the specific context of each television programme. It isperhaps a shame that the author did not correlate her main findings with the data from the Prague Spoken Corpus and the Brno Spoken Corpus of theCzech National Corpus with a view to defining more clearly any regional differences in usage. The fact thatCzech 782 SEER, 85, 4, OCTOBER 2007 television is largely Prague-based may have a more profound bearing on the use of language than has generally been appreciated. Not surprisingly, this study has many of the strengths and some of the weaknesses of a typical doctoral thesis. It offers a comprehensive summary and evaluation of existing research and provides very useful cross-references. It also highlights the complexity of language usage in a linguistic settingwhere stylisticallyand functionally divergent forms coexist, and where theprestigious 'standard' variant is not the spoken norm. Most importantly, it offers new statistical information to add to the existing body of data on morphological, phonological and lexical variation, and to substantiate claims that language choice always depends to a significant extent on the purpose of the dialogue and the formality of the situation.However, minor problems with editing and proof-reading detract from the overall quality of thework. Furthermore, the selection of television broadcasts inevitably contains a degree of subjectivity and is not indicative of the speech of the population as a whole. Finally, itwould appear that a...
XARA is a rule-based PropBank labeler for Alpino XML files, written in Java. I used XARA in my research on semantic role labeling in a Dutch corpus to bootstrap a dependency treebank with semantic roles. Rules in XARA are based on XPath expressions, which makes it a versatile tool that is applicable to other treebanks as well.
This paper aims to explore the norms, strategies and procedures of translating lexical doublets in Arabic literary discourse. Lexical doublets are sets of two (near-) synonyms connected with ﻮ ‘and’, ﺃﻮ ‘or’, or the zero article. The empirical basis material for this study consists of a three-part autobiography ( al-Ayyām, ‘The Days’) and a narrative ( Hadīth ´Īsā ibn Hishām, ‘´Īsā ibn Hishām’s Tale’). Findings show that patterns of repetition are shifted in the English translations, and various translation strategies are applied, the most common being grammatical transposition and reduction. A quantitative analysis of the translation of lexical doublets in three samples is also conducted. The samples are about 2500 words each, randomly selected from the three parts of the autobiography. The figures indicate that one translator (that of Part One) adopts a source text-oriented strategy while the other two translators prefer a shifting strategy. This may be seen as a useful indicator of the translations’ orientation towards either adequacy or acceptability (Toury 1995).
Proceedings of the Sixth International Workshop on Treebanks and \nLinguistic Theories. \nEditors: Koenraad De Smedt, Jan Hajič and Sandra Kübler. \nNEALT Proceedings Series, Vol. 1 (2007), 19-30. \n© 2007 The editors and contributors. \nPublished by \nNorthern European Association for Language \nTechnology (NEALT) \nhttp://omilia.uio.no/nealt. \nElectronically published at \nTartu University Library (Estonia) \nhttp://hdl.handle.net/10062/4476.
In morphologically rich languages, should morphological and syntactic disambiguation be treated sequentially or as a single problem? We describe several efficient, probabilisticallyinterpretable ways to apply joint inference to morphological and syntactic disambiguation using lattice parsing. Joint inference is shown to compare favorably to pipeline parsing methods across a variety of component models. State-of-the-art performance on Hebrew Treebank parsing is demonstrated using the new method. The benefits of joint inference are modest with the current component models, but appear to increase as components themselves improve. 1
This paper presents a new research and development project called Papillon [Planas00]. It is a French-Japanese cooperation between laboratories GETA/CLIPS (Grenoble, France) and NII (Tokyo, Japan). Its goal is to build a French-English-Japanese multilingual lexical database by using interlingual links and to extract from it digital bilingual French-Japanese and Japanese-French dictionaries. These dictionaries will be available under the terms of an open source license. This project, initiated by some computational linguists, aims at being useful and open to all those who are interested in Japanese and French. A seminar was organized on the 10-12 of August 2000 in Tokyo [Planas00]. It was devoted to discussions aiming at reaching a general consensus on the structure and content of the database, and to decide some technical aspects of database development, i.e. database configuration, contents of the entries and link between the entries. Introduction There are few French-Japanese usa...
Criteria for Manual Clustering of Verb Senses Cecily Jill Duffield, Jena D. Hwang, Susan Windisch Brown, Dmitriy Dligach, Sarah E. Vieweg, Jenny Davis, Martha Palmer ({cecily.duffield, hwangd, susan.brown, dmitriy.dligach, sarah.vieweg, jennifer.davis, martha.palmer}@colorado.edu) Departments of Linguistics and Computer Science University of Colorado Boulder, C0 80039-0295, USA Key words: word sense disambiguation, annotation, inter-annotator agreement, syntactic/semantic features Introduction Word sense ambiguity poses significant obstacles to accurate and efficient information extraction and automatic translation. Successful disambiguation of polysemous words in NLP applications depends on determining an appropriate level of granularity of sense distinctions, especially for verbs. WordNet, an important and widely used lexical resource, uses fine-grained distinctions that provide subtle information about the particular usages of various lexical items (Felbaum, 1998). When used as a resource for annotation of various genres of text, this fine level of granularity has not been conducive to high rates of inter-annotator agreement (ITA) or high automatic tagging performance. Annotation of verb senses as described by coarse-grained Proposition Bank framesets may result in higher ITA scores, but the blurring of distinctions between verb senses with similar argument structures may fail to alleviate the problems posed by ambiguity. Our goal in this project is to create verb sense distinctions at a middle level of granularity that allow us to capture as much information as possible from a lexical item while still attaining high ITA scores and high system performance in automatic sense disambiguation. We have demonstrated that clear sense distinctions improve annotator productivity and accuracy, which results in a corresponding improvement in system performance. Training on this new data, Chen, Schein, Ungar and Palmer, (2006) report 86.7% accuracy for verbs using a smoothed maximum entropy model and rich linguistic features (just over 70% for fine-grained senses). This paper focuses on the methodology used to create the sense groupings, with a particular emphasis on the types of features that are most accessible to human annotators who are not linguists. Various criteria are considered when disambiguating senses and creating sense groupings for the verbs, including frequent lexical usages and collocations, syntactic features and alternations, and semantic features, similarly to the groupings for Senseval2 (Palmer, Dang & Felbaum, 2007). Our highest priority is to create clear distinctions among sense groupings that will be easily understood by the annotators and consequently result in high rates of inter- annotator agreement. We have found that the most successful approach is to cluster senses intuitively on a verb-by-verb basis, distinguishing sense groupings with features that are easily grasped by all annotators. Such features include specific domain usages, as in legal, financial, and social uses (distinguishing two senses of integrate in, “Over two-thirds of the teachers report they integrated the arts into their subjects,” and “The movie is set in 1971, when TC Williams High integrated blacks and whites.”); a specific syntactic construction, such as a required locative prepositional phrase (distinguishing the sense of open in, “The master bedroom opens to a large terrace,” from that in “The door won't open.”); and the features of nominal arguments, such as an agentive subject (separating senses of indicate in “These symptoms indicate a serious illness,” and “He indicated the right road by nodding towards it.”) More theoretical features for distinguishing groupings have proven to be less successful. Annotators not familiar with linguistics were confused by concrete/abstract distinctions and such aspectual features as continuative or stative. Therefore, they are now rarely used to label sense groupings. Such concepts, when used, are more likely to be described in prose commentary. Certain compositional features of verbs such as manner and path have also proven to be confusing for annotators, and resulted in decreased annotator agreement. Verb sense groupings that do not receive high ITA scores in initial rounds of annotation are revised, often prioritizing the use of the more successful features illustrated above with examples. Acknowledgements We gratefully acknowledge the support of the National Science Foundation Grant NSF-0415923, Word Sense Disambiguation, and the Defense Advanced Research Projects Agency GALE program, Contract No. HR0011-06- C-0022, a subcontract from the BBN-AGILE Team. References Chen, J., Schein, A., Ungar, L., & Palmer, M. (2006). An empirical study of the behavior of word sense disambiguation. Proceedings of NAACL-HLT 2006. New York City, NY. Fellbaum, C. (Ed.). (1998). WordNet: An on-line lexical database and some of its applications. MIT Press, Cambridge, MA. Palmer, M., Dang, H.T., and Fellbaum, C. (to appear, 2007). Making fine-grained and coarse-grained sense distinctions, both manually and automatically. Journal of Natural Language Engineering
Application of egalitarian and prioritarian accounts of health resource allocation in low-income countries have both been criticized for implying distribution outcomes that allow decreasing/undermining health gains and for tolerating unacceptable standards of health care and health status that result from such allocation schemes. Insufficient health care and severe deprivation of health resources are difficult to accept even when justified by aggregative efficiency or legitimized by fair deliberative process in pursuing equality and priority oriented outcomes. I affirm the sufficientarian argument that, given extreme scarcity of public health resources in low-income countries, neither health status equality between populations nor priority for the worse off is normatively adequate. Nevertheless, the threshold norm alone need not be the sole consideration when a country's total health budget is extremely scarce. Threshold considerations are necessary in developing a theory of fair distribution of health resources that is sensitive to the lexically prior norm of sufficiency. Based on the intuition that shares must not be taken away from those who barely achieve a minimal level of health, I argue that assessments based on standards of minimal physical/mental health must be developed to evaluate the sufficiency of the total resources of health systems in low-income countries prior to pursuing equality, priority, and efficiency based resource allocation. I also begin to examine how threshold sensitive health resource assessment could be used in the Philippines.
This article is about analysis of data obtained in repeated measures designs in psycholinguistics and related disciplines with items (words) nested within treatment (5 type of words). Statistics tested in a series of computer simulations are:F 1,F 2,F 1 &F 2,F′, minF′, plus two decision procedures, the one suggested by Forster and Dickinson (1976) and one suggested by the authors of this article. The most common test statistic,F 1 &F 2, turns out to be wrong, but all alternative statistics suggested in the literature have problems too. The two decision procedures perform much better, especially the new one, because it systematically takes into account the subject by treatment interaction and the degree of word variability.
This paper addresses the question of how to obtain consistent semantic annotation on the basis of a set of noisy texts. Many potential real-world applications of semantic computing are faced with the need to handle texts which are not well-edited, and for which a resource-intensive treebanking effort is not feasible. Student-produced short answers contain many grammatical and lexical errors, making consistent annotation a challenge. Nevertheless, this paper demonstrates that semantic role annotation can be done in a consistent and useful manner even under these constraints.
The Penn Treebank does not annotate within base noun phrases (NPs), committing only to flat structures that ignore the complexity of English NPs. This means that tools trained on Treebank data cannot learn the correct internal structure of NPs. This paper details the process of adding gold-standard bracketing within each noun phrase in the Penn Treebank. We then examine the consistency and reliability of our annotations. Finally, we use this resource to determine NP structure using several statistical approaches, thus demonstrating the utility of the corpus. This adds detail to the Penn Treebank that is necessary for many NLP applications.
We modified the traditional (verbal) digit span task for administration via computer and the Internet. This online version collects data on the floor and ceiling of a subject’s span capacity, rather than generating a rough estimate of capacity based on a 50% success rate, as the traditional version does. We compared the two versions within adult subjects in two cohorts: college-age normal readers and college-age reading-disabled readers. To explore the reliability of the online version as a research tool, we employed the Bland-Altman approach to examine agreement between instruments. The online version yielded spans similar to those yielded by the traditional version, tending toward smaller values at the high end and larger values at the low end of span sizes, in the typical readers. It differentiated between the better and poorer readers reliably, and to the same extent as does the verbal version. The online version of the digit span task is comparable to the traditional version in assessing verbal span capacity; the code with which to implement the task is available at www.psychonomic.org/archive.
This work is about multimodal and expressive synthesis on virtual agents, based on the analysis of actions performed by human users. As input we consider the image sequence of the recorded human behavior. Computer vision and image processing techniques are incorporated in order to detect cues needed for expressivity features extraction. The multimodality of the approach lies in the fact that both facial and gestural aspects of the user’s behavior are analyzed and processed. The mimicry consists of perception, interpretation, planning and animation of the expressions shown by the human, resulting not in an exact duplicate rather than an expressive model of the user’s original behavior.
In this paper we discuss a spatiotemporal layer of annotation to be added to an existing (syntactic) treebank. Although our system, called MiniSTEx, was developed for Dutch, it will also work for other EU-languages. This may, however, ask for some adaptations to the database which is the centre of our system. Next to adaptations for other languages, we may need adaptations for specific situations, even when only one language is covered. 1
ion over Parts of Speech. Lexical units are grouped into frames irrespective of their parts of speech. This allows to easily map, e. g., two text fragments onto each other that carry essentially the same meaning, but where one is headed by a verb and the other by a noun, such as ‘A bought B’ vs. ‘(the) acquisition of B by A’. In GermaNet, this mapping requires additional knowledge in the form of derivation relations (see above). Semantic Role Labelling. By semantic role labelling (‘frame elements’), syntactic variations are abstracted over: As such syntactic variants are mapped onto the same FrameNet representations, no additional relabelling mechanism is required. Frame-to-Frame Relations. Frame-to-frame relations, as recorded in the FrameNet database, list correspondences between frame elements. ‘A sold B to C.’ can be directly mapped onto ‘C bought B from A.’: The frames COMMERCE_SELL and COMMERCE_BUY are properly related, as are the participating frame elements BUYER, SELLER and GOODS. For every subtree for which a FrameNet representation can be found (based on the lemma of the node and the argument realisation), the corresponding FrameNet labels will be added: The name of the frame is added as a supplementary label to the node corresponding to the frame evoking element; the edges are labelled with the corresponding frame elements. Thus, the subtree is annotated as representing an instance of the respective frame. An example is shown in fig. 5.11. Note that frame structures are not fully disambiguated. Syntactic differences are used for disambiguation. For example, the reflexive use of a verb may be associated with a different frame than the intransitive one. In these cases, disambiguation is done. In other cases no disambiguation is performed, for example, where the correct frame can only be identified through sortal preferences on arguments. 5.2. THE LINGUISTIC KNOWLEDGE-BASE 213 Frame information provides an additional level of normalisation: syntactically different realisations, as, e. g., occasioned by dative shift, will receive the same FrameNet representation. For ‘John gave the book to Mary.’ and ‘John gave Mary the book.’, a GIVING frame with the same frame elements is derived. In particular, ‘Mary’ is identified as the RECIPIENT in both cases. Frame Relations as Sources of Inferences. The FrameNet lexical database not only defines frames as abstract semantic predicates and frame elements as abstract semantics role labels, but also a hierarchy based upon different frameto-frame relations defined both between frames and frame elements. We translate frame-to-frame relations directly into relabelling relations with corresponding relevance values. We currently use all available FrameNet frame-to-frame relations, except for the SEE_ALSO relation, even though some of them are only rather vaguely defined (cf. 4.3). It is therefore not always possible to foresee whether or not using a frame-to-frame relation will or will not result in a valid inference relation. For example, when two words evoke the same frame, this does not mean that they stand in the classical synonymy relation: ‘Good’ and ‘bad’ both evoke the DESIRABILITY frame, even though they would be considered antonyms in terms of classical lexical relations. We have decided to exploit all frame-to-frame relations as sources of inferences. From the definition of the relations (cf. 4.3), we considered that this would in most cases produce interesting, if possibly sometimes unlikely inferences. We considered, however, that the additional step of answer checking should detect and properly mark those cases. In our experiments, we did not observe any serious problems with this approach (cf. 7.2.2.3). This may to a large extent be due to the limited overall current coverage of the frame lexicon that we use (cf. 4.3.3). We expect clearer definitions of the relations to emerge together with growing coverage; at some point, it may turn out to be advisable to remove all but some core relations from consideration. This is how the relations are currently utilised for inferences: ‘Same Frame’. This is not strictly a frame-to-frame relation: Two subtrees labelled with the same frame (and frame elements) match during direct answer matching, simply because the labels are identical. No additional inference rules are required. Inheritance. We treat inheritance like a classical hyponymy relation, that is, we use it to introduce inferences in both directions (cf. 3.5.2.3). 214 CHAPTER 5. MATCHING STRUCTURED REPRESENTATIONS
The dependency relation is the most essential ingredient in a dependency-based theory of syntax. This paper presents some statistical findings on the dependency relation extracted from a Chinese dependency treebank. A sentence in the proposed treebank can easily be converted into a SSyntS graph in Meaning-Text Theory. The statistics on the dependency relation show that modifiers make up 55% of all dependencies and actants have a lower proportion of 45%. The paper demonstrates it is possible to extract from the treebank active and passive valence information of a word (or word class). The paper gives a formula to calculate the mean dependency distance (MDD) for a specific type of dependency relation in a language and obtains MDD of all dependency types in Chinese. These figures show that some dependencies tend to be much farther apart than others, and demonstrate that dependency distance tends to minimization and different dependency types have varying preference on the direction of dependency.
The purpose of this chapter is to provide a two-dimensional approach to language documentation (Hi mmelman 1998). In addition to building a database, we also conducted a sociolinguistic survey des igned to document the state of health of a language in a particular spatio-temporal frame. Our goa l is to share our fieldwork experience of documenting Kavalan, a seriously endangered language in sou theastern Taiwan now spoken by fewer than just a few dozen speakers. We first discuss our field exp eriences in working with speakers of Kavalan in Sinshe village, the only significant Kavalan set tlement left in Taiwan, and the state of the Kavalan language, based in part on Huang and Cha ng’s (19 95) earlier sociolinguistic survey, and in part on a recent more in-depth village-wide survey of lan guage use in the community. Next, we introduce the NTU Corpus of Formosan Languages, part of which incorporates our corpus data in Kavalan. The NTU Corpus of Formosan Languages aims to establish a standard for the creation of linguistic corpus databases through the application of information technology to linguistic research. The creation of this linguistic database enables us both to preserve valuable linguistic data and to provide a systematic recording of these languages, for the benefit of future linguistic research.
We study the correlations in the connectivity patterns of large scale syntactic dependency networks. These networks are induced from treebanks: their vertices denote word forms which occur as nuclei of dependency trees. Their edges connect pairs of vertices if at least two instance nuclei of these vertices are linked in the dependency structure of a sentence. We examine the syntactic dependency networks of seven languages. In all these cases, we consistently obtain three findings. Firstly, clustering, i.e., the probability that two vertices which are linked to a common vertex are linked on their part, is much higher than expected by chance. Secondly, the mean clustering of vertices decreases with their degree — this finding suggests the presence of a hierarchical network organization. Thirdly, the mean degree of the nearest neighbors of a vertex x tends to decrease as the degree of x grows—this finding indicates disassortative mixing in the sense that links tend to connect vertices of dissimilar degrees. Our results indicate the existence of common patterns in the large scale organization of syntactic dependency networks.
We argue in favor of the the use of labeled directed graph to represent various types of linguistic structures, and illustrate how this allows one to view NLP tasks as graph transformations. We present a general method for learning such transformations from an annotated corpus and describe experiments with two applications of the method: identification of non-local depenencies (using Penn Treebank data) and semantic role labeling (using Proposition
Proceedings of the 16th Nordic Conference \nof Computational Linguistics NODALIDA-2007. \nEditors: Joakim Nivre, Heiki-Jaan Kaalep, Kadri Muischnek and Mare Koit. \nUniversity of Tartu, Tartu, 2007. \nISBN 978-9985-4-0513-0 (online) \nISBN 978-9985-4-0514-7 (CD-ROM) \npp. 152-159.
In this article, we introduce a new representation based on lattice theory for lexical data from a lexical-database embodying the frame-semantic approach to language description, FrameNet. We present proof of the abundance of Concept Lattices as proposed in Formal Concept Analysis both in the theory of frames and in its present-day incarnation, the FrameNet resource, by constructing several types of these. We further argue for the adequacy of such lattices in representing linguistic data with contributions that range from data-visualization to the fine-tuning of some frame-theoretical concepts. We argue finally that FrameNet is better thought of as being a lexical resource rather than an ontology, but we make the case throughout the article for Concept Lattices being a linguistically adequate, formally effective intermediate representations from which knowledge representation languages may draw knowledge-rich, linguistic facts from FrameNet at their convenience.
This is a study of German compound nouns, used metaphorically to refer to three life style manifestations characteristic of postmodern society: consumption, health and fitness orientation, and the pursuit of pleasure. The lexical items are classified according to the source domains of the respective metaphors, thus demonstrating the variety of perspectives that come into play and providing an insight into the ways people think as well as their attitudes and values with respect to the phenomena focussed on in the study. Common to all lexemes is an element of excess, which suggests that the phenomena in question are regarded as violations of societal norms.
This paper presents a semi-automatic approach for extraction of collocations from corpora which uses the results of Conceptual Vectors as a semantic filter. First, this method estimates the ability of each co-occurrence to be a collocation, using a statistical measure based on the fact that it occurs more often than by chance. Then the results are automatically filtered (with conceptual vectors) to retain only one given semantic kind of collocations. Finally we perform a new filtering based on manually entered data. Our evaluation on monolingual and bilingual experiments shows the interest to combine automatic extraction and manual intervention to extract collocations (to fill multilingual lexical databases). It proves especially that the use of conceptual vectors to filter the candidates allows us to increase the precision noticeably.
The article foregrounds some basic explanations about syntactic corpus annotation and the methodology of building treebanks. It draws attention to the linguistic theoretical models on the basis of which the majority of treebank annotation approaches are formed, and to the applicability of syntactic corpus annotation for the studies of descriptive and theoretical linguistics, and for natural language processing. In addition the article presents the results of three studies connected with multi-lingual inductive dependency parsing on the first treebank for the Slovene language, the Slovene Dependency Treebank.
A parallel treebank consists of syntactically annotated sentences in two or more languages, taken from translated documents. These parallel sentences are linked through alignment. This paper explores the use of word n-gram alignment, computed for statistical machine translation, to create syntactic phrase alignment. We achieve a weighted F0.5 -score of over 65%.
GeometryNet can be considered as a lexical database for geometric entities and concepts. The idea is borrowed from WordNet, a popular knowledge repository often used for natural language processing tasks in AI applications. The basic objective behind the construction of a GeometryNet is to analyze and understand the geometric problems and draw relevant geometric figures automatically. Initial emphasis is put on machine understanding of problem statements that involve geometric constructions. School level geometry problems practiced by students of age group 13-16 are targeted. This paper explains different aspects of a GeometryNet, issues behind its construction, and its possible use for machine understanding of geometric problem statements.
Proceedings of the Sixth International Workshop on Treebanks and \nLinguistic Theories. \nEditors: Koenraad De Smedt, Jan Hajič and Sandra Kübler. \nNEALT Proceedings Series, Vol. 1 (2007), 61-72. \n© 2007 The editors and contributors. \nPublished by \nNorthern European Association for Language \nTechnology (NEALT) \nhttp://omilia.uio.no/nealt. \nElectronically published at \nTartu University Library (Estonia) \nhttp://hdl.handle.net/10062/4476.
The aim of this paper is to investigate whether a treebank grammar can be used to automatically classify and annotate German phrases contained in a MT lexicon. Phrases from the lexicon appear in their citation form and may differ structurally from the phrase tokens found in the corpus. We describe the grammar extraction process for a formalism called Tree-Generating Binary Grammar and evaluate the performance of subsets of the obtained grammar on a set of four types of lexical phrases.
Our paper reports an attempt to apply an unsupervised clustering algorithm to a Hungarian treebank in order to obtain semantic verb classes. Starting from the hypothesis that semantic metapredicates underlie verbs' syntactic realization, we investigate how one can obtain semantically motivated verb classes by automatic means. The 150 most frequent Hungarian verbs were clustered on the basis of their complementation patterns, yielding a set of basic classes and hints about the features that determine verbal subcategorization. The resulting classes serve as a basis for the subsequent analysis of their alternation behavior.
This paper describes a tool for aligning and searching parallel treebanks. Such treebanks are a new type of parallel corpora that come with syntactic annotation on both languages plus sub-sentential alignment. Our tool allows the visualization of tree pairs and the comfortable annotation of word and phrase alignments. It also allows monolingual and bilingual searches including the specification of alignment constraints. We show that the TIGER-Search query language can easily be combined with such alignment constraints to obtain a powerful cross-lingual query language.
An approach for identifying the human source of a text by leveraging the significance of synonyms in language is presented. While others have attempted to identify authors in the past, they have focused on purely statistical approaches such as word length distribution, number of distinct words, and language models. We claim that an author's choice of synonyms is idiosyncratic and can be used in determining the identity of an author, which we demonstrate via our algorithm for recognizing authors. This algorithm uses synonym sets from the WordNet lexical database to give more weight to words that have many common synonyms. The results of this method applied to the task of identifying the authors of classic literature show that there is a correlation between an author's synonym choice and the author's identity. With this new author recognition technology, we may now explore new avenues of intelligent and meaningful interaction with users.
This article presents a transfer-based statistical model for Chinese to Taiwanese sign-language (TSL) translation. Two sets of probabilistic context-free grammars (PCFGs) are derived from a Chinese Treebank and a bilingual parallel corpus. In this approach, a three-stage translation model is proposed. First, the input Chinese sentence is parsed into possible phrase structure trees (PSTs) based on the Chinese PCFGs. Second, the Chinese PSTs are then transferred into TSL PSTs according to the transfer probabilities between the context-free grammar (CFG) rules of Chinese and TSL derived from the bilingual parallel corpus. Finally, the TSL PSTs are used to generate the possible translation results. The Viterbi algorithm is adopted to obtain the best translation result via the three-stage translation. For evaluation, three objective evaluation metrics including AER, Top-N, and BLUE and one subjective evaluation metric using MOS were used. Experimental results show that the proposed approach outperforms the IBM Model 3 in the task of Chinese to sign-language translation.