Papers reviewed and determined not to be word norm studies. Use the flag icon to report errors or suggest re-inclusion.
16504 papers
The dependency relation is the most essential ingredient in a dependency-based theory of syntax. This paper presents some statistical findings on the dependency relation extracted from a Chinese dependency treebank. A sentence in the proposed treebank can easily be converted into a SSyntS graph in Meaning-Text Theory. The statistics on the dependency relation show that modifiers make up 55% of all dependencies and actants have a lower proportion of 45%. The paper demonstrates it is possible to extract from the treebank active and passive valence information of a word (or word class). The paper gives a formula to calculate the mean dependency distance (MDD) for a specific type of dependency relation in a language and obtains MDD of all dependency types in Chinese. These figures show that some dependencies tend to be much farther apart than others, and demonstrate that dependency distance tends to minimization and different dependency types have varying preference on the direction of dependency.
Proceedings of the 16th Nordic Conference \nof Computational Linguistics NODALIDA-2007. \nEditors: Joakim Nivre, Heiki-Jaan Kaalep, Kadri Muischnek and Mare Koit. \nUniversity of Tartu, Tartu, 2007. \nISBN 978-9985-4-0513-0 (online) \nISBN 978-9985-4-0514-7 (CD-ROM) \npp. 81-88.
French expressions like FRUIT DE MER 'seafood', MONTER UN BATEAU [a qqn] 'to fool someone' and A PROPOS [de qqch] 'about something' are idioms. Set phrases like these are considered in Meaning-Text Theory (MTT) as lexical units and are therefore described by an autonomous lexical entry in the dictionary. However, since each phrase consists of more than one word-form, these lexical units display complex behavior which has not yet been rigorously described in the Meaning-Text lexicography. In this paper, we present a MTT treatment of idioms focusing on their presentation in the dictionary, and propose tools for constructing a lexical database of idioms, which can efficiently represent form flexibility and combinatorial particularities of French idioms.
An alternative to the view that during evolution the human brain became specialized to preferentially attend to threat-related stimuli is to assume that all classes of stimuli that have high biological significance are prioritized by the attention system. Newborns are highly biologically relevant stimuli for members of a species, as their survival is important for reproductive success. The authors examined whether the Kindchenschema (baby schema) as described by Lorenz (1943) captures attention in the dot probe task. The results confirm attentional capture by photos of human infants presented to the left visual field, suggesting right hemisphere advantage. The magnitude of the attentional modulation was highly correlated with subjective arousal ratings of the photos. The findings show that biologically significant positive stimuli are prioritized by the attention system.
About a very normative but not very popular field in linguistics: Spelling in school Spelling in french is a complex business and sometimes seen as the most arbitrary part of the linguistic norm. But beyond this characterization, it is a good indicator of the capabilities required by all sorts of learning, and by language-associated skills particularly. good spelling is valued by society and it is the schools' mission to teach it. this article comments on the results of measuring the spelling skills of 10 to 16 year-old pupils in the mandatory french school system at a twenty-year interval. Within a single generation, the level has markedly fallen, and the drop is imputable especially to the grammatical part of orthography (agreements, conjugations) that demands applying regular rules. it is likely that what accounts for the phenomenon are the transformations that have taken place in schooling, but also, more generally, the way society has evolved with relation to norms.
How far can we get with unsupervised parsing if we make our training corpus several orders of magnitude larger than has hitherto be attempted? We present a new algorithm for unsupervised parsing using an all-subtrees model, termed U-DOP*, which parses directly with packed forests of all binary trees. We train both on Penn’s WSJ data and on the (much larger) NANC corpus, showing that U-DOP * outperforms a treebank-PCFG on the standard WSJ test set. While U-DOP * performs worse than state-of-the-art supervised parsers on handannotated sentences, we show that the model outperforms supervised parsers when evaluated as a language model in syntax-based machine translation on Europarl. We argue that supervised parsers miss the fluidity between constituents and non-constituents and that in the field of syntax-based language modeling the end of supervised parsing has come in sight. 1
Multiobjective evolutionary algorithms (MOEA) are an effective tool for solving search and optimization problems containing several incommensurable and possibly conflicting objectives. Unfortunately, many MOEAs face difficulties in solving problems when the number of objectives increases. In this paper, we investigate the efficacy of spatially structured MOEAs for scalable multiobjective problems. The algorithm is an extension of the standard cellular evolutionary algorithm, where the population is mapped to nodes of alternative complex networks. A selection regime based on a non-dominance rating and a crowding mechanism guides the evolutionary trajectory and an ε-dominance external archive is used to maintain a spread of solutions across the Pareto-optimal front. An important outcome of this work is the classification of the network models based on their impact on convergence speed and solution quality as the number of objectives increases for a given problem.
A new method to solve Chinese text chunking was introduced as conditional random fields (CRF) model, by which Chinese text chunking transformed into labeling the words with their chunk tags and establishing a model for tagged corpus according to conditional random fields so as to predict the chunk tag of each word. An F1 score of 85.5% is achieved by using the evaluation dataset of Chinese treebank of Beijing university, and obviously better than those of hidden Markov model and maximum entropy Markov model. Experimental results show that conditional random fields model is an effective way on Chinese text chunking and the strict Independence hypothesis and the label bias problem are avoided.
Based on the language of 17th century Bosnian Franciscan literature and enriched with features of the Neo-Štokavian folklore koine, the language of the 18th century writers represents a consistent system. Although it was not subject to willful codification, the language of the 18th century writers has codification elements. This is primarily implied by functional distribution, pronounced independence from common speech, especially on the syntactic level, compulsory use for all users and specific prescriptiveness in grammar handbooks of the time.
An improved k-means clustering method is proposed to identify Chinese phrases with the purpose of avoiding data sparseness and taking think of the relationship of neighbor part of speech and the cohesion of all part of speeches within one phrase.The proposed method regards each phrase as a cluster whose kernel is headword,which richly used the constituent disciplinarian of one phrase.It also integrates supervised statistical method and unsupervised clustering method by setting the original center of each class according the data from small Chinese corpus,which not only improves the accuracy of clustering but also avoids data sparseness.Through testing on Chinese Penn Treebank, the F score of seven types of Chinese phrase achieves to 92.94%.So,it is effective for Chinese text chunking.
This paper presents a semi-automatic approach for extraction of collocations from corpora which uses the results of Conceptual Vectors as a semantic filter. First, this method estimates the ability of each co-occurrence to be a collocation, using a statistical measure based on the fact that it occurs more often than by chance. Then the results are automatically filtered (with conceptual vectors) to retain only one given semantic kind of collocations. Finally we perform a new filtering based on manually entered data. Our evaluation on monolingual and bilingual experiments shows the interest to combine automatic extraction and manual intervention to extract collocations (to fill multilingual lexical databases). It proves especially that the use of conceptual vectors to filter the candidates allows us to increase the precision noticeably.
Reinforcing value of a behavior refers to the motivation to engage in the behavior. A reinforcing behavior will support more work to obtain the behavior. Individual differences in the reinforcing value of physical activity predict the usual physical activity in children. Another factor that may influence physical activity is liking of physical activity. Liking or hedonics refers to an affective rating associated with the behavior, and people are more likely to engage in physical activities that they like than ones that they do not like. Liking correlates with physical activity in youth. Although the independence of reinforcing value and liking of physical activity has not yet been tested, the motivation to gain access to a behavior and liking for that behavior are likely different constructs. PURPOSE: To determine whether liking and relative reinforcing value (RRV) of physical activity independently predict time youth spend in moderate-to-vigorous physical activity (MVPA). METHODS: Boys (n = 21) and girls (n = 15) age 8 to 12 years were measured for height, weight, aerobic fitness, liking and RRV of physical activity, and minutes in MVPA using accelerometers. RESULTS: Using multiple regression to control for individual differences in age, sex, BMI percentile, aerobic fitness, and time the accelerometer was worn, liking (P < 0.05) and RRV (P < 0.01) of physical activity independently predicted time in MVPA. When using median splits of the RRV and liking data to form subject groups, the group of children with both a high liking and RRV of physical activity participated in greater (P < 0.05) minutes per week of MVPA (1340 + 72 min) than groups with high RRV-low liking (1040 + 95 min), low RRV-high liking (978 + 89 min), or low RRV-low liking (1007 + 70 min) of physical activity. CONCLUSIONS: RRV and liking of MVPA are separate constructs as they independently predict MVPA of children. Those children who find physical activity the most reinforcing and also have a high liking of physical activity engage in 33% more MVPA than children who either find physical activity highly reinforcing or have a high liking of physical activity. Interventions that concurrently increase the reinforcing value and liking of physical activity may be the most effective for increasing youth participation in free-living MVPA. Supported by NIH Grant RO1 HD42766.
An increased interest in body image (more specifically, becoming or staying thin) is a common trend that is increasing among children and adolescents. Previous research shows that the interest to become or remain thin peaks during early adolescence, particularly among females. PURPOSE: Because research on this topic is limited in younger populations, the purpose of this study was to examine the interest in weight control for a population of preadolescent children. METHODS: Subjects included 261 (122 female and 139 male) children from third (n=92), fourth (n=80), and fifth (n=89) grades. Average age of the participants was 9.5 years. The primary investigator met one-on-one with each child. Height was measured to the nearest centimeter and mass to the nearest 1/2 kilogram. Each child was asked if they judged themselves to be overweight (fat), underweight (skinny) or in-between. Also, each child was asked if they would like to lose weight, gain weight, or stay the same. Categorical data were separated through cross tabulation and significant differences were assessed with Chi-Square analyses. RESULTS: Average body mass index (BMI) for girls was 18.2 (just below the 75th percentile for age and gender) and for boys was 18.7 (the 75th percentile for age and gender). Thirty nine percent of all participants wanted to lose weight or remain thin. For boys and girls respectively, significant percentages 23.7% and 30.3% wanted to loose weight. Moreover, 33.8% and 43.4% of boys and girls respectively wanted to remain or become thin in spite of “in-between” self image ratings and normal BMIs (X2=19.742, p=0.001 and X2=16.418, p=0.003 for boys and girls respectively). No significant differences were noted between grades 3, 4, and 5 for boys or girls. CONCLUSIONS: The results of this study suggest that preadolescent children are focusing on body image and specifically on wanting to become or remain thin. This is especially of interest as these children had normal BMIs and generally viewed themselves as having “in-between” body images. The data also suggest that while interest in body image may peak in early adolescence, it clearly begins for both boys and girls in early prepubescent years.
this paper, is not a language-specific one, in spite of the fact that most of the languages have their own repository for marking the role of the addressee in communicative utterances. In our opinion this linguistic phenomenon needs its adequate treatment because of two main reasons: 1. The vocative is supposed to be present on two levels: syntax and pragmatics. Therefore it needs more elaborate interpretation on the interface side, which, in HPSG, is more developed for morphology /syntax and syntax/semantics than syntax/pragmatics; 2. It will be useful for HPSG-oriented implementations, especially treebanks and dialogue systems. The paper is structured as follows: in the next section the status of the vocative in Bulgarian is discussed. In section 3 we propose our ideas on a unified treatment of vocatives. In section 4 the HPSG model is given. Section 5 outlines the conclusions and future work. 2 The Status of the Vocative in Bulgarian Vocatives are assumed to be restricted to the second person usage only. Usually they subsume the following two subtypes: calls (hey you) and addresses (Madam) [Levinson 1987, p. 71]. Bulgarian vocative role is usually treated within the opposition: vocative form (a remnant of the case paradigm) vs. base nominative form, i.e. with respect to the presence or loss of the special vocative inflections. Hence, The work reported here is done within the BulTreeBank project. The project is funded by the Volkswagen Stiftung, Federal Republic of Germany under the Programme &quot;Cooperation with Natural and Engineering Scientists in Central and Eastern Europe&quot; contract I/76 887. The authors wish to thank the Seminar fur Sprachwissenschaft of the Eberhard-Karls-Universitat, Tubingen, for hosting the writing of this paper, and the Internationales Zentr...
Iraj Mirza’s poetry occupies a special place in Persian literature as compared to the works of other poets of his time, due to his almost unrivalled use of language and rhetorics. Deviating from syntactic, semantic and pragmatic norms, he creates a new atmosphere with the simple language he uses, which draws his poetry close to the language of nature. The present paper examines Iraj Mirza’s poetry in terms of language function and his expert play with language. He deviates from the accepted linguistic norms of syntax, semantics and pragmatics, with an artistic courage, creating a new atmosphere in literary language: It is worth mentioning that he does so with such a simple language that one can claim, without unnecessary exaggeration that his poems are closer to the language of the nature than those of his contemporary poets. This article studies some of the language functions of Iraj's poems, revealing a small part of his skill in playing with the language.
The PARC 700 dependency bank is a potentially very useful resource for parser evaluation that has, so to speak, a high barrier to entry, because of tokenisation that is quite different from the source of the data, the Penn Treebank, and because there is no representation of word order, producing an uncertainty factor of some 15%. There is also a small, but perhaps not insignificant, number of errors. When using the dependency bank for evaluation, it seems likely that these things will cause inflated counts for mismatches, so to obtain more accurate measurements, it is desirable to eliminate them. The work reported here consists of an automatic conversion of the dependency bank into a Prolog representation where the word order is explicit, as well as graphical representations of the dependency trees for all 700 sentences, automatically generated from the Prolog data. As a side effect of the transformation, errors were detected and corrected. It is hoped that this work will lead to more widespread use of the PARC 700 dependency bank for parser evaluation.
In this paper,we propose a SVM-combined generative statistical model for Chinese dependency analysis that trains SVM classifier using erroneous results generated by generative statistical model.To further improve the precision of dependency analysis,two measures were taken,first,dynamic programming algorithm that extends the range of finding the best local solution was used to estimate the error rate of generative model;second,a ranging factor was introduced to make the solutions adaptive on the practical situation.All those efforts make it possible for the new method to largely decrease the number of negative support vectors without sacrificing classification ability in training.Comparative experiments on Hit Chinese Treebank corpus show that the new method shows better performance than current Chinese dependency methods,with precision reaching to 86.4%.
We study the correlations in the connectivity patterns of large scale syntactic dependency networks. These networks are induced from treebanks: their vertices denote word forms which occur as nuclei of dependency trees. Their edges connect pairs of vertices if at least two instance nuclei of these vertices are linked in the dependency structure of a sentence. We examine the syntactic dependency networks of seven languages. In all these cases, we consistently obtain three findings. Firstly, clustering, i.e., the probability that two vertices which are linked to a common vertex are linked on their part, is much higher than expected by chance. Secondly, the mean clustering of vertices decreases with their degree — this finding suggests the presence of a hierarchical network organization. Thirdly, the mean degree of the nearest neighbors of a vertex x tends to decrease as the degree of x grows—this finding indicates disassortative mixing in the sense that links tend to connect vertices of dissimilar degrees. Our results indicate the existence of common patterns in the large scale organization of syntactic dependency networks.
CCGbank is an automatic conversion of the Penn Treebank to Combinatory Categorial Grammar (CCG). We present two extensions to CCGbank which involve manipulating its derivation and category structure. We discuss approaches for the automatic re-insertion of removed quote symbols and evaluate their impact on the performance of the C&C CCG parser. We also analyse CCGbank to extract a multi-modal CCG lexicon, which will allow the removal of hardcoded language-specific constraints from the C&C parser, granting benefits to parsing speed and accuracy.
Functional Arabic Morphology is a formulation of the Arabic inflectional system seeking the working interface between morphology and syntax. ElixirFM is its high-level implementation that reuses and extends the Functional Morphology library for Haskell. Inflection and derivation are modeled in terms of paradigms, grammatical categories, lexemes and word classes. The computation of analysis or generation is conceptually distinguished from the general-purpose linguistic model. The lexicon of ElixirFM is designed with respect to abstraction, yet is no more complicated than printed dictionaries. It is derived from the open-source Buckwalter lexicon and is enhanced with information sourcing from the syntactic annotations of the Prague Arabic Dependency Treebank.
This paper tests three factors that have been held to be responsible for the variable stress behavior of noun-noun constructs in English: argument structure, semantics, and analogy. In a large-scale investigation of some 4500 compounds extracted from the CELEX lexical database (Baayen et al. 1995), we show that traditional claims about noun-noun stress cannot be upheld. Argument structure plays a role only with synthetic compounds ending in the agentive suffix - er. The semantic categories and relations assumed in the literature to trigger rightward stress do not show the expected effects. As an alternative to the rule-based approaches, the data were modeled computationally and probabilistically using a memory-based analogical algorithm (TiMBL 5.1) and logistic regression, respectively. It turns out that probabilistic models and the analogical algorithm are more successful in predicting stress assignment correctly than any of the rules proposed in the literature. Furthermore, the results of the analogical modeling suggest that the left and right constituent are the most important factor in compound stress assignment. This is in line with recent findings on the semi-regular behavior of compounds in other languages.
This paper investigates how the use of machine learning techniques can significantly predict the three major dimensions of learner-s emotions (pleasure, arousal and dominance) from brainwaves. This study has adopted an experimentation in which participants were exposed to a set of pictures from the International Affective Picture System (IAPS) while their electrical brain activity was recorded with an electroencephalogram (EEG). The pictures were already rated in a previous study via the affective rating system Self-Assessment Manikin (SAM) to assess the three dimensions of pleasure, arousal, and dominance. For each picture, we took the mean of these values for all subjects used in this previous study and associated them to the recorded brainwaves of the participants in our study. Correlation and regression analyses confirmed the hypothesis that brainwave measures could significantly predict emotional dimensions. This can be very useful in the case of impassive, taciturn or disabled learners. Standard classification techniques were used to assess the reliability of the automatic detection of learners- three major dimensions from the brainwaves. We discuss the results and the pertinence of such a method to assess learner-s emotions and integrate it into a brainwavesensing Intelligent Tutoring System.
Deterministic dependency parsing has often been regarded as an efficient algorithm while its parsing accuracy is a little lower than the best results reported by more complex methods. In this paper, we compare deterministic dependency parsers with complex parsing methods such as generative and discriminative parsers on the standard data set of Penn Chinese Treebank. The results show that, for Chinese dependency parsing, deterministic parsers outperform generative and discriminative parsers. Furthermore, basing on the observation that deterministic parsing is a greedy algorithm which chooses the most probable parsing action at every step, we propose three kinds of ungreedy deterministic dependency parsing algorithms to globally model parsing actions. We take the original deterministic parsers as baseline systems. Results show that ungreedy deterministic dependency parsers perform better than the baseline systems while maintaining the same time complexity, and our best result improve much over baseline.
Because of the wide variety of contemporary practices used in the automatic syntactic parsing of natural languages, it has become necessary to analyze and evaluate the strengths and weaknesses of different approaches. This research is all the more necessary because there are currently no genre- and domain-independent parsers that are able to analyze unrestricted text with 100% preciseness (I use this term to refer to the correctness of analyses assigned by a parser). All these factors create a need for methods and resources that can be used to evaluate and compare parsing systems. This research describes: (1) A theoretical analysis of current achievements in parsing and parser evaluation. (2) A framework (called FEPa) that can be used to carry out practical parser evaluations and comparisons. (3) A set of new evaluation resources: FiEval is a Finnish treebank under construction, and MGTS and RobSet are parser evaluation resources in English. (4) The results of experiments in which the developed evaluation framework and the two resources for English were used for evaluating a set of selected parsers.
Reviewed by: Indian and British English: A handbook of usage and pronunciationby Paroo Nihalni, R. K. Tongue, Priya Hosali, and Jonathan Crowther Niladri Sekhar Dash Indian and British English: A handbook of usage and pronunciation. 2ndedn. By Paroo Nihalni, R. K. Tongue, Priya Hosali, and Jonathan Crowther. New Delhi: Oxford University Press, 2004. Pp. x, 260. ISBN 0195666569. $15.95. The present handbook is divided into two main parts. The first part (‘Lexicon of usage’) is designed to provide English users with information about the way in which certain words, idioms, collocations, phrases, and similar expressions of English used in India differ from British Standard English (BSE)—a model that has the closest affinity to Indian English. This part includes a thousand English words, which are used in a distinctive manner by large numbers of educated Indian speakers of English irrespective of their place, profession, education, gender, or other sociolinguistic factors. The words included in the handbook are selected from the speech or writing samples of the persons (such as university and school teachers, journalists, and radio commentators) who are likely to influence the English use of Indian learners. The handbook also contains many European words that have been Indianized over the years. Thus, it serves as a handy resource for Indian speakers of English, illustrating the many, often quite subtle, ways in which Indian English differs from standard British English usage, and where these differences are regarded as acceptable or substandard in the subcontinent. Examples in the handbook, which supplement the texts, are helpful to Indian users of English who are uncertain about the ‘correctness’ of their speech and writing, and serve those scholars who want to explore the differences between Indian and British uses of English. The book also has the potential to address special problems faced by learners of English, who are often impeded by the difficulties of recognizing finer nuances of meaning and usage. The second part of the handbook includes a brief report on the development of the pronunciation dictionary in India and abroad, followed by insightful discussions on standards of pronunciation in second/ foreign language teaching, the phonological systems of the British Received Pronunciation (BRP) and Educated Indian English (EIE), and the role of supraseg-mental properties (i.e. word stress, sentence stress, rhythm, intonation, etc.) in Indian English. The introduction contains a list of keywords for phonetic symbols used in the following part, ‘Dictionary of pronunciation’. Two types of pronunciation (Indian Recommended Pronunciation and the BRP) are supplied for more than two thousand words collected from the original lexical database of Michael West’s General service list of English wordstogether with a few additions compiled from the language resources available to the compilers. Each entry of the dictionary is tagged with relevant phonological information. This second edition also includes additional information on lexical collocation (the tendency of words to be used together in fixed phrases). In essence, the handbook not only serves as an invaluable reference guide for students and teachers of English, but also makes a valuable contribution for applied linguists, lexicographers, journalists, and scholars who write in Indian English. [End Page 465] Niladri Sekhar Dash Indian Statistical Institute, Kolkata Copyright © 2007 Linguistic Society of America
Digital Humanities (DH) 2006 was the first conference hosted by the newly constituted Association of Digital Humanities Organizations (ADHO). It represents a continuation of the ALLC and ACH joint conferences, and this continuation can be seen in the topics discussed in the various papers, panels and posters. In February 1989 readers of the Humanist Mailing List, then into its second year, and then as now edited by Willard McCarty, could read the following list of topics that would be presented in the first joint ACH/ALLC conference in Toronto later that year, which included: archaeology; lexical databases; authorship attribution; manuscript bibliographies; computational linguistics; music; humanistic research; national research funding; computer-assisted learning; content analysis; narrative analysis; databases; scanning; discourse analysis; stylistics; editorial problems; text archives; the French novel; funding issues; text encoding; hypertext.
Abstract. Despite their widespread use in Natural Language Processing applications, lexical databases and wordnets in particular do not yet contribute satisfactorily to the difficult problem of automatic word sense discrimination. Having built a number of lexical databases ourselves, we are keenly aware of still unresolved fundamental theoretical issues. In this paper we examine some of these questions and suggests preliminary answers concerning the nature of lexical elements and the conceptualsemantic and lexical relations that interconnect them. Our perspective is multilingual, and our goal is to formulate a proposal for a “Global Wordnet Grid ” that will meet the challenge of mapping the lexicons of many languages in interesting and useful ways. 1
The ability to detect similarity in conjunct heads is potentially a useful tool in helping to disambiguate coordination structures - a difficult task for parsers. We propose a distributional measure of similarity designed for such a task. We then compare several different measures of word similarity by testing whether they can empirically detect similarity in the head nouns of noun phrase conjuncts in the Wall Street Journal (WSJ) treebank. We demonstrate that several measures of word similarity can successfully detect conjunct head similarity and suggest that the measure proposed in this paper is the most appropriate for this task.
We compare the accuracy of a statistical parse ranking model trained from a fully-annotated portion of the Susanne treebank with one trained from unlabeled partially-bracketed sentences derived from this treebank and from the Penn Treebank. We demonstrate that confidence-based semi-supervised techniques similar to self-training outperform expectation maximization when both are constrained by partial bracketing. Both methods based on partially-bracketed training data outperform the fully supervised technique, and both can, in principle, be applied to any statistical parser whose output is consistent with such partial-bracketing. We also explore tuning the model to a different domain and the effect of in-domain data in the semi-supervised training processes.
Collocation is of great importance in dictionary compilation and natural language processing.Collocation extraction is one of the principal applications of corpus linguistics.Automatic extraction of bi-grams as candidate collocations is studied on Penn Treebank using the criteria of log likelihood,chi square and mutual information as association measure.The experimental results show the feasibility of the statistical methods.On the other hand,collocations extracted show different characteristics because of the different distribution assumptions by the three criteria.
Proceedings of the Sixth International Workshop on Treebanks and \nLinguistic Theories. \nEditors: Koenraad De Smedt, Jan Hajič and Sandra Kübler. \nNEALT Proceedings Series, Vol. 1 (2007), 31-42. \n© 2007 The editors and contributors. \nPublished by \nNorthern European Association for Language \nTechnology (NEALT) \nhttp://omilia.uio.no/nealt. \nElectronically published at \nTartu University Library (Estonia) \nhttp://hdl.handle.net/10062/4476.
One of the goals of natural language processing (NLP) systems is determining the meaning of what is being transmitted. Although much work has been accomplished in traditional written and spoken language domains, little has been performed in the newer computer-mediated communication domain enabled by the Internet, to include text-based chat. This is due in part to the fact that there are no annotated chat corpora available to the broader research community. The purpose of our research is to build a chat corpus, initially tagged with lexical and discourse information. Such a corpus could be used to develop stochastic NLP applications that perform tasks such as conversation thread topic detection, author profiling, entity identification, and social network analysis. During the course of our research, we preserved 477,835 chat posts and associated user profiles in an XML format for future investigation. We privacy-masked 10,567 of those posts and part-of-speech tagged a total of 45,068 tokens. Using the Penn Treebank and annotated chat data, we achieved part-of-speech tagging accuracy of 90.8%. We also annotated each of the privacy-masked corpus's 10,567 posts with a chat dialog act. Using a neural network with 23 input features, we achieved 83.2% dialog act classification accuracy.
As the Introduction to this volume observes, sixteenth-century France is marked by ‘une vaste réflexion sur le bien dire’. This not only impacted upon theory and practice across the different literary genres using French but also promoted considerable debate on the form and basis of the emerging standard form of the vernacular. Given the wide-ranging nature of this réflexion, covering its different manifestations in one volume poses an almost insuperable challenge, but the twenty-eight contributions to the colloquium collected here certainly address an ambitiously broad span of topics and add usefully to our understanding of cultural developments in a period of major change. The papers are organized into three general sub-sections, ‘Interroger la norme’, ‘Évolutions de la norme’ and ‘Normes et société’. However, such is the fluid nature of the subject matter treated in certain papers that their classification under one or other of these headings can sometimes seem of doubtful appropriateness. The focus of the first sub-section is predominantly literary. The various contributions address the creation or adaptation of norms across a considerable number of different genres some of which are perhaps rather less familiar, for instance, oracular writings (Dubois), accounts of pilgrimages (Gomez-Géraud) and Jesuit letter-writing (Laborie). Particularly interesting is the close study by Duché of the influential approach which Nicolas Herberay adopted for translation. Herberay, an acknowledged master of French prose (‘un vray Cicero françois’, according to Jean Martin), wrote with a female as well as a male readership in mind, developing a prose style that was eloquent and natural that would set an example for bien dire in this area. The second sub-section begins with a cogent overview (Baddeley) of a familiar field, developments in orthography and the interplay between orthography and spelling, and is followed by a series of studies which explore revealingly topics such as the evolving relationship between poetics and grammar (Monferran), the increasing limitation on the use of metaphor in literary works (Cernogora) and developments in historiography (Dumontet). Perhaps the most interesting paper is the examination of the fortunes of the alexandrine in the early part of the century (Halévy). Particular attention is given to the writings of Jean Lemaire de Belges and Geoffroy Tory both of whom, on the basis of fanciful argumentation, sought to invest the alexandrine with special prestige and nationalistic symbolism matching the terza rima in Italian. Their exercises in myth-making were to contribute indirectly, it is argued, to the rapid rise in the alexandrine's use from around 1555. The final sub-section of the volume contains contributions that more particularly address linguistic issues. Notable amongst these are two items: a re-evaluation of the system of vers mesurés devised by Baïf which is seen as an attempt not only to reproduce the metrical patterns of ancient Greek but also to contribute towards the norms of spoken French by reflecting the élite ‘usage des Bons’ (Vignes); and a meticulous examination by Morin of change in the pronunciation norms presented by Peletier du Mans in his earlier works (1550, 1555) as against his 1581 Euvres poëtiques, the new norm correlating with that presented later in the works of Lanoue (1596) and La Touche (1696). Alongside these are a number of other attractive essays including a study of the linguistic norms in the speeches made at the formal opening of the Paris Parlement, with eloquence and high rhetoric dominating over practicality and clarity between 1560 and 1600 before a reversal occurred in the early seventeenth century (Petey-Girard), and an investigation of sixteenth-century liminaires (any text preceding a written work) composed by women where a complex set of norms operate involving humility, simplicity of style, the practice of dedicating the work to another woman and, in the light of the lack of image for the female writer, an attempt to ‘socialiser l'auteur’ (Gauthier). Completing the text is an Index Nominum and a table of contents. The diversity and scholarly depth of the volume should ensure that all seiziémistes will derive benefit from a close reading.
Nietzsche writes that literary styles must be taken into account when determining meaning. To not do so vulgarizes language. The importance of style to meaning is best understood by examining the relationship between language and experience. Style requires readers to experience the particularity of texts if they are to understand their meaning, and style is experienced and evaluated in terms of the audience’s ethos. Through style’s ability to provide alternative perspectives on the audience’s linguistic norms, language is not limited to reiterating common generalisations but expresses the unfamiliar, rare and evolving.
To date, work on Non-Local Dependencies (NLDs) has focused almost exclusively on English and it is an open research question how well these approaches migrate to other languages. This paper surveys non-local dependency constructions in Chinese as represented in the Penn Chinese Treebank (CTB) and provides an approach for generating proper predicate-argument-modifier structures including NLDs from surface contextfree phrase structure trees. Our approach recovers non-local dependencies at the level of Lexical-Functional Grammar f-structures, using automatically acquired subcategorisation frames and f-structure paths linking antecedents and traces in NLDs. Currently our algorithm achieves 92.2 % f-score for trace insertion and 84.3 % for antecedent recovery evaluating on gold-standard CTB trees, and 64.7 % and 54.7%, respectively, on CTBtrained state-of-the-art parser output trees. 1
In this report an unsupervised and knowledge-based algorithm for concept sense disambiguation in concept maps is proposed. Concept maps are graphical tools for organizing and representing knowledge, based on concepts and labeled interconnections among them, forming propositions. The disambiguation process is carried combining Magnini’s domain, context information and the gloss. It’s supported in the Spanish WordNet lexical database and the lexical relations hypernyms-hyponyms, meronyms-holonyms and instance.
In order to determine novel information from raw text documents, a novelty detection recommender system was developed to explore the method of comparing various types of entities within sentences. We first detected novel sentences using named entity recognition to extract the entity types of person, place, time, and organization. In addition, part-of-speech tagging was performed to tag each word in the documents, allowing syntactic structures of noun, verb, and adjective to be used for comparisons. WordNet, an English lexical database of concepts and relations, was also incorporated to generate synonyms for the entities and parts of speech, as well as to determine the similarity of sentences. The novelty score of each sentence was determined by using two different metrics, UniqueComparison and ImportanceValue. UniqueComparison calculated the number of matched entities, whereas ImportanceValue took into account the total weight of matched words that coexisted in both the test and history sentences. The results look promising when compared to the benchmark scores for the Text Retrieval Conference’s (TREC) Novelty Track 2004. This demonstrated that the combination of named entity recognition and part-of-speech tagging is capable of detecting novelty with good results.
Background Irritable bowel syndrome (IBS) is conceptualized as a syndrome of enhanced central stress circuit responsiveness, likely associated with altered adrenergic and autonomic responses. Aims (1) To determine if baseline autonomic nervous system (ANS) differences exist in IBS versus control subjects; (2) to determine group differences in ANS response to yohimbine (YOH) and clonidine (CLO); and (3) to determine group differences in the effects of YOH and CLO on affective states. Methods IBS and control subjects were enrolled. ANS and affective measures were taken before and after drug ingestion. ANS was measured by HRV (high frequency, HF, indicating cardiovagal and low/high frequency ratio, LF/HF, indicating sympathetic) and systolic BP. Affective ratings were made with the Stress Symptom Rating Questionnaire. Results Mean baseline HF was lower in IBS versus controls (34.29 nu vs 62.88 nu, p =.013). Mean baseline LF/HF was higher in IBS versus controls (2.59 vs 0.91, p =.057). No significant HRV changes were seen in response to CLO or YOH in either group. YOH significantly increased BP (p =.01) and CLO significantly reduced BP (p =.02) in pooled subjects; no group difference in BP was seen. No group differences in affective ratings were seen. In combined groups, YOH increased anxiety (p =.05) and CLO led to increased fatigue (p = 0.01) with decreased arousal (p =.02). Conclusion These findings confirm that there is greater baseline sympathetic and lower parasympathetic activity in IBS patients compared with controls. A larger sample of patients is likely needed to elucidate group differences in drug responses.
BACKGROUND: The aim of this study was to assess the accuracy of visual image rating as compared to parametric analysis of regional cerebral blood flow (rCBF) measured with SPECT in patients referred to a memory clinic for diagnostic evaluation of cognitive symptoms. METHODS: SPECT with (99m)Tc-HMPAO was used to determine rCBF in 47 patients and 26 healthy control subjects. The 47 patients (30 F/17 M) had a mean age of 74.6 years (range = 62-88) and mild or questionable dementia with an MMSE score of 24.8 (range = 20-30). Two experienced image readers blinded to the classifications and identity of subjects performed visual rating in consensus and the global and regional CBF patterns were evaluated and graded according to severity of hypoperfusion. Correlation coefficients were calculated using results from the parametric analyses as gold standard. RESULTS: The sensitivity and specificity of global visual rating (normal vs. abnormal SPECT) was 92 and 86%, respectively, yielding an overall accuracy of 89% for visual rating compared to parametric analysis. The correlation between visual rating and parametric analysis was highly significant (p < 0.001). CONCLUSION: Visual rating is a valid method for analyzing SPECT images in patients with mild or questionable dementia.
Interactionists interested in second language acquisition postulate that learners’ competences are sensitive to the context in which they are put into play. Here we explore the language practices displayed, in a bilingual socio-educational milieu, by three dyads of English learners while carrying out oral communicative pair-work. In particular, we examine the role language choice plays in each task. A first analysis of our data indicates that the learners’ language choices seem to reveal the linguistic norms operating in the community of practice they belong to. A second analysis reveals that they exploited their linguistic repertoires according to their interpretation of the task and to their willingness to complete it in English. Thus, in the first two tasks students relied on code-switching as a mechanism to solve communication failures, whereas the third task generated the use of a mixed repertoire as a means to complete the task in the target language.
We consider the problem of efficiently storing n-gram counts for large n over very large corpora. In such cases, the efficient storage of sufficient statistics can have a dramatic impact on system performance. One popular model for storing such data derived from tabular data sets with many attributes is the ADtree. Here, we adapt the ADtree to benefit from the sequential structure of corpora-type data. We demonstrate the usefulness of our approach on a portion of the well-known Wall Street Journal corpus from the Penn Treebank and show that our approach is exponentially more efficient than the naive approach to storing n-grams and is also significantly more efficient than a traditional prefix tree.
This paper presents the work that we have carried out in inves tigating the purpose of discourse structure forwhy-question answering (why-QA). We developed a system for answer- ing why-questions that employs the discourse relations in a pre-an notated document collection (the RST Treebank). With this method, we obtain a recall of 53.3% with a mean reciprocal rank (MRR) of 0.662. We argue that the maximum recall that can be ob tained from the use of RST relations as proposed in the present paper is 58.0%. If we dis card the questions that require world knowledge, maximum recall is 73.9%. We conclude that d iscourse structure can play an important role in complex question answering, but that more forms of linguistic processing are needed for increasing recall.