Papers reviewed and determined not to be word norm studies. Use the flag icon to report errors or suggest re-inclusion.
16504 papers
Believing in Louis Althusser's axiom that writing is a campaign against absence, Shahrokh Meskoob struggled, throughout his literary and intellectual life, against all manners of absence. His language was a most potent weapon in, and the clearest manifestation of, his struggles. Many adjectives have been used by literary critics to describe his linguistic style. Some have called it a literary creation of first order. Others have characterized it as no less than a miracle; a phenomenon that defies definition. Indeed, Meskoob's language has an existence and identity of its own, palpably independent of the meanings and ideas that it projects. One can claim that his language a gate to a town; but a gate that deserves to be observed and admired in its own right. Meskoob was forever preoccupied with form. In a sense he constantly strove to flesh out and delineate abstract and diffused concepts. He wanted to embody and encase the flow of water, the fleeting time and the unbounded nature. In Meskoob's language, one does not come across the allegories common to classical and even modern Persian literature since he does not want to present his inner experiences through the filter of common linguistic norms or constructs and thereby pare away their original zest and liveliness. The sheer beauty, elegance and power of Meskoobs language is not only the natural byproduct of his mastery of Persian classic literature and western modern culture. It is also the concrete manifestation of his deep fascination with the language itself. In his own words: Language is the panacea of life and antidote to death. Ferdowsi used it to construct a lofty edifice never to be felled by the ravages of time.
The development of technologies for monitoring the welfare of crewmembers is a critical requirement for extended spaceflight. Behavior analytic methodologies provide a framework for studying the performance of individuals and groups, and brief computerized tests have been used successfully to examine the impairing effects of sleep, drug, and nutrition manipulations on human behavior. The purpose of the present study was to evaluate the feasibility and sensitivity of repeated performance testing during spaceflight. Four National Aeronautics and Space Administration crewmembers were trained to complete computerized questionnaires and performance tasks at repeated regular intervals before and after a 10-day shuttle mission and at times that interfered minimally with other mission activities during spaceflight. Two types of performance, Digit-Symbol Substitution trial completion rates and response times during the most complex Number Recognition trials, were altered slightly during spaceflight. All other dimensions of the performance tasks remained essentially unchanged over the course of the study. Verbal ratings of Fatigue increased slightly during spaceflight and decreased during the postflight test sessions. Arousal ratings increased during spaceflight and decreased postflight. No other consistent changes in rating-scale measures were observed over the course of the study. Crewmembers completed all mission requirements in an efficient manner with no indication of clinically significant behavioral impairment during the 10-day spaceflight. These results support the feasibility and utility of computerized task performances and questionnaire rating scales for repeated measurement of behavior during spaceflight.
We investigated the performance efficacy of beam search parsing and deep parsing techniques in probabilistic HPSG parsing using the Penn treebank. We first tested the beam thresholding and iterative parsing developed for PCFG parsing with an HPSG. Next, we tested three techniques originally developed for deep parsing: quick check, large constituent inhibition, and hybrid parsing with a CFG chunk parser. The contributions of the large constituent inhibition and global thresholding were not significant, while the quick check and chunk parser greatly contributed to total parsing performance. The precision, recall and average parsing time for the Penn treebank (Section 23) were 87.85%, 86.85%, and 360 ms, respectively.
Norms may also be understood as social realization of correctness notions and linguistic norms as performance instructions. This article reports on a case study of the application of Toury' s norm theory, particularly his operative translation norms, in subtitle translation. The authors believe that is a norm-governed communicative activity between two or more languages, and that subtitle translation, governed by linguistic and textual norms, should seek invisibility of subtitling as the ultimate goal.
We introduce a method for transferring annotation from a syntactically annotated corpus in a source language to a target language. Our approach assumes only that an (unannotated) text corpus exists for the target language, and does not require that the parameters of the mapping between the two languages are known. We outline a general probabilistic approach based on Data Augmentation, discuss the algorithmic challenges, and present a novel algorithm for sampling from a posterior distribution over trees.
This paper presents a lexicalized HMM-based approach to Chinese text chunking. To tackle the problem of unknown words, we formalize Chinese text chunking as a tagging task on a sequence of known words. To do this, we employ the uniformly lexicalized HMMs and develop a lattice-based tagger to assign each known word a proper hybrid tag, which involves four types of information: word boundary, POS, chunk boundary and chunk type. In comparison with most previous approaches, our approach is able to integrate different features such as part-of-speech information, chunk-internal cues and contextual information for text chunking under the framework of HMMs. As a result, the performance of the system can be improved without losing its efficiency in training and tagging. Our preliminary experiments on the PolyU Shallow Treebank show that the use of lexicalization technique can substantially improve the performance of a HMM-based chunking system.
Toempower thegeneral massthrough access toinformation andknowledge, organized efforts arebeing madetodevelop relevant content inlocallanguages andprovide local language capabilities toutility software. Wehavedeveloped a Question Answering (QA)System forHindidocuments that wouldberelevant formassesusingHindiasprimary language ofeducation. Theusershould beabletoaccess information fromE-learning documents ina userfriendly way,that isbyquestioning thesystem intheir native language Hindi andthesystem will return theintended answer (also in Hindi) bysearching incontext fromtherepository ofHindi documents. Thelanguage constructs, querystructure, commonwords, etc.arecompletely different inHindias compared toEnglish. A novelstrategy, inaddition to conventional search andNLP techniques, wasusedto construct theHindi QAsystem. Thefocus isoncontext based retrieval ofinformation. Forthis purpose weimplemented a Hindi search engine that works onlocality-based similarity heuristics toretrieve relevant passages fromthecollection. It alsoincorporates language analysis modules like stemmer andmorphological analyzer aswellasself constructed lexical database ofsynonyms. Theexperimental results over corpus oftwoimportant domains ofagriculture andscience showeffectiveness ofourapproach.
Traditional Chinese text chunking approach is to identify phrases using only one model and same features. It is shown that one model couldn't comprise each phrase's characteristics, and same features are not suitable to all phrases, data sparseness also appears. Multi-agent strategy uses several model and sensitive features of each phrase to identify different phrases. This paper describes the multi-agent strategy applied in the identification of Chinese phrases whose main features are: 1) easy and quick communication between phrases; 2) avoidance of data sparseness. Through testing on Chinese Penn Treebank, F score of Chinese text chunking using multi-agent strategy achieves to 95.82%, which is higher than the best result that has been reported.
The web has caused an explosion of documents, requiring the need for an automated text categorization system. This paper explores the notion of semantic feature selection by employing WordNet [Introduction to WordNet: An On-line Lexical Database], a lexical database. The proposed semantic approach employs noun synonyms and word senses for feature selection to select terms that are semantically representative of a category of documents. The categorical sense disambiguation extends the use of WordNet, which has been typically used for text retrieval and word sense disambiguation [A WordNet-based Algorithm for Word Sense Disambiguation]. Our experiments on the Reuters-21578 dataset have shown that automated semantic feature selection is able to perform better than well known statistical feature selection methods, Information Gain and Chi-Square as a feature selection method.
This dissertation is a study on the works of the Swedish author Carl Jonas Love Almqvist during the final years of his exile in America. Focusing on the monumental 1438-page unpublished manuscript 'About Swedish Rhymes', the study first presents the textual material and then discusses the text from different formal and content-based aspects essential to an understanding of Almqvist's works in exile. In the manuscripts preserved from his last years of exile, i.e. the period after 1860, Almqvist refers to 'Mr Hugo's Academy, established in the year 1838' introduced in one of the volumes of The Book of the Wild Rose (1839). In comparison with his earlier fiction about academic "cabinet meetings", this fiction of such an academy, conceived in exile, is in some ways extraordinary. A close reading of the texts reveals that the aging Almqvist, contrary to previous opinions about him, maintained strict control over the activity: the extension and division of the record, as well as its references to time and space, all indicate a complete consistency and an exact mimetic order. The consideration of 'About Swedish Rhymes' starts out from exterior qualities. The observations are first considered in relation to the author's statements on the importance of the manuscript for the literary work of art. Subsequently, the genesis of the "exile" texts is re-examined. One key question here is whether the manuscript was completed in Philadelphia, or was continued in Bremen during the final year of his life. The content of the conversations in the records of the cabinet meetings is also analyzed. Although questions of metre and versification dominate, the text also deals with a variety of widely differing subjects, including discussions about the use of language and linguistic norms. The fictitious frame that the cabinet meeting provides for the purpose of discussing metre and rhyme is also considered. Here we find various improvised verses composed at the cabinet meeting and put into the mouth of the authentic versifier H.J. Seseman. One important question is whether the cabinet-meeting discussions about the metre in these verses are intended to be a serious contribution to scholarly debate, or whether they in fact have ironic undertones. Next, the narration of the "exile" texts is discussed from the point of view provided by its own fictitious perspective, together with the author’s relation to irony, satire and parody. The concluding chapter deals with verse-making in the record of rhyming. The emphasis is laid on the analysis and characterization of the various rhymed verses collected under the title Sesemana. One essential question concerns the 'rubbishy' or 'plain' character of these poems. The present analysis indicates that questions of rubbish, textual triviality and the like must bow to the broader question of the character of the poems in a deeper sense. Seseman's poetry is considered in relation to the Songes collection. Finally the question of how rhythm manifests itself as 'free verse' in a number of these poems with more serious content is also discussed.
The structural theory of metaphor (STM) uses techniques from possible worlds semantics to generate and interpret metaphors. STM is presented in detail in The Logic of Metaphor: Analogous Parts of Possible Worlds (Steinhart, 2001). STM is based on Kittay’s semantic field theory of metaphor (1987) and ultimately on Black’s interactionist theory (1962, 1979). STM uses an intensional calculus to specify truth-conditions for many grammatical forms of metaphor. The truth-conditional analysis in STM is inspired in part by Miller (1979) and Hintikka & Sandu (1994). STM is by no means a toy theory. It has been successfully tested on dozens of large texts taken from real authors. Its methods can be applied in very large linguistic databases like WordNet or MindNet.
Synthesised stimuli were used to investigate how two notionally\nseparable dimensions of tone-of-voice? voice quality and\nfundamental frequency? are involved in the expression of\naffect. Listeners were presented with three series of stimuli:\n(1) stimuli exemplifying different voice qualities, (2) stimuli\nall with modal voice quality but with different affect-related f0\ncontours, and (3) stimuli incorporating variation in both voice\nquality and affect-related f0 contours. A total of 15 stimuli\nwere rated for 12 different affective attributes. Voice quality\ndifferentiation appears to account for the highest affect ratings\noverall, as indicated by the scores obtained for stimuli series\n(1) and (3). The relatively weaker affect signalling of stimuli\ndifferentiated by f0 alone corroborates findings in [2]. It also\nsuggests that for the generation of expressive, affectively\ncoloured speech synthesis, it is not sufficient to manipulate\nonly f0; we also need to capture the voice quality dimension\nof the voice source.
We formalize weighted dependency parsing as searching for maximum spanning trees (MSTs) in directed graphs. Using this representation, the parsing algorithm of Eisner (1996) is sufficient for searching over all projective trees in O(n3) time. More surprisingly, the representation is extended naturally to non-projective parsing using Chu-Liu-Edmonds (Chu and Liu, 1965; Edmonds, 1967) MST algorithm, yielding an O(n2) parsing algorithm. We evaluate these methods on the Prague Dependency Treebank using online large-margin learning techniques (Crammer et al., 2003; McDonald et al., 2005) and show that MST parsing increases efficiency and accuracy for languages with non-projective dependencies.
Recent empirical experiments on surface realizers have shown that grammars for generation can be effectively evaluated using large corpora. Evaluation metrics are usually reported as single averages across all possible types of errors and syntactic forms. But the causes of these errors are diverse, and the extent to which the accuracy of generation over individual syntactic phenomena is unknown. This article explores the types of errors, both computational and linguistic, inherent in the evaluation of a surface realizer when using large corpora. We analyze data from an earlier wide coverage experiment on the FUF/SURGE surface realizer with the Penn TreeBank in order to empirically classify the sources of errors and describe their frequency and distribution. This both provides a baseline for future evaluations and allows designers of NLG applications needing off-the-shelf surface realizers to choose on a quantitative basis. 1
In order to realize the full potential of dependency-based syntactic parsing, it is desirable to allow non-projective dependency structures. We show how a data-driven deterministic dependency parser, in itself restricted to projective structures, can be combined with graph transformation techniques to produce non-projective structures. Experiments using data from the Prague Dependency Treebank show that the combined system can handle non-projective constructions with a precision sufficient to yield a significant improvement in overall parsing accuracy. This leads to the best reported performance for robust non-projective parsing of Czech.
This paper reports the corpus-oriented development of a wide-coverage Japanese HPSG parser. We first created an HPSG treebank from the EDR corpus by using heuristic conversion rules, and then extracted lexical entries from the treebank. The grammar developed using this method attained wide coverage that could hardly be obtained by conventional manual development. We also trained a statistical parser for the grammar on the treebank, and evaluated the parser in terms of the accuracy of semantic-role identification and dependency analysis.
In this paper, we present a deterministic dependency structure analyzer for Chinese. This analyzer implements two algorithms – Yamada and Nivre models – and two sorts of classifiers – Support Vector Machines and Maximum Entropy methods. We compare the performance of these 2x2 combinations. We evaluate the method on a dependency tagged corpus derived from the CKIP Treebank corpus. Then, we analyzed the errors in the experiments and found that some errors were caused by mistakes of nominal compounds analysis. Therefore we adopt an NP-chunker to solve this problem.
It is widely believed that the difference between regular and irregular verbs is restricted to form. This study questions that belief. We report a series of lexical statistics showing that irregular verbs cluster in denser regions in semantic space. Compared to regular verbs, irregular verbs tend to have more semantic neighbors that in turn have relatively many other semantic neighbors that are morphologically irregular. We show that this greater semantic density for irregulars is reflected in association norms, familiarity ratings, visual lexical-decision latencies, and word-naming latencies. Meta-analyses of the materials of two neuroimaging studies show that in these studies, regularity is confounded with differences in semantic density. Our results challenge the hypothesis of the supposed formal encapsulation of rules of inflection and support lines of research in which sensitivity to probability is recognized as intrinsic to human language.
By Stéphane Chaudier. (Recherches proustiennes, 2). Paris, Champion, 2003. 549 pp. Hb €85.00. A reservoir of potent symbolic referents, religious vocabulary is deflected by Proust onto such secular contexts as society, love, sexuality, and literary creation. The tensions born of this juxtaposition of sacred and profane lie at the heart of Stéphane Chaudier's far-reaching analysis. Sensitive close textual readings illuminate the mechanisms of this ‘détournement’ of religious terms in the first part of the study. Playing with the hierarchy of sublime and trivial, high and low styles, Proust's often ludic handling of this intertext is interpreted as a means of transcending one's own universe, of perceiving its comic strangeness. Yet Chaudier rightly detects no anti-religious philosophical system in Proust's work, discovering instead, in his desacralization of religious terms, a stylistic transgression of social and linguistic norms and, more generally, an ironic perspective on all forms of cultural domination, whether organized religion or, as Chaudier persuasively discusses, a social orthodoxy. Examining the beliefs — akin to, but prevailing over religious convictions — that support the shifting edifice of social hierarchies, the study demonstrates how social power, in Proust, is based on ‘la gestion de biens symboliques’ (pp. 14–15) such as aristocratic prestige, wit, or the expression of avant-garde aesthetic opinions. Proust's presentation of this socio-religious hegemony is ambivalent, however, for, as Chaudier's analysis shows, Proust both recognizes, even admires, the strength of social groupings and simultaneously reveals social power to be hollow. Ontologically ‘empty’ signs such as Mme Verdurin's laugh are revealingly examined here. Reflecting the subtleties of Proust's position, Chaudier ultimately defines society as a ‘néant’, but a captivating one, and if idealization is followed by reality, ‘croyance’ by ‘savoir’, the former is not entirely negated by the latter — poetic truth remains. The study consistently foregrounds the material world, recognizing that the body is the conduit for experience. The role of religious metaphor in conceptualizing the erotic thus provides the focus of one part of the analysis. Presented in these terms, the erotic also becomes a metaphor for knowledge and creativity: Sodom, for example, offers a vision of fruitful union in terms of aesthetic — if not physical — creation, whilst Gomorrah, as the epitome of alterity, is the poetic principle that threatens language with extinction, that assigns its limits. Chaudier's subsequent examination of the function of Judaeo-Christian references within aesthetic reflections identifies the body as the artist's means of encountering the world rather than as an obstacle to the act of creation. Sensual and spiritual are thus reconciled: the artist-demiurge creates matter from nothing by an act of the spirit. The metaphor of the novel as cathedral is also teased out in all its intricacy — at times stable and coherent, at others vacillating and contradictory. Like the work of art, it is the incarnation of the infinite within the finite; ultimately, however, Proust's profane cathedral is a godless one. The study ends with a series of useful indexes and an extensive bibliography of French-language criticism. Opening up many new perspectives, it represents a valuable addition to recent English-language scholarship in this field.
Data-driven parsing techniques have a number of advantages over rule-based parsing techniques, such as fast development time, broad-coverage and robustness. Treebanks, collections of syntactically annotated sentences, are important resources for data-driven parsers. When developing a parser for Swedish one needs a treebank containing Swedish sentences, but currently there is a lack of Swedish treebanks of substantial size. This holds for the other Nordic languages too, with Danish as an exception. The absence of Swedish treebanks is remarkable considering that two corpora of Swedish text augmented with syntactic annotation have been created, one as early as 1974 named Talbanken (Einarsson 1976), and another in the 80's named Syntag (Järborg 1980). Unfortunately, the annotation formats of these resources make them cumbersome to use for modern treebank tools and parsers. In a way, Sweden can be regarded as a pioneer in this area, but thereafter the work with creating new treebanks has decreased considerably.
MONA is an automata toolkit providing a compiler for compiling formulae of monadic second order logic on strings or trees into string automata or tree automata. In this paper, we evaluate the option of using MONA as a treebank query tool. Unfortunately, we find that MONA is not an option. There are several reasons why the main being unsustainable query answer times. If the treebank contains larger trees with more than 100 nodes, then even the processing of simple queries may take hours.
Chunk parsing is conceptually appealing but its performance has not been satisfactory for practical use. In this paper we show that chunk parsing can perform significantly better than previously reported by using a simple sliding-window method and maximum entropy classifiers for phrase recognition in each level of chunking. Experimental results with the Penn Treebank corpus show that our chunk parser can give high-precision parsing outputs with very high speed (14 msec/sentence). We also present a parsing method for searching the best parse by considering the probabilities output by the maximum entropy classifiers, and show that the search method can further improve the parsing accuracy.
The goal of this article is to account for the resolution of vowel sequences across word boundaries in Catalan. Specifically, the paper accounts for the distribution of hiatuses and syllable contraction cases between two lexical words in this language. The article argues that V1 (the last vowel of the first word) does not undergo any change if it is followed by a vowel V2 bearing nuclear stress (or phrasal stress) prominence. The blocking of V1 glide formation will be seen in relation with the systematic maintenance of schwa in this position (canti ara [i »a] ‘you sing.imp now’, canto ara [u »a] ‘I sing now’, tallo ungles [u »u] ‘I cut nails’, canta ara [ »a] ‘he/she sings now’). Blocking of glide formation or schwa deletion is thus not due to rhythmic reasons (stress clash), as some previous studies have contended, but rather to the presence of a nuclear stress prominence on V2. This phenomenon will be interpreted as the instantiation of an alignment constraint which aligns the word-initial nuclear stressed foot to the left edge of the prosodic word. This alignment constraint triggers a ‘prosodic isolation’ phenomenon which prevents vowel gliding or deletion from applying. Finally, the paper also accounts for vowel sandhi in contexts where V2 is not stressed: in these contexts, syllable contraction is the norm. * Earlier versions of this work were presented at PaPI 2003 (Phonetics and Phonology in Iberia, Lisbon), at the Toulouse International Conference “From representations to constraints” (Toulouse, July 2003) and at the XVth International Congress of Phonetic Sciences (Barcelona, August 2003). We are grateful to those who attended these meetings for interesting observations and comments, and especially to Eulalia Bonet, Sonia Colina, Sonia Frota, Jose Ignacio Hualde, Michael Kenstowicz, John Kingston, Maria Rosa Lloret, Joan Mascaro, John McCarthy, Daniel Recasens, Elizabeth Selkirk, Donca Steriade, Hubert Truckenbrodt, Marina Vigario, and Max Wheeler for discussion of some parts of the material included in the article. Thanks are also due to Nuria Riera for transcribing the vowel contacts present in 5 spontaneous conversations of the Corpus Oral de Catala and to Marta Paya and Lluis Payrato for kindly providing us with a copy of this database before its publication. Finally, we thank Teresa Barenys, Julia Cufi, Teresa Espinal, Anna Gavarro, Nuria Marti, Jaume Sola, and Xavier Vall, who patiently responded to our questionnaire. All remaining errors are of course ours. This research was funded by grants 2002XT-00032 and 2001SGR 00150 from the Generalitat de Catalunya and BFF2003-06590 and BFF2003-09453-C02-C02 from the Ministry of Science and Technology of Spain.
During the past few years, the information representation and retrieval sector in the area of Documentation and Biblioteconomy has had to assume the important repercussions of the Internet and its associated technologies, and in particular, the World Wide Web (WWW). Technological modifications arising from these important changes are leading to the gradual digitalisation of the information representation and retrieval sector, affecting information artefacts, representation and retrieval tools and user requirements.\nIn the light of this growing context of digitalisation, diverse information representation and retrieval tools exist, which must be studied in addition to diverse fields of knowledge in which these tools have originated: Linguistics, Artificial Intelligence, Documentation, Linguistic Engineering... Hence, in specialised literature, analyses are performed on information representation and retrieval tools, taxonomies, classification systems, computational lexicons, lexical databases, thesauruses, titles lists, knowledge bases, conceptual maps, ontologies, synonym rings and semantic networks, among others. Among this wide spectrum of information representation and retrieval tools are thesauruses and ontologies, which are most often linked in bibliography, even though they come from completely different disciplinary areas. However, the conceptualisation applied by authors to the terms "thesaurus" and "ontology" is quite diverse, and sometimes authors confuse, oppose, complement or overlap both these concepts.\nThe overall objective of the present article [1] is to establish the relationship between the concepts of thesaurus and ontology in the Documentation and Biblioteconomy field. Two specific objectives have been established for this purpose. Firstly, to make an analysis of the thesaurus-based concept with a view to defining its most important characteristics and to verify the similarities and differences it shares with ontologies. And secondly, to establish a definition for the ontology concept, also for the purpose of verifying its characteristics and analysing the similarities and differences it has with thesauruses.
Translation tests are widely used for high school term tests and entrance examinations as well as university entrance examinations. Although a considerable number of papers point out the possibility of low reliability for scoring, little is known about the factors raters play in the reliability of scoring (Watanabe, 1994). This study examines how the professional backgrounds of raters affect rating criteria. The results indicated that novice raters tended to over-estimate examinees' comprehension whereas experienced raters were more likely to focus on the correctness of the Japanese sentence. In addition, it turned out that the difficulty of sentences affected the scoring of both experienced and novice raters. After administering a sorting task, the difficulty of the sentences showed that the perception of sentence difficulty did not correspond to the difficulty of examinees' translation. The paper closes by suggesting several pedagogical implications for administering translation tests. Of particular importance is that test developers should consider not only the complexity of sentence structures and vocabulary, but also the examinees' topic familiarity of the sentences to be translated.
A core function of the olfactory system is to determine the valence of odors. In humans, central processing of odor valence perception has been shown to take form already within the olfactory bulb (OB), but the neural mechanisms by which this important information is communicated to, and from, the olfactory cortex (piriform cortex, PC) are not known. To assess communication between the 2 nodes, we simultaneously measured odor-dependent neural activity in the OB and PC from human participants while obtaining trial-by-trial valence ratings. By doing so, we could determine when subjective valence information was communicated, what kind of information was transferred, and how the information was transferred (i.e., in which frequency band). Support vector machine (SVM) learning was used on the coherence spectrum and frequency-resolved Granger causality to identify valence-dependent differences in functional and effective connectivity between the OB and PC. We found that the OB communicates subjective odor valence to the PC in the gamma band shortly after odor onset, while the PC subsequently feeds broader valence-related information back to the OB in the beta band. Decoding accuracy was better for negative than positive valence, suggesting a focus on negative valence. Critically, we replicated these findings in an independent data set using additional odors across a larger perceived valence range. Combined, these results demonstrate that the OB and PC communicate levels of subjective odor pleasantness across multiple frequencies, at specific time points, in a direction-dependent pattern in accordance with a two-stage model of odor processing.
This paper presents a high performance method to identify English proper nouns (PNs) based on maximum entropy model (MaxEnt). Most traditional PNs recognition systems use lexical resources such as name list, as new names are constantly coming into existence, these are necessarily incomplete. Therefore machine learning methods are used to identify PNs automatically. In the framework of MaxEnt model, semantic and lexical information of surrounding words and word itself acting as atomic features comprises feature templates and forms feature without requiring extra expert knowledge. The test on WSJ of Penn Treebank II shows that this method guarantees high precision and recall, and at the same time it can reduce the quantity of features dramatically, downsize system space consumption, and decrease the time of training and testing, so as to improve the efficiency considerably. The method in this paper can be transformed to identify other specific noun easily because the principle of methods is universal.
In the year 2001, the French government made the Creole languages of Guadeloupe, Guyane, Martinique and Reunion, one single “Regional Language of France”. The main reason for this policy is educational and should lead to results in creating one single teacher’s assessment exam for that subject. Disregarding local differences, underestimating the complexity of the settling of regional linguistic norms, and blindly following the all Creole activist discourse, the French authorities launched a project that has had no significant positive results. This article leads to the conclusion that there is a need for a differentiating branch of linguistics and recommends the creation of normative commissions working in the field with a goal to implement efficient policies.
Most existing statistical surface realizers either make use of hand-crafted grammars to provide coverage or are tuned to specific applications. This paper describes an initial effort toward building a statistical surface realization model that provides both precision and coverage. We trained a Maximum Entropy model that given a predicate-argument semantic representation, predicts the surface form for realizing a semantic concept and the ordering of sibling semantic concepts and their parent, on the Penn TreeBank and Proposition Bank corpora. Initial results have shown that the precisions for predicting surface forms and orderings reached 80% and 90% respectively, on a held-out part of Penn TreeBank. We use the model to generate sentences from our domain representations. We are in the process of evaluating the model on a corpus collected for our in-car applications.
This master’s thesis describes a deterministic dependency parser using a memorybased learning approach to parse unrestricted English text. A converter transforms the Wall Street Journal section of the Penn Treebank to an intermediate dependency representation which is used to train the parser using the TiMBL (Daelemans, Zavrel, Sloot, & Bosch, 2003) library. The output of the parser is labeled dependency graphs, using as arc labels a combination of bracket labels and grammatical role labels constructed from the Penn Treebank II annotation scheme (Marcus, Kim, et al., 1994). The parser reaches a maximum unlabeled attachment score of 87.1% and produces labeled dependency graphs with an accuracy of of 86.0% with the correct head and arc label recognised. The results are close to the state of the art in dependency parsing, and the parser also outputs arc labels that other parsers do not produce.
Knowledge acquisition is always regarded as a bottleneck in many NLP tasks, such as machine translation, information extraction. Treebank-based statistical parsing is not an exceptant. The latent linguistic knowledge in treebank is very rich, which, however, cant be acquired directly.In our model, the following three ways are used to incorporate such rich linguistic features for Chinese statistical parsing. First of all, non-recursive noun and verb phrases are annotated in the Penn Chinese Treebank because of their strong mark of boundaries. Second, a new head percolation table is designed based on Xias table. The last linguistic feature our model uses is the context configuration frame which provides a stronger representation of bilexical dependency structures. All these three linguistic features gain an improvement of remarkable 2.37% in terms of F1 measure, 5.36% in terms of complete match ratio.
In this paper, we propose a feature-based Korean grammar utilizing the learned constraint rules in order to improve parsing efficiency. The proposed grammar consists of feature structures, feature operations, and constraint rules; and it has the following characteristics. First, a feature structure includes several features to express useful linguistic information for Korean parsing. Second, a feature operation generating a new feature structure is restricted to the binary-branching form which can deal with Korean properties such as variable word order and constituent ellipsis. Third, constraint rules improve efficiency by preventing feature operations from generating spurious feature structures. Moreover, these rules are learned from a Korean treebank by a decision tree learning algorithm. The experimental results show that the feature-based Korean grammar can reduce the number of candidates by a third of candidates at most and it runs 1.5 ∼ 2 times faster than a CFG on a statistical parser.
An automatic method for annotating the Penn-II Treebank (Marcus et al., 1994) with high-level Lexical Functional Grammar (Kaplan and Bresnan, 1982; Bresnan, 2001; Dalrymple, 2001) f-structure representations is presented by Burke et al. (2004b). The annotation algorithm is the basis for the automatic acquisition of wide-coverage and robust probabilistic approximations of LFG grammars (Cahill et al., 2004) and for the induction of subcategorisation frames (O’Donovan et al., 2004; O’Donovan et al., 2005). Annotation quality is, therefore, extremely important and to date has been measured against the DCU 105 and the PARC 700 Dependency Bank (King et al., 2003). The annotation algorithm achieves f-scores of 96.73% for complete f-structures and 94.28% for preds-only f-structures against the DCU 105 and 87.07% against the PARC 700 using the feature set of Kaplan et al. (2004). Burke et al. (2004a) provides detailed analysis of these results. \nThis paper presents an evaluation of the annotation algorithm against PropBank (Kingsbury and Palmer, \n2002). PropBank identifies the semantic arguments of each predicate in the Penn-II treebank and annotates their semantic roles. As PropBank was developed independently of any grammar formalism it provides a platform for making more meaningful comparisons between parsing technologies than was previously possible. PropBank also allows a much larger scale evaluation than the smaller DCU 105 and PARC 700 gold standards. In order to perform the evaluation, first, we automatically converted the PropBank annotations \ninto a dependency format. Second, we developed conversion software to produce PropBank-style semantic annotations in dependency format from the f-structures automatically acquired by the annotation algorithm from Penn-II. The evaluation was performed using the evaluation software of Crouch et al. (2002) and Riezler et al. (2002). Using the Penn-II Wall Street Journal Section 24 as the development set, currently we achieve an f-score of 76.58% against PropBank for the Section 23 test set.
We describe a history-based generative parsing model which uses a k-nearest neighbour (k-NN) technique to estimate the model's parameters. Taking the output of a base n-best parser we use our model to re-estimate the log probability of each parse tree in the n-best list for sentences from the Penn Wall Street Journal treebank. By further decomposing the local probability distributions of the base model, enriching the set of conditioning features used to estimate the model's parameters, and using k-NN as opposed to the Witten-Bell estimation of the base model, we achieve an f-score of 89.2%, representing a 4% relative decrease in f-score error over the 1-best output of the base parser.
This paper presents a Chinese parsing method which takes data-oriented parsing technique as the basic framework and utilizes the similarity-based probability estimate technique. Through the initial selection process, the fragment-combination forms of the input sentence are acquired on the constructed knowledge source including treebank, fragment-bank and fragment-combination-bank. Then by using the similarity-based probability estimate technique, the combination parsing process can be completed successfully. To prove the method efficiency, the knowledge source is constructed on the real-world Chinese corpus, and the other corpus is used as the test set. The experiment results show that every test parameter is satisfied.