Papers reviewed and determined not to be word norm studies. Use the flag icon to report errors or suggest re-inclusion.
16504 papers
With the growing access to heterogeneous and independent data repositories, determining the semantic difference of two ontologies is critical in information retrieval, information integration and semantic web applications. In this paper, we propose an ontology comparison tool based on a novel senses refinement algorithm, which builds a senses set to accurately represent the semantics of the input ontology. The senses refinement algorithm automatically extracts senses from the electronic lexical database WordNet (locally installed or online), removes unnecessary senses based on the relationship among the entity classes of the ontology, and specifies relations and constraints of the concepts in the refined senses set. The senses refinement converts the measurement of ontology difference into simple set operations based on set theory, thus ensures the efficiency and accuracy of the ontology comparison. Our experimental studies show that the proposed senses refinement algorithm outperforms the naive senses set construction algorithm in terms of efficiency and accuracy.
Purpose: Strategy selection and effort allocation in cognitive tasks are often mediated by metacognitive assessments of self-performance (e.g., Nelson & Leonesio, 1988). The current experiment examined the accuracy and role of predictive metacognitive judgments in a real-world visual search task. Method: Procedure was modeled after the judgments-of-learning task (e.g., Dunlosky & Nelson, 1992) commonly used to study metacognitive accuracy and control. In a pair of experiments, subjects performed a simulated luggage x-ray screening task, searching for knives hidden among varying numbers of background objects in passenger bags. Before performing the search task, subjects viewed each stimulus image without the embedded target item and rated how likely they would be to find the knife if it were hidden somewhere in that image. Ratings were made on a 5-point Likert scale. Experiment 1 used a 2IFC procedure for the visual search task. Experiment 2 used a speeded response procedure. Results: Experiment 1 produced a statistically significant but weak correlation between predicted and observed target detection performance (mean Goodman-Kruskal gamma =.161), indicating that subjects' predictive metacognitive judgments were only modestly accurate. Experiment 2 found that target-present RTs were similar across levels of predicted detection likelihood, but that target-absent RTs were longer for stimulus images in which target detection was predicted to be difficult. Conclusions: Results suggest that predictive metacognitive judgments are only modestly accurate, but that the information used in metacognitive judgments is nonetheless used to regulate criterion for terminating search when no target is detected.
Linguistic politeness is intimately connected with social norms. Estoniansociety has gone through considerable change over the last ten years. It hasregained independence and, at the same time, switched from a planned toa market economy as well as from dictatorship to democracy. A decade ismost probably not long enough for linguistic norms to change drastically:as we know, the structure of a language often takes much longer to change.Politeness, however, may to some extent be subject to deliberate influence,as witnessed, for example, by the reform of Swedish du (you, sg.) where therecommendations of some left-wing organisations on the usage of mutualdu (T) have won general social acceptance. It is, thus, not unlikely that change is taking place in Estonian politeness at present....
We present a method for improving dependency structure analysis of Chinese. Our bottom-up deterministic analyzer adopt Nivre’s algorithm (Nivre and Scholz, 2004). Support Vector Machines (SVMs) are utilized to determine the word dependency relations. We find that there are two problems in our analyzer and propose two methods to solve them. One problem is that some operations cannot be solved only using local feature. We utilize the global features to solve this. The other problem is that this bottom-up analyzer doesn’t use top-down information. We supply the top-down information by constructing SVMs based root node finder to solve this problem. Experimental evaluation on the Penn Chinese Treebank Corpus shows that the proposed extensions improve the parsing accuracy significantly. 1
During the past few years, the information representation and retrieval sector in the area of Documentation and Biblioteconomy has had to assume the important repercussions of the Internet and its associated technologies, and in particular, the World Wide Web (WWW). Technological modifications arising from these important changes are leading to the gradual digitalisation of the information representation and retrieval sector, affecting information artefacts, representation and retrieval tools and user requirements.\nIn the light of this growing context of digitalisation, diverse information representation and retrieval tools exist, which must be studied in addition to diverse fields of knowledge in which these tools have originated: Linguistics, Artificial Intelligence, Documentation, Linguistic Engineering... Hence, in specialised literature, analyses are performed on information representation and retrieval tools, taxonomies, classification systems, computational lexicons, lexical databases, thesauruses, titles lists, knowledge bases, conceptual maps, ontologies, synonym rings and semantic networks, among others. Among this wide spectrum of information representation and retrieval tools are thesauruses and ontologies, which are most often linked in bibliography, even though they come from completely different disciplinary areas. However, the conceptualisation applied by authors to the terms "thesaurus" and "ontology" is quite diverse, and sometimes authors confuse, oppose, complement or overlap both these concepts.\nThe overall objective of the present article [1] is to establish the relationship between the concepts of thesaurus and ontology in the Documentation and Biblioteconomy field. Two specific objectives have been established for this purpose. Firstly, to make an analysis of the thesaurus-based concept with a view to defining its most important characteristics and to verify the similarities and differences it shares with ontologies. And secondly, to establish a definition for the ontology concept, also for the purpose of verifying its characteristics and analysing the similarities and differences it has with thesauruses.
Agent technologies represent a promising approach for the integration of interorganizational capabilities across distributed, networked environments. However, knowledge sharing interoperability problems can arise when agents incorporating differing ontologies try to synchronize their internal information. Moreover, in practice, agents may not have a common or global consensus ontology that will facilitate knowledge sharing and integration of functional capabilities. We propose a method to enable agents to develop a local consensus ontology during operation time as needed. By identifying similarities in the ontologies of their peer agents, a set of agents can discover new concepts/relations and integrate them into a local consensus ontology on demand. We evaluate this method, both syntactically and semantically, when forming local consensus ontologies with and without the use of a lexical database. We also report on the effects when several factors, such as the similarity measure, the relation search level depth, and the merge order, are varied. Finally, experimenting in the domain of agent-supported Web service composition, we demonstrate how our method allows us to successfully autonomously form service-oriented local consensus ontologies.
The paper investigates the use of richer syntactic dependencies in the structured language model (SLM). We present two simple methods of enriching the dependencies in the syntactic parse trees used for initializing the SLM. We evaluate the impact of both methods on the perplexity (PPL) and word-error-rate (WER, N-best rescoring) performance of the SLM. We show that the new model achieves an improvement in PPL and WER over the baseline results reported using the SLM on the UPenn Treebank and Wall Street Journal (WSJ) corpora, respectively.
Translation tests are widely used for high school term tests and entrance examinations as well as university entrance examinations. Although a considerable number of papers point out the possibility of low reliability for scoring, little is known about the factors raters play in the reliability of scoring (Watanabe, 1994). This study examines how the professional backgrounds of raters affect rating criteria. The results indicated that novice raters tended to over-estimate examinees' comprehension whereas experienced raters were more likely to focus on the correctness of the Japanese sentence. In addition, it turned out that the difficulty of sentences affected the scoring of both experienced and novice raters. After administering a sorting task, the difficulty of the sentences showed that the perception of sentence difficulty did not correspond to the difficulty of examinees' translation. The paper closes by suggesting several pedagogical implications for administering translation tests. Of particular importance is that test developers should consider not only the complexity of sentence structures and vocabulary, but also the examinees' topic familiarity of the sentences to be translated.
Believing in Louis Althusser's axiom that writing is a campaign against absence, Shahrokh Meskoob struggled, throughout his literary and intellectual life, against all manners of absence. His language was a most potent weapon in, and the clearest manifestation of, his struggles. Many adjectives have been used by literary critics to describe his linguistic style. Some have called it a literary creation of first order. Others have characterized it as no less than a miracle; a phenomenon that defies definition. Indeed, Meskoob's language has an existence and identity of its own, palpably independent of the meanings and ideas that it projects. One can claim that his language a gate to a town; but a gate that deserves to be observed and admired in its own right. Meskoob was forever preoccupied with form. In a sense he constantly strove to flesh out and delineate abstract and diffused concepts. He wanted to embody and encase the flow of water, the fleeting time and the unbounded nature. In Meskoob's language, one does not come across the allegories common to classical and even modern Persian literature since he does not want to present his inner experiences through the filter of common linguistic norms or constructs and thereby pare away their original zest and liveliness. The sheer beauty, elegance and power of Meskoobs language is not only the natural byproduct of his mastery of Persian classic literature and western modern culture. It is also the concrete manifestation of his deep fascination with the language itself. In his own words: Language is the panacea of life and antidote to death. Ferdowsi used it to construct a lofty edifice never to be felled by the ravages of time.
A core function of the olfactory system is to determine the valence of odors. In humans, central processing of odor valence perception has been shown to take form already within the olfactory bulb (OB), but the neural mechanisms by which this important information is communicated to, and from, the olfactory cortex (piriform cortex, PC) are not known. To assess communication between the 2 nodes, we simultaneously measured odor-dependent neural activity in the OB and PC from human participants while obtaining trial-by-trial valence ratings. By doing so, we could determine when subjective valence information was communicated, what kind of information was transferred, and how the information was transferred (i.e., in which frequency band). Support vector machine (SVM) learning was used on the coherence spectrum and frequency-resolved Granger causality to identify valence-dependent differences in functional and effective connectivity between the OB and PC. We found that the OB communicates subjective odor valence to the PC in the gamma band shortly after odor onset, while the PC subsequently feeds broader valence-related information back to the OB in the beta band. Decoding accuracy was better for negative than positive valence, suggesting a focus on negative valence. Critically, we replicated these findings in an independent data set using additional odors across a larger perceived valence range. Combined, these results demonstrate that the OB and PC communicate levels of subjective odor pleasantness across multiple frequencies, at specific time points, in a direction-dependent pattern in accordance with a two-stage model of odor processing.
It is not clear a priori how well parsers trained on the Penn Treebank will parse significantly different corpora without retraining. We carried out a competitive evaluation of three leading treebank parsers on an annotated corpus from the human molecular biology domain, and on an extract from the Penn Treebank for comparison, performing a detailed analysis of the kinds of errors each parser made, along with a quantitative comparison of syntax usage between the two corpora. Our results suggest that these tools are becoming somewhat over-specialised on their training domain at the expense of portability, but also indicate that some of the errors encountered are of doubtful importance for information extraction tasks.
Synthesised stimuli were used to investigate how two notionally\nseparable dimensions of tone-of-voice? voice quality and\nfundamental frequency? are involved in the expression of\naffect. Listeners were presented with three series of stimuli:\n(1) stimuli exemplifying different voice qualities, (2) stimuli\nall with modal voice quality but with different affect-related f0\ncontours, and (3) stimuli incorporating variation in both voice\nquality and affect-related f0 contours. A total of 15 stimuli\nwere rated for 12 different affective attributes. Voice quality\ndifferentiation appears to account for the highest affect ratings\noverall, as indicated by the scores obtained for stimuli series\n(1) and (3). The relatively weaker affect signalling of stimuli\ndifferentiated by f0 alone corroborates findings in [2]. It also\nsuggests that for the generation of expressive, affectively\ncoloured speech synthesis, it is not sufficient to manipulate\nonly f0; we also need to capture the voice quality dimension\nof the voice source.
Toempower thegeneral massthrough access toinformation andknowledge, organized efforts arebeing madetodevelop relevant content inlocallanguages andprovide local language capabilities toutility software. Wehavedeveloped a Question Answering (QA)System forHindidocuments that wouldberelevant formassesusingHindiasprimary language ofeducation. Theusershould beabletoaccess information fromE-learning documents ina userfriendly way,that isbyquestioning thesystem intheir native language Hindi andthesystem will return theintended answer (also in Hindi) bysearching incontext fromtherepository ofHindi documents. Thelanguage constructs, querystructure, commonwords, etc.arecompletely different inHindias compared toEnglish. A novelstrategy, inaddition to conventional search andNLP techniques, wasusedto construct theHindi QAsystem. Thefocus isoncontext based retrieval ofinformation. Forthis purpose weimplemented a Hindi search engine that works onlocality-based similarity heuristics toretrieve relevant passages fromthecollection. It alsoincorporates language analysis modules like stemmer andmorphological analyzer aswellasself constructed lexical database ofsynonyms. Theexperimental results over corpus oftwoimportant domains ofagriculture andscience showeffectiveness ofourapproach.
This paper reports the corpus-oriented development of a wide-coverage Japanese HPSG parser. We first created an HPSG treebank from the EDR corpus by using heuristic conversion rules, and then extracted lexical entries from the treebank. The grammar developed using this method attained wide coverage that could hardly be obtained by conventional manual development. We also trained a statistical parser for the grammar on the treebank, and evaluated the parser in terms of the accuracy of semantic-role identification and dependency analysis.
The structural theory of metaphor (STM) uses techniques from possible worlds semantics to generate and interpret metaphors. STM is presented in detail in The Logic of Metaphor: Analogous Parts of Possible Worlds (Steinhart, 2001). STM is based on Kittay’s semantic field theory of metaphor (1987) and ultimately on Black’s interactionist theory (1962, 1979). STM uses an intensional calculus to specify truth-conditions for many grammatical forms of metaphor. The truth-conditional analysis in STM is inspired in part by Miller (1979) and Hintikka & Sandu (1994). STM is by no means a toy theory. It has been successfully tested on dozens of large texts taken from real authors. Its methods can be applied in very large linguistic databases like WordNet or MindNet.
Recently, shallow parsing has been applied to various information processing systems, such as information retrieval, information extraction, question answering, and automatic document summarization. A shallow parser is suitable for online applications, because it is much more efficient and less demanding than a full parser. In this research, we formulate shallow parsing as a sequential tagging problem and use a supervised machine learning technique, Maximum Entropy (ME), to build a Chinese shallow parser. The major features of the ME-based shallow parser are POSs and the context words in a sentence. We adopt the shallow parsing results of Sinica Treebank as our standard, and select 30,000 and 10,000 sentences from Sinica Treebank as the training set and test set respectively. We then test the robustness of the shallow parser with noisy data. The experiment results show that the proposed shallow parser is quite robust for sentences with unknown proper nouns. 1.
One of the main motivations for building treebanks is that they facilitate the development of syntactic parsers, by providing realistic data for evaluation as well as inductive learning. In this paper we present what we believe to be the first robust data-driven parser for Bulgarian, trained and evaluated on data from BulTreeBank (Simov et al., 2002). The parser uses dependency-based representations and employs a deterministic algorithm to construct dependency structures in a single pass over the input string, guided by a memory-based classifier at each nondeterministic choice point, as described in Nivre et al. (2004). Since the original BulTreeBank annotation is based on HPSG, it has been necessary to extract dependency structures from the original annotation both for training and evaluating the parser. The paper is structured as follows. Section 2 introduces the MaltParser system used in the experiments to induce parsers from treebank data. Section 3 presents BulTreeBank, and section 4 shows how the BulTreeBank annotation can be converted into a dependency structure annotation of the kind required by MaltParser. Section 5 describes the experimental conditions, and section 6 discusses the results of the experiments. Section 7 contains our conclusions.
This paper presents a high performance method to identify English proper nouns (PNs) based on maximum entropy model (MaxEnt). Most traditional PNs recognition systems use lexical resources such as name list, as new names are constantly coming into existence, these are necessarily incomplete. Therefore machine learning methods are used to identify PNs automatically. In the framework of MaxEnt model, semantic and lexical information of surrounding words and word itself acting as atomic features comprises feature templates and forms feature without requiring extra expert knowledge. The test on WSJ of Penn Treebank II shows that this method guarantees high precision and recall, and at the same time it can reduce the quantity of features dramatically, downsize system space consumption, and decrease the time of training and testing, so as to improve the efficiency considerably. The method in this paper can be transformed to identify other specific noun easily because the principle of methods is universal.
Current trends in language technology require treebanks that do not stop at the level of constituent structure, but include deeper and richer levels of analysis, including appropriate meaning structures. Capturing sufficient detail at different levels of linguistic description is too complex a task to be practically achievable by manual annotation or shallow parsing; rather it requires sophisticated tools that help secure the consistency of parallel but different structures. We are constructing a multilevel treebanking tool that incorporates a deep parser and grammar for Norwegian. Thus, we are tightly linking our treebank to grammar development so as to achieve a sound embedding in grammatical theory and yield more useful results for applications.
Building on work showing the harmfulness of annotation errors for both the training and evaluation of natural language processing technologies, this thesis develops a method for detecting and correcting errors in corpora with linguistic annotation. The so-called variation n-gram method relies on the recurrence of identical strings with varying annotation to find erroneous mark-up. We show that the method is applicable for varying complexities of annotation. The method is most readily applied to positional annotation, such as part-of-speech annotation, but can be extended to structural annotation, both for tree structures— as with syntactic annotation—and for graph structures—as with syntactic annotation allowing discontinuous constituents, or crossing branches. Furthermore, we demonstrate that the notion of variation for detecting errors is a powerful one, by searching for grammar rules in a treebank which have the same daughters but different mothers. We also show that such errors impact the effectiveness of a grammar induction algorithm and subsequent parsing. After detecting errors in the different corpora, we turn to correcting such errors, through the use of more general classification techniques. Our results indicate that the particular classification algorithm is less important than understanding the nature of the errors and altering the classifiers to deal with these errors. With such alterations, we can automatically correct errors with 85% accuracy. By sorting the errors, we can
To empower the general mass through access to information and knowledge, organized efforts are being made to develop relevant content in local languages and provide local language capabilities to utility software. We have developed a Question Answering (QA) System for Hindi documents that would be relevant for masses using Hindi as primary language of education. The user should be able to access information from E-learning documents in a user friendly way, that is by questioning the system in their native language Hindi and the system will return the intended answer (also in Hindi) by searching in context from the repository of Hindi documents. The language constructs, query structure, common words, etc. are completely different in Hindi as compared to English. A novel strategy, in addition to conventional search and NLP techniques, was used to construct the Hindi QA system. The focus is on context based retrieval of information. For this purpose we implemented a Hindi search engine that works on locality-based similarity heuristics to retrieve relevant passages from the collection. It also incorporates language analysis modules like stemmer and morphological analyzer as well as self constructed lexical database of synonyms. The experimental results over corpus of two important domains of agriculture and science show effectiveness of our approach.
This paper presents an integrated interactive user interface for teaching grammatical analysis on the Internet (Visual Interactive Syntax Learning), developed at the University of Southern Denmark, offering a unified system of analysis for 22 different languages, 7 of which are supported by live grammatical analysis of running text. For reasons of robustness, efficiency and correctness, the system's internal tools are based on the Constraint Grammar Formalism (Karlsson 1990), but users are free to choose from a variety of notional filters, supporting different descriptional paradigms, with a current teaching focus on syntactic tree structures, language independent grammatical categories and the form-function dichotomy. VISL's core NLP-programs use hybrid multi-level parsers (Bick 2003), while graphical teaching applications and corpus searching tools are implemented as platform independent Java-programs and Perlegi's. Though lexica and parsing rules are developed individually for each language, a common CG and treebank data format facilitates source data transfer into grammar teaching games, structural or color based visualisation, and linguistic revision of corpus data
Vese of the third generation,with its vehement passion,strives to break through the confines of entire discourse and to transcend the summit of poetic creation newly instituted by hazy poetry.In a word,it aims to grapple with language by disdaining and revolting against all fixed linguistic norms so as to shake off the grave sense of social and historical responsibility shouldered by poets of former generations,the ideological restraints imposed by tragic social reality,and above all to point to the very essence of life and inquire into the minute sense of existence as acutely felt by individuals.The crux lies in language,however and what verse of the third generation has endeavored to attain remains utopian in linguistic aspects in that the goal thus targeted would,in consequence,make verse readily understood by few readers and merely by the poet him/herself,hence the somewhat blocked channel of communication between poets and readers.Such a phenomenon has sounded the alarm for literature as follows: adequate attention must be paid to language in literary creation of any kind.
In the year 2001, the French government made the Creole languages of Guadeloupe, Guyane, Martinique and Reunion, one single “Regional Language of France”. The main reason for this policy is educational and should lead to results in creating one single teacher’s assessment exam for that subject. Disregarding local differences, underestimating the complexity of the settling of regional linguistic norms, and blindly following the all Creole activist discourse, the French authorities launched a project that has had no significant positive results. This article leads to the conclusion that there is a need for a differentiating branch of linguistics and recommends the creation of normative commissions working in the field with a goal to implement efficient policies.
By Stéphane Chaudier. (Recherches proustiennes, 2). Paris, Champion, 2003. 549 pp. Hb €85.00. A reservoir of potent symbolic referents, religious vocabulary is deflected by Proust onto such secular contexts as society, love, sexuality, and literary creation. The tensions born of this juxtaposition of sacred and profane lie at the heart of Stéphane Chaudier's far-reaching analysis. Sensitive close textual readings illuminate the mechanisms of this ‘détournement’ of religious terms in the first part of the study. Playing with the hierarchy of sublime and trivial, high and low styles, Proust's often ludic handling of this intertext is interpreted as a means of transcending one's own universe, of perceiving its comic strangeness. Yet Chaudier rightly detects no anti-religious philosophical system in Proust's work, discovering instead, in his desacralization of religious terms, a stylistic transgression of social and linguistic norms and, more generally, an ironic perspective on all forms of cultural domination, whether organized religion or, as Chaudier persuasively discusses, a social orthodoxy. Examining the beliefs — akin to, but prevailing over religious convictions — that support the shifting edifice of social hierarchies, the study demonstrates how social power, in Proust, is based on ‘la gestion de biens symboliques’ (pp. 14–15) such as aristocratic prestige, wit, or the expression of avant-garde aesthetic opinions. Proust's presentation of this socio-religious hegemony is ambivalent, however, for, as Chaudier's analysis shows, Proust both recognizes, even admires, the strength of social groupings and simultaneously reveals social power to be hollow. Ontologically ‘empty’ signs such as Mme Verdurin's laugh are revealingly examined here. Reflecting the subtleties of Proust's position, Chaudier ultimately defines society as a ‘néant’, but a captivating one, and if idealization is followed by reality, ‘croyance’ by ‘savoir’, the former is not entirely negated by the latter — poetic truth remains. The study consistently foregrounds the material world, recognizing that the body is the conduit for experience. The role of religious metaphor in conceptualizing the erotic thus provides the focus of one part of the analysis. Presented in these terms, the erotic also becomes a metaphor for knowledge and creativity: Sodom, for example, offers a vision of fruitful union in terms of aesthetic — if not physical — creation, whilst Gomorrah, as the epitome of alterity, is the poetic principle that threatens language with extinction, that assigns its limits. Chaudier's subsequent examination of the function of Judaeo-Christian references within aesthetic reflections identifies the body as the artist's means of encountering the world rather than as an obstacle to the act of creation. Sensual and spiritual are thus reconciled: the artist-demiurge creates matter from nothing by an act of the spirit. The metaphor of the novel as cathedral is also teased out in all its intricacy — at times stable and coherent, at others vacillating and contradictory. Like the work of art, it is the incarnation of the infinite within the finite; ultimately, however, Proust's profane cathedral is a godless one. The study ends with a series of useful indexes and an extensive bibliography of French-language criticism. Opening up many new perspectives, it represents a valuable addition to recent English-language scholarship in this field.
Introduction: This exploratory study investigated the impact of altered physiological states on emotion processing in quadriplegic patients, endurance athletes and healthy controls. High cervical spinal cord injury patients differ from both other groups in their lack of peripheral feedback and their inability to show motor approach or withdrawal behavior. However, both patients and athletes at rest are distinguished from non-athletic controls by their reduced sympathetic nervous system activity which may lead to less intense emotional experiences. Methods: We measured both physiological (event-related potentials, skin conductance, heart rate) and subjective responses (valence and arousal ratings) to stimuli from the International Affective Picture System. Results: While there were no group differences in arousal, both athletes and quadriplegics tended to rate positive slides as less pleasant and negative slides as less unpleasant than ablebodied controls. Skin conductance response was reduced in athletes and absent in quadriplegics, and both groups had greater heart rate deceleration than controls. The three groups showed similar electrocortical responses to both pleasant and unpleasant pictures. Discussion: The similarities between patients and athletes suggest that feedback of peripheral sympathicoadrenergic processes contributes to emotion processing. Further research is warranted to elucidate the contributions of autonomic nervous system activity to emotional experience.
Most existing statistical surface realizers either make use of hand-crafted grammars to provide coverage or are tuned to specific applications. This paper describes an initial effort toward building a statistical surface realization model that provides both precision and coverage. We trained a Maximum Entropy model that given a predicate-argument semantic representation, predicts the surface form for realizing a semantic concept and the ordering of sibling semantic concepts and their parent, on the Penn TreeBank and Proposition Bank corpora. Initial results have shown that the precisions for predicting surface forms and orderings reached 80% and 90% respectively, on a held-out part of Penn TreeBank. We use the model to generate sentences from our domain representations. We are in the process of evaluating the model on a corpus collected for our in-car applications.
This master’s thesis describes a deterministic dependency parser using a memorybased learning approach to parse unrestricted English text. A converter transforms the Wall Street Journal section of the Penn Treebank to an intermediate dependency representation which is used to train the parser using the TiMBL (Daelemans, Zavrel, Sloot, & Bosch, 2003) library. The output of the parser is labeled dependency graphs, using as arc labels a combination of bracket labels and grammatical role labels constructed from the Penn Treebank II annotation scheme (Marcus, Kim, et al., 1994). The parser reaches a maximum unlabeled attachment score of 87.1% and produces labeled dependency graphs with an accuracy of of 86.0% with the correct head and arc label recognised. The results are close to the state of the art in dependency parsing, and the parser also outputs arc labels that other parsers do not produce.
Knowledge acquisition is always regarded as a bottleneck in many NLP tasks, such as machine translation, information extraction. Treebank-based statistical parsing is not an exceptant. The latent linguistic knowledge in treebank is very rich, which, however, cant be acquired directly.In our model, the following three ways are used to incorporate such rich linguistic features for Chinese statistical parsing. First of all, non-recursive noun and verb phrases are annotated in the Penn Chinese Treebank because of their strong mark of boundaries. Second, a new head percolation table is designed based on Xias table. The last linguistic feature our model uses is the context configuration frame which provides a stronger representation of bilexical dependency structures. All these three linguistic features gain an improvement of remarkable 2.37% in terms of F1 measure, 5.36% in terms of complete match ratio.
This paper investigates automatic identification of Information Structure (IS) in texts. The experiments use the Prague Dependency Treebank which is annotated with IS following the Praguian approach of Topic Focus Articulation. We automatically detect t(opic) and f(ocus), using node attributes from the treebank as basic features and derived features inspired by the annotation guidelines. We present the performance of decision trees (C4.5), maximum entropy, and rule induction (RIPPER) classifiers on all tectogrammatical nodes. We compare the results against a baseline system that always assigns f(ocus) and against a rule-based system. The best system achieves an accuracy of 90.69%, which is a 44.73% improvement over the baseline (62.66%).
Dieser Beitrag beschäftigt sich mit sprachlichen Mischformen im frankophonen Kanada. Im Mittelpunkt steht die hybride Varietät des Chiac, eine in Moncton/Acadie gesprochene urbane Mischvarietät, die aus dem Sprachkontakt von Englisch und Französisch entstanden ist und auf einer französischen Grammatik mit englischen Lexikelementen basiert. Gegenstand der Betrachtung ist die in den späten 1960er Jahren entstandene littérature acadienne und ihr Umgang mit sprachlicher Variation im Spannungsfeld von gesellschaftlicher Norm und Formen ihrer Transgression. Am Beispiel der zeitgenössischen Autorin France Daigle und ihrem Umgang mit der Varietät des Chiac soll verdeutlicht werden, in welcher Weise hybride Formen von Sprache, die existente Sprach- und Kulturgrenzen in Frage stellen, dazu beitragen, individuelle und soziale Widersprüche zu erfassen und zu bearbeiten.