Papers reviewed and determined not to be word norm studies. Use the flag icon to report errors or suggest re-inclusion.
16504 papers
CT resulted in variable functional and structural changes in dementia, and conclusions are limited by heterogeneity and study quality. Larger, more robust studies are required to correlate these findings with clinical benefits from CT.
In her article Magdalena Edut presented a new type of the lexicographical description which is based on the methodology of the object-oriented approach (approche orientee-objets). In the first part of the text she discussed, on the example of the word mżawka (drizzle), which belongs to the category of the “natural phenomena”, the principles and aims of the proposed approach. In the following parts she compared the received description with the description suggested by the lexical database WordNet and with the contemporary lexicographical theories dealing with the issue of the automatic translation: - prof. G. Gross’s idea of an electronic dictionary, - I. Mel’cuk’s and A. Zolkovsky’s model Sens-Texte, - Pustejovsky’s and Boguraev’s structure qualia.
Multilingual dependency parsing is gaining popularity in recent years for several reasons. Dependency structures are more adequate for languages with freer word order than the traditional constituency notion. There is a growing availability of dependency treebanks for new languages. Broad coverage statistical dependency parsers are available and easily portable to new languages. Dependency parsing can provide useful contributions in areas such as information extraction, machine translation and question answering, among others. In addition, syntactic head-dependent pairs are a good interface between the traditional phrase structures and semantic theta roles. In this paper we present the learning curves of a statistical dependency parser for four languages: Arabic, Bulgarian, Italian and Slovene. We discuss issues that mostly concern the employed annotation scheme for each treebank with an emphasis on coordinated structures. Povzetek: Opisano je večjezično odvisnostno skladenjsko razčlenjevanje štirih jezikov. 1
Most of the work on treebank-based statistical parsing exclusively uses the Wall-Street-Journal part of the Penn treebank for evaluation purposes. Due to the presence of this quasi-standard, the question of to which degree parsing results depend on the properties of treebanks was often ignored. In this paper, we use two similar German treebanks, TüBa-D/Z and NeGra, and investigate the role that different annotation decisions play for parsing. For these purposes, we approximate the two treebanks by gradually taking out or inserting the corresponding annotation components and test the performance of a standard PCFG parser on all treebank versions. Our results give an indication of which structures are favorable for parsing and which ones are not.
The Hamburg implementation of the Weighted Constraint Dependency Grammar formalism (WCDG) includes an example grammar with comprehensive coverage for written German. This manual is the annotation guideline that was used to define the goals of the grammar and to create the Hamburg Dependency Treebank also published in the course of this project.
Each year the Conference on Computational Natural Language Learning (CoNLL) features a shared task, in which participants train and test their systems on exactly the same data sets, in order to better compare systems. The tenth CoNLL (CoNLL-X) saw a shared task on Multilingual Dependency Parsing. In this paper, we describe how treebanks for 13 languages were converted into the same dependency format and how parsing performance was measured. We also give an overview of the parsing approaches that participants took and the results that they achieved. Finally, we try to draw general conclusions about multi-lingual parsing: What makes a particular language, treebank or annotation scheme easier or harder to parse and which phenomena are challenging for any dependency parser?
Summary form only given. We present first some general remarks on challenges faced by modern information technology, notably when a human being is a relevant factor. These challenges are mainly related to inherent difficulties in solving some "meta-problems", in particular broadly perceived decision making. We assume, on the one hand, business intelligence related perspective, augmented with elements of Web intelligence, to fully use all available tools and resources. On the other hand, we assume a human centric computing perspective in the spirit of, for instance, Dertouzos's ideas. First, we present a brief account of modem approaches to real world decision making, emphasize the concept of a decision making process that involves more factors and aspects like: the use of own and external knowledge, involvement of various,actors, aspects, etc., individual habitual domains, non-trivial rationality, different paradigms. As an example we mention Checkland's deliberative decision making (which is an important elements of his soft approach to systems analysis). After an analysis of specifics and difficulties encountered in many real world decision-making situations, we strongly advocate the use of computer based decision support systems. First, we briefly review the history of decision support systems, and then present a popular classification, starting from data driven to Web based and inter-organizational. We indicate that decision support systems should incorporate some sort of "intelligence", and we first briefly mention some views of what intelligence may mean in this concept, and then assume some more pragmatic, though limited, view of intelligent decision support systems. We indicate possible advantages of using elements of fuzzy logic and soft computing, notably, Zadeh's computing with words to be able to somehow merge the ideas presented like: human centric computing, decision making processes, intelligent decision support, etc. Finally, we present an example of implementation in which the above-mentioned ideas have been to some extent implemented. This concern a data and document driven decision support system for a small to medium company in which, first, Zadeh's computing with words and perceptions paradigm is employed via linguistic database summaries, elements of Web intelligence are used to derive additional information, and the ideas of an intelligent decision support and human centric computing are shown to be synergistically combined. We finish with some general remarks emphasizing that fuzzy logic and soft computing, notably as exemplified by Zadeh's computing with words and perceptions may be viewed as providing just the right tools to solve the problems considered
This presentation reports the methodology followed and the results attained on an on-going project aiming at building a large lexical database of corpus-extracted multiword (MW) expressions for the Portuguese language. MW expressions were automatically extracted from a balanced 50 million word corpus compiled for this project, furthermore statistically interpreted using lexical association measures and are undergoing a manual validation process. The lexical database covers different types of MW expressions, from named entities to lexical associations with different degrees of cohesion, ranging from totally frozen idioms to favoured co-occurring forms, like collocations. We aim to achieve two main objectives with this resource: to build on the large set of data of different types of MW expressions to revise existing typologies of collocations and to integrate them in a larger theory of MW units; to use the extensive hand-checked data as training data to evaluate existing statistical lexical association measures.
Preface Part I: Background Chapter 1: Factors Influencing the Acquisition and Refinement of Communication Skills. Chapter 2: Communication Access: Overview and Issues Chapter 3: Issues in Assessment and Intervention Part II Chapter 4: Pre-Language Communication Chapter 5: Pre-Language Assessment and Intervention Chapter 6: Later Language Development Chapter 7: Adolescent Language: Assessment and Intervention Appendix A: Resource List for Children with Auditory Processing Disorders (APD) Appendix B: Troubleshooting Technology Appendix C: Technology Resources Appendix D: Recommendations for Comprehensive Assessment: Pragmatics and Semantics. Appendix E: Familiarity and Transparency Ratings for 100 Idioms Appendix F: Familiarity Ratings for 107 Proverbs Appendix G: Pronunciation Skill Inventory Index
Motivationally relevant stimuli have been shown to receive prioritized processing compared to neutral stimuli at distinct processing stages. This effect has been related to the evolutionary importance of rapidly detecting dangers and potential rewards and has been shown to be modulated by the distance between an organism and a faced stimulus. Similarly, recent studies showed degrees of emotional modulation of autonomic responses and subjective arousal ratings depending on stimulus size. In the present study, affective modulation of pictures presented in different sizes was investigated by measuring event-related potentials during a two-choice categorization task. Results showed significant emotional modulation across all sizes at both earlier and later stages of processing. Moreover, affective modulation of earlier processes was reduced in smaller compared to larger sizes, whereas no changes in affective modulation were observed at later stages.
This paper explores techniques to take advantage of the fundamental difference in structure between hidden Markov models (HMM) and hierarchical hidden Markov models (HHMM). The HHMM structure allows repeated parts of the model to be merged together. A merged model takes advantage of the recurring patterns within the hierarchy, and the clusters that exist in some sequences of observations, in order to increase the extraction accuracy. This paper also presents a new technique for reconstructing grammar rules automatically. This work builds on the idea of combining a phrase extraction method with HHMM to expose patterns within English text. The reconstruction is then used to simplify the complex structure of an HHMM The models discussed here are evaluated by applying them to natural language tasks based on CoNLL-2004 1 and a sub-corpus of the Lancaster Treebank 2.
Body image has been shown to be influenced by weight loss. Little attention however has been devoted to personal evaluations of physical fitness (fitness evaluation) or the extent to which individuals psychologically invest in improving fitness (fitness orientation) during periods of weight loss. PURPOSE: The purpose of the present study was to examine the relation between fitness evaluation (FE) and fitness orientation (FO) subscales with other body image ratings and physiological measures of fitness during a twelve-week behavioral weight loss program. METHODS: Thirty overweight, sedentary women (age= 42.5 ± 8.5 years, BMI= 29.3 kg/m2 ± 3.2) participated in a twelve-week behavioral weight loss program which reduced energy intake to 1200–1500 kcal/day and dietary fat to <30% of total calories. Subjects were progressively increased to 40 min/day, 5 days/week of home-based walking exercise. Body image was assessed using Multidimensional Body-Self Relations Questionnaire (MBSRQ) subscales. Cardiorespiratory fitness was measured using time to reach 85% of age-predicted maximal heart rate during a submaximal test on a treadmill. RESULTS: Time to reach 85% of age-predicted maximal heart rate significantly increased over the 12 week intervention (10.3+3.0 min vs. 12.5+3.4 min, p< 0.00). Baseline FE and FO were positively associated with baseline cardiorespiratory fitness (r=0.55 and 0.53, respectively). Change in FO was positively associated with baseline FO, changes in appearance evaluation and orientation and change in cardiorespiratory fitness (r=0.40, 0.65, 0.37 and 0.39, respectively). Furthermore, positive change in FO was associated with 12-week scores of appearance orientation, health orientation and overweight preoccupation (r=0.38, 0.37, and 0.51, respectively). CONCLUSION: These findings demonstrate that individuals with higher scores of FO at baseline had the greatest change in FO at 12 weeks, which was positively associated with changes in other perceived personal fitness constructs and fitness improvements. Exercise interventions should target strategies for improving FO and FE as this may have implications for exercise adoption and maintenance during weight loss. Supported by a Research Incentive Grant from the University of Louisville
PURPOSE: To determine observer performance in the detection of multiple sclerosis (MS) lesions on magnetic resonance (MR) images of the brain and to assess the dependence of observer performance on lesion size, parenchymal location, pulse sequence, and supratentorial versus infratentorial level. MATERIALS AND METHODS: This HIPAA-compliant protocol was approved by the institutional review board, and previously acquired MR data from a healthy volunteer and a patient with MS were used to derive parameter maps, with waiver of informed consent. Parameter maps and image simulator software were used to generate 320 phantom brain images with simulated supratentorial and infratentorial MS lesions. Images were displayed with T2-weighting or fluid-attenuated inversion recovery (FLAIR) contrast. Four readers independently evaluated the images, rating lesions on a five-point certainty scale. Observer performance was measured by using the area under the alternative free-response receiver operating characteristic curve (A(1)), and significance was determined with the z test. RESULTS: Pooled A(1) scores were significantly better for FLAIR imaging (0.96 +/- 0.01 [standard error]) than for T2-weighted MR imaging (0.89 +/- 0.04) supratentorially (P =.05) but were similar for FLAIR imaging (0.90 +/- 0.06) and T2-weighted MR imaging (0.88 +/- 0.05) infratentorially. A(1) scores for cortical, deep white matter, and periventricular lesions were 0.93 +/- 0.05, 0.97 +/- 0.02, and 0.89 +/- 0.04, respectively, for FLAIR imaging and 0.77 +/- 0.06, 0.99 +/- 0.01, and 0.89 +/- 0.05, respectively, for T2-weighted MR imaging. FLAIR scores were significantly higher than T2-weighted scores for cortical lesions. Linear correlation was found between A(1) and lesion size (r = 0.5). CONCLUSION: Supratentorially, performance was better with FLAIR imaging than with T2-weighted MR imaging. Infratentorially, performance was moderate with both modalities. Observers did better with FLAIR imaging in the detection of cortical lesions, and performance improved with increasing lesion size.
Both classroom instruction and lexical database development stand to benefit from applied research on sign language, which takes into consideration American Sign Language rules, pedagogical issues, and teacher characteristics. In this study of technical science signs, teachers' experience with signing and, especially, knowledge of content, were found to be essential for the identification of signs appropriate for instruction. The results of this study also indicate a need for a systematic approach to examine both sign selection and its impact on learning by deaf students. Recommendations are made for the development of lexical databases and areas of research for optimizing the use of sign language in instruction.
We present a novel PCFG-based architecture for robust probabilistic generation based on wide-coverage LFG approximations (Cahill et al., 2004) automatically extracted from treebanks, maximising the probability of a tree given an f-structure. We evaluate our approach using string-based evaluation. We currently achieve coverage of 95.26%, a BLEU score of 0.7227 and string accuracy of 0.7476 on the Penn-II WSJ Section 23 sentences of length ≤20.
DepAnn is an interactive annotation tool for dependency treebanks, providing both graphical and text-based annotation interfaces. The tool is aimed for semi-automatic creation of treebanks. It aids the manual inspection and correction of automatically created parses, making the annotation process faster and less error-prone. A novel feature of the tool is that it enables the user to view outputs from several parsers as the basis for creating the final tree to be saved to the treebank. DepAnn uses TIGER-XML, an XML-based general encoding format for both, representing the parser outputs and saving the annotated treebank. The tool includes an automatic consistency checker for sentence structures. In addition, the tool enables users to build structures manually, add comments on the annotations, modify the tagsets, and mark sentences for further revision.
We built a morphological analyzer, which can be freely used by anyone for research purpose. In order to build a pratical system, a dictionary with reasonable size is necessary. The initial dictionary is built from the Penn Chinese Treebank corpus v4.0 and contains only 33,438 entries. Since the initial dictionary is quite small, unknown word detection methods are applied to a huge raw text in order to extract new words to be added into the system dictionary. We have successfully constructed a dictionary with 120,769 entries. Finally, we propose a two-layer morphological analyzer to cater for two sets of outputs. The first layer produces the minimal segmentation units defined by us, and the second layer transforms the output of the first layer to the original segmentation units defined by Penn Chinese Treebank.
This paper explores several important issues in developing syntactically annotated Korean corpora for higher-level language processing, including semantic-discourse parsing, question-answering, machine translation, information retrieval, etc.In particular, we compare the Penn Korean Treebank (PKT) and the Korean Treebank of the 21st Century Sejong Project (ST) and discuss four critical issues in syntactic annotation.We argue for the use of more sophisticated morphosyntactic information, and based on our comparative study, we propose revisions in the syntactic annotation schemes of the existing Korean Treebanks in order to improve the quality of annotated corpora and their usability both for conducting theoretical research and for developing computational tools.The results of our comparative study reveal four significant issues in syntactic annotations: the syntactic analysis of verbal complexes, the hierarchical structure of noun phrases, the representation of traces, and the marking of zero elements.These factors may trigger erroneous syntactic representations for certain linguistic phenomena, and they may increase difficulties in data search and lessen reliability in computational processing.Thus, evaluating and improving the syntactic annotation of Treebanks is an important task for aspects of both theoretical and computational linguistics.
In this work we learn clusters of contextual annotations for non-terminals in the Penn Treebank. Perhaps the best way to think about this problem is to contrast our work with that of That research used treetransformations to create various grammars with different contextual annotations on the non-terminals. These grammars were then used in conjunction with a CKY parser. The authors explored the space of different annotation combinations by hand. Here we try to automate the process -to learn the "right" combination automatically. Our results are not quite as good as those carefully created by hand, but they are close (84.8 vs 85.7).
The paper focuses on the description of the process of “mining” lexical semantic information from published dictionaries of Brazilian Portuguese language. Specifically, it is described the manual approach of compiling and filtering hyperonymy/hyponymy and holonymy/meronym logic-conceptual relations of concrete nouns. These relations, filtering from the dictionaries, will be use to organize part of the nouns of Wordnet.Br lexical database.
In this paper, current dependencybased treebanks are introduced and analyzed.The methods used for building the resources, the annotation schemes applied, and the tools used (such as POS taggers, parsers and annotation software) are discussed.
Abstract Facial masculinity may be used as a cue in female mate choice, as it reflects the success of the male genotype in its developmental environment. Women may maximize reproductive success by using a conditional strategy favoring highly masculine facial features for short‐term relationships and feminized facial features in men for long‐term relationships. Three studies examine reactions to masculinized and feminized male facial composites. Properties of the original composite image affect ratings of critical attributes and the magnitude of the differences in ratings between versions undergoing identical processes of geometric manipulation (Study 1). Both men and women attribute personality, behavior, and mating strategies consistent with predictions derived from the good genes and mating trade‐off hypotheses (Study 2). Participants accurately grouped behavioral tendencies related to high mating effort/risky strategies and high parenting effort/risk adverse strategies and associated mating effort more so with masculinized faces and parenting effort more so with feminized faces (Study 3). These results indicate that male facial masculinity serves as a visual cue for inferring personality and reproductive strategy.
An emerging task in text understanding and generation is to categorize information as fact or opinion and to further attribute it to the appropriate source. Corpus annotation schemes aim to encode such distinctions for NLP applications concerned with such tasks, such as information extraction, question answering, summarization, and generation. We describe an annotation scheme for marking the attribution of abstract objects such as propositions, facts and eventualities associated with discourse relations and their arguments annotated in the Penn Discourse TreeBank. The scheme aims to capture the source and degrees of factuality of the abstract objects. Key aspects of the scheme are annotation of the text spans signalling the attribution, and annotation of features recording the source, type, scopal polarity, and determinacy of attribution.
Recent research on causal learning found (a) that causal judgments reflect either the current predictive value of a conditional stimulus (CS) or an integration across the experimental contingencies used in the entire experiment and (b) that postexperimental judgments, rather than the CS's current predictive value, are likely to reflect this integration. In the current study, the authors examined whether verbal valence ratings were subject to similar integration. Assessments of stimulus valence and contingencies responded similarly to variations of reporting requirements, contingency reversal, and extinction, reflecting either current or integrated values. However, affective learning required more trials to reflect a contingency change than did contingency judgments. The integration of valence assessments across training and the fact that affective learning is slow to reflect contingency changes can provide an alternative interpretation for researchers' previous failures to find an effect of extinction training on verbal reports of CS valence.
The Dutch spelling system, like other European spelling systems, represents a certain balance between preserving the spelling of morphemes (the morphological principle) and obeying letter-to-sound regularities (the phonological principle). We present experimental results with artificial learners that show a competition effect between the two principles: adhering more to one principle leads to more violations of the other. The artificial learners, memory-based learning algorithms, are trained (1) to convert written words to their phonemic counterparts and (2) to analyze written words on their morphological composition, based on data extracted from the CELEX lexical database. As an exception to the competition effect we show that introducing the schwa as a letter in the spelling system causes both morphology and phonology to be learnt better by the artificial learners. In general we argue that artificial learning studies are a tool in obtaining objective measurements on a spelling system that may be of help in spelling reform processes.
Data-driven grammatical function tag assignment has been studied for English using the Penn-II Treebank data. In this paper we address the question of whether such methods can be applied successfully to other languages and treebank resources. In addition to tag assignment accuracy and f-scores we also present results of a task-based evaluation. We use three machine-learning methods to assign Cast3LB function tags to sentences parsed with Bikel's parser trained on the Cast3LB treebank. The best performing method, SVM, achieves an f-score of 86.87% on gold-standard trees and 66.67% on parser output - a statistically significant improvement of 6.74% over the baseline. In a task-based evaluation we generate LFG functional-structures from the function-tag-enriched trees. On this task we achive an f-score of 75.67%, a statistically significant 3.4% improvement over the baseline.
In this paper, we present a generic analysis of etymological data intended to provide a uniform framework for modeling such information in lexical databases. Based on the explicit reification of etymons and of the links between them, our proposal provides means to state additional constraints for both, as well as a preliminary set of standard descriptors to be used to this end. The model has been extensively tested on a variety of concrete cases extracted from the TLFi (Trésor de la Langue Française informatisé), allowing us to identify further mechanisms (alternatives, multiple links, composition, etc.) needed for etymological data representation. Taking into account the current standardization efforts within ISO committee TC 37/SC 4 to define a specification platform for lexical data (aka LMF, Lexical Markup Framework), we try to show that our proposal could be a possible contribution to this project
Previous articleNext article FreeCurrent ApplicationsLinguistic AnthropologyV.ChandV.Chand Search for more articles by this author PDFPDF PLUSFull Text Add to favoritesDownload CitationTrack CitationsPermissionsReprints Share onFacebookTwitterLinked InRedditEmailQR Code SectionsMoreImmigration Practices in Belgium: African AsylumSeeker DiscourseJan Blommaert, a Belgian anthropologist currently at the Institute of Education, University of London, is known in Belgium as a public campaigner on immigration issues. He became involved in African asylumseekers rights in 1998, when the death of Semira Adamu during her forced repatriation provoked a public outcry over the implementation of Belgian immigration policies, in which more than 95% of applicants for asylum are rejected. Adamu had fled Nigeria in her teens to avoid entering into a polygamous marriage with a 65yearold man. Shackled at the ankles and vigorously resisting, she was suffocated by Belgian police attempting to restrain her.Poster for a 2003 commemoration of Adamu's death.View Large ImageDownload PowerPointAdamus death was a catalyst for the formation of new action groups, and existing organizations also became involved: churches opened their doors to asylum seekers, and NGOs such as OXFAM and the League of Human Rights campaigned for asylum seekers rights. Early on, Blommaert contributed to this organic collaborative effort: I gave tons of public lectures for any audience willing to listen, wrote opeds in major newspapers, campaigned with MPs close to the government, participated in public debates on these matters, and wrote expert articles for a wider audience. Additionally, he saw that his understanding of the legal and sociolinguistic issues was directly applicable to the problem. He was motivated, he explains, by awareness of the real stakes and real people involved.African asylum seekers face circumstances not shared with those from Europe and the Gulf because they come from wartorn areas with unclear state boundaries, may be illiterate and lack documentation, have long migration paths to Belgium, and do not share a language with immigration authorities. In particular, Blommaert points out, their choices of language, background texts, and genres for storytelling affect their chances of acquiring refugee status. African asylumseeker language is stereotypically filled with language mixing, language impurities, varying degrees of literacy, and different styles of storytelling and discourse. When asylum seekers present their stories, they are unaware of how the officials are judging them on their linguistic habits. The officials note these details and find them inconsistent with Belgian expectations; they judge the asylum seekers in terms of these nave choices and reject them. In many such cases, rejection is a matter of life or death for the refugees, given the risks associated with deportation and with repatriation into the home country that originally motivated them to seek asylum.Blommaert has mobilized an informal network of Belgian academics and institutions that has produced many academically informed public statements. Additionally, by documenting and analyzing oral immigration interviews with an eye to understanding the disconnect between asylum seekers presentations and immigration officials expectations, he has improved practice in the Immigration Department. He has used his analysis to train members of the department, helping to adjust views of what can be gathered in an immigration interview, to improve interview techniques, and to promote an awareness of variability in sociocultural linguistic norms and presentation styles.He continues to work as an official expert for Belgian legal, government, and security authorities, translating and analyzing documentation and corroboratory written texts provided for immigration procedures and advising on specific issues. He has been able to influence outcomes for some individual asylum seekers. His public campaigning has raised awareness of the issues surrounding African asylum seekers, and his publications have provoked work within academia that may lead to further interventions by academics with regard to the interview process.Blommaert hopes that public campaigning by NGOs and academic scholarship will promote continued dialogue with immigration officials, eventually producing policies and procedures that take into account the politics of migration and displacement and the way in which asylum seekers frame their life stories. Previous articleNext article DetailsFiguresReferencesCited by Current Anthropology Volume 47, Number 3June 2006 Sponsored by the Wenner-Gren Foundation for Anthropological Research Article DOIhttps://doi.org/10.1086/504161 Views: 287Total views on this site Citations: 1Citations are reported from Crossref PDF download Crossref reports the following articles citing this article:Kevin D. O'Gorman, Cailein Gillespie The mythological power of hospitality leaders?, International Journal of Contemporary Hospitality Management 22, no.55 (Jul 2010): 659–680.https://doi.org/10.1108/09596111011053792
research-article Free Access Share on The treebanks used in the shared taskCoNLL-X '06: Proceedings of the Tenth Conference on Computational Natural Language LearningJune 2006 Pages 165Online:08 June 2006Publication History 0citation107DownloadsMetricsTotal Citations0Total Downloads107Last 12 Months35Last 6 weeks0 Get Citation AlertsNew Citation Alert added!This alert has been successfully added and will be sent to:You will be notified whenever a record that you have chosen has been cited.To manage your alert preferences, click on the button below.Manage my AlertsNew Citation Alert!Please log in to your account Save to BinderSave to BinderCreate a New BinderNameCancelCreateExport CitationPublisher SiteeReaderPDF
Automatic text analysis systems can lexically recognize a word only if it already exists in the electronic dictionary. The same thing is true for the NOOJ system analysis programs. One understands here by electronic dictionaries the lexical databases where all information is explicit because they are intended for computer programs use. These bases aim at the modelling of the language, which distinguishes them from the electronic lexicons created for particular applications needs. For technical languages or of speciality, work remains to be made to build dictionaries. Our work applies to the NOOJ system, with for immediate objective the installation of an electronic dictionary of French computer science terms (compound words), with an aim of analysing, automatically indexing texts. The linguistic aspects of the terminology are retained.
We present the implementation of a system which extracts not only lexicalized grammars but also feature-based lexicalized grammars from Korean Sejong Treebank. We report on some practical experiments where we extract TAG grammars and tree schemata. Above all, full-scale syntactic tags and well-formed morphological analysis in Sejong Treebank allow us to extract syntactic features. In addition, we modify Treebank for extracting lexicalized grammars and convert lexicalized grammars into tree schemata to resolve limited lexical coverage problem of extracted lexicalized grammars.
Recently, the Prague Dependency Treebank 2.0 (PDT 2.0) has emerged as the largest text corpora annotated on the level of tectogrammatical representation (“linguistic meaning”) described in Sgall et al. (2004) and containing about 0.8 milion words (see Hajič (2004)). We hope that this level of annotation is so close to the meaning of the utterances contained in the corpora that it should enable us to automatically transform texts contained in the corpora to the form of knowledge base, usable for information extraction, question answering, summarization, etc. We can use Multilayered Extended Semantic Networks (MultiNet) described in Helbig (2006) as the target formalism. In this paper we discuss the suitability of such approach and some of the main issues that will arise in the process. In section 1. we introduce formalisms underlying PDT 2.0 and MultiNet, in section 2. we describe the role MultiNet can play in the system of Functional Generative Description (FGD), section 3. discusses issues of automatic conversion to MultiNet and section 4. gives some conclusions. 1.
Statistical parsers trained and tested on the Penn Wall Street Journal (WSJ) treebank have shown vast improvements over the last 10 years. Much of this improvement, however, is based upon an ever-increasing number of features to be trained on (typically) the WSJ treebank data. This has led to concern that such parsers may be too finely tuned to this corpus at the expense of portability to other genres. Such worries have merit. The standard "Charniak parser" checks in at a labeled precision-recall f-measure of 89.7% on the Penn WSJ test set, but only 82.9% on the test set from the Brown treebank corpus.This paper should allay these fears. In particular, we show that the reranking parser described in Charniak and Johnson (2005) improves performance of the parser on Brown to 85.2%. Furthermore, use of the self-training techniques described in (McClosky et al., 2006) raise this to 87.8% (an error reduction of 28%) again without any use of labeled Brown data. This is remarkable since training the parser and reranker on labeled Brown data achieves only 88.4%.
Deterministic parsing guided by treebank-induced classifiers has emerged as a simple and efficient alternative to more complex models for data-driven parsing. We present a systematic comparison of memory-based learning (MBL) and support vector machines (SVM) for inducing classifiers for deterministic dependency parsing, using data from Chinese, English and Swedish, together with a variety of different feature models. The comparison shows that SVM gives higher accuracy for richly articulated feature models across all languages, albeit with considerably longer training times. The results also confirm that classifier-based deterministic parsing can achieve parsing accuracy very close to the best results reported for more complex parsing models.
Purpose This paper examined employee perceptions of the rewards associated with their participation in a six sigma program. Six sigma is an approach to organizational change that incorporates elements of total quality management, business process reengineering, and employee involvement. Design/methodology/approach A survey was completed by 215 employees (34 percent response rate). Respondents rated the extent to which they felt their participation in six sigma was “instrumental” for a range of outcomes, as well as valence (desirability) of each outcome (based on the VIE concept of instrumentality). The outcomes were classified into four categories: extrinsic, intrinsic, social, and organizational. Findings Valence ratings revealed that all 12 outcomes were perceived as desirable. Instrumentality ratings showed that extrinsic outcomes were rated significantly lower than intrinsic, social, and organizational outcomes. Additional analyses revealed significant differences on all four outcome categories between participants and non‐participants in the six sigma program. Practical implications The positive valence and instrumentality ratings for participants indicate they believe their participation will lead to valued outcomes for themselves and their organizations. However, employees who choose not to get involved in six sigma do not perceive that their participation would have led to desired outcomes. The results also show that while participants value extrinsic rewards, they do not see six sigma as instrumental in their receipt. These perceptions have important implications for attracting and retaining program participants. Originality/value While much has been written about the use of reward systems in supporting a successful six sigma effort, this study empirically examines how employees actually perceive the rewards associated with their participation. It also identifies which types of rewards are most instrumental for participants and non‐participants.