Papers reviewed and determined not to be word norm studies. Use the flag icon to report errors or suggest re-inclusion.
18265 papers
Internet has opened the access to an overwhelming amount of data, requiring the development of new applications to automatically recognize, process and manage informationavailable in web sites or web-based applications. The standardSemantic Web architecture exploits ontologies to give a shared(and known) meaning to each web source elements.In this context, we developed MELIS (Meaning Elicitation and Lexical Integration System). MELIS couples the lexical annotation module of the MOMIS system with some components from CTXMATCH2.0, a tool for eliciting meaning from severaltypes of schemas and match them. MELIS uses the MOMIS WNEditor and CTXMATCH2.0 to support two main tasks in theMOMIS ontology generation methodology: the source annotationprocess, i.e. the operation of associating an element of a lexicaldatabase to each source element, and the extraction of lexicalrelationships among elements of different data sources.
In this paper, we construct a biomedical semantic role labeling (SRL) system that can be used to facilitate relation extraction. First, we construct a proposition bank on top of the popular biomedical GENIA treebank following the PropBank annotation scheme. We only annotate the predicate-argument structures (PAS's) of thirty frequently used biomedical predicates and their corresponding arguments. Second, we use our proposition bank to train a biomedical SRL system, which uses a maximum entropy (ME) model. Thirdly, we automatically generate argument-type templates which can be used to improve classification of biomedical argument types. Our experimental results show that a newswire SRL system that achieves an F-score of 86.29% in the newswire domain can maintain an F-score of 64.64% when ported to the biomedical domain. By using our annotated biomedical corpus, we can increase that F-score by 22.9%. Adding automatically generated template features further increases overall F-score by 0.47% and adjunct arguments (AM) F-score by 1.57%, respectively.
This paper presents a Chinese dependency syntax for treebanking. The syntax contains 13 word classes and 34 dependency types. A format of treebank based on the syntax is also proposed for the applications of computational and general linguistic research. Some experiments show that the treebank based on the proposed dependency syntax can be used for training and evaluating the dependency parser and for quantitative analysis of Chinese syntax.
There are many methods to improve performances of statistical parsers. Among them, resolving structural ambiguities is a major task. In our approach, the parser produces a set of n-best trees based on a feature-extended PCFG grammar and then selects the best tree structure based on association strengths of dependency word-pairs. However, there is no sufficiently large Treebank producing reliable statistical distributions of all word-pairs. This paper aims to provide a self-learning method to resolve the problems. The word association strengths were automatically extracted and learned by parsing a giga-word corpus. Although the automatically learned word associations were not perfect, the built structure evaluation model improved the bracketed f-score from 83.09 % to 86.59%. We believe that the above iterative learning processes can improve parsing performances automatically by learning word-dependence knowledge continuously from web. 1.
The aim of this research was to study the influence of both the emotional content and the physical characteristics of affective stimuli on the psychophysiological, behavioral and cognitive indexes of the emotional response. We selected 54 pictures from the IAPS, depicting unpleasant, neutral, and pleasant contents, and used two picture sizes as experimental conditions (120 x 90 cm and 52 x 42 cm). Sixty-one subjects were randomly assigned to each experimental condition. We recorded the startle blink reflex, skin conductance response, heart rate, free viewing time, and picture valence and arousal ratings. In line with previous research (e.g., Bradley, Codispoti, Cuthbert, and Lang, 2001), our data showed an effect of the affective content on all the measurements recorded. Importantly, effects of the size of the affective pictures on emotional responses were not found, indicating that the emotional content is more important than the formal properties of the stimuli in evoking the emotional response.
While syntactically annotated corpora known as treebanks have been available for many years, along with a variety of customized tools for querying these annota-
TransBooster is a wrapper technology designed to improve the performance of wide-coverage machine translation \nsystems. Using linguistically motivated syntactic information, it automatically decomposes source language sentences into shorter and syntactically simpler chunks, and recomposes their translation to form target language sentences. This generally improves both the word order \nand lexical selection of the translation. To date, TransBooster has been successfully applied to rule-based MT, statistical MT, and multi-engine MT. This paper presents \nthe application of TransBooster to Example-Based Machine Translation. In an experiment conducted on test sets \nextracted from Europarl and the Penn II Treebank we show that our method can raise the BLEU score up to 3.8% relative \nto the EBMT baseline. We also conduct a manual evaluation, showing that TransBooster-enhanced EBMT produces \na better output in terms of fluency than the baseline EBMT in 55% of the cases and in terms of accuracy in 53% of the \ncases.
Based on text chunking using HMM, transformation-based learning is made use of to improve the precision of chunk tags further. The training data and the test data are from Penn treebank 4.0, and 13 text chunks are used. Rules are learned automatically according to the rule templates. The precision is improved 4.48%. The detailed analysis that affects the text chunking is given. Different threshold, different scale of training data, different learning equation and different rule templates can affect the precision of the text chunking.
While the processing of verbal and psychophysiological indices of emotional arousal have been investigated extensively in relation to the left and right cerebral hemispheres, it remains poorly understood how both hemispheres normally function together to generate emotional responses to stimuli. Drawing on a unique sample of nine high-functioning subjects with complete agenesis of the corpus callosum (AgCC), we investigated this issue using standardized emotional visual stimuli. Compared to healthy controls, subjects with AgCC showed a larger variance in their cognitive ratings of valence and arousal, and an insensitivity to the emotion category of the stimuli, especially for negatively-valenced stimuli, and especially for their arousal. Despite their impaired cognitive ratings of arousal, some subjects with AgCC showed large skin-conductance responses, and in general skin-conductance responses discriminated emotion categories and correlated with stimulus arousal ratings. We suggest that largely intact right hemisphere mechanisms can support psychophysiological emotional responses, but that the lack of interhemispheric communication between the hemispheres, perhaps together with dysfunction of the anterior cingulate cortex, interferes with normal verbal ratings of arousal, a mechanism in line with some models of alexithymia.
The Surrey Cross-linguistic Database on Deponency encodes information on the presence of morphological mismatches in a controlled sample of genetically and geographically diverse languages (based on the 100-language sample from the World Atlas of Language Structures). The database was created for the project 'Extended Deponency: The right morphology in the wrong place', funded by the Economic and Social Research Council under grant number RES000230375.
In the past, a divide could be seen between ’deep ’ parsers on the one hand, which construct a semantic representation out of their input, but usually have significant coverage problems, and more robust parsers on the other hand, which are usually based on a (statistical) model derived from a treebank and have larger coverage,
We present an improved approach for learning dependency parsers from tree-bank data. Our technique is based on two ideas for improving large margin training in the context of dependency parsing. First, we incorporate local constraints that enforce the correctness of each individual link, rather than just scoring the global parse tree. Second, to cope with sparse data, we smooth the lexical parameters according to their underlying word similarities using Laplacian Regularization. To demonstrate the benefits of our approach, we consider the problem of parsing Chinese treebank data using only lexical features, that is, without part-of-speech tags or grammatical categories. We achieve state of the art performance, improving upon current large margin approaches.
In this paper, we present a generic analysis of etymological data intended to provide a uniform framework for modeling such information in lexical databases. Based on the explicit reification of etymons and of the links between them, our proposal provides means to state additional constraints for both, as well as a preliminary set of standard descriptors to be used to this end. The model has been extensively tested on a variety of concrete cases extracted from the TLFi (Trésor de la Langue Française informatisé), allowing us to identify further mechanisms (alternatives, multiple links, composition, etc.) needed for etymological data representation. Taking into account the current standardization efforts within ISO committee TC 37/SC 4 to define a specification platform for lexical data (aka LMF, Lexical Markup Framework), we try to show that our proposal could be a possible contribution to this project
This study uses a spatial logit model to evaluate the statistical effect of conditions of communities on municipal bond ratings. It finds that private (non-farm) earnings in the community positively explain bond ratings with statistical significance, while earnings from personal transfers negatively affect ratings. Own source of revenues of local governments (local taxes) increase ratings and inter-governmental revenues (transfers to local governments) negatively impact ratings. Outstanding debt fails to significantly explain ratings. The composition of the local economy (e.g., the service sector) weights heavily in a rating, and proximity of the local government to areas with high municipal bond ratings increases ratings.
WordNet, a lexical database for English that is extensively used by computational linguists, has not previously distinguished hyponyms that are classes from hyponyms that are instances. This work describes an attempt to draw this distinction and reports the way in which the results were incorporated in the last version (2.1) of WordNet.
Recently the rational principle and the customary principle have been regarded as the main judging basis for linguistic norm.But the practice proves that the rational principle is hard to conduct and the customary principle less popularized.The establishment of the judging basis for linguistic norm should take the consideration of those target's characteristics.Langue and parole must be divided and ruled by their dissimilarity.Stability,sociality and pragmatic function should be the basis for langue,while effect principle and ethic principle should be the basis for parole.When dealing with the transitional area between langue and parole,the normal and the abnormal,we should not go to extremes.
In this paper, we describe our hybrid approach to two key NLP technologies: biomedical named entity recognition (Bio-NER) and (Bio-SRL). In Bio-NER, our system successfully integrates linguistic features into the CRF framework. In addition, we employ web lexicons and template-based post-processing to further boost its performance. Through these broad linguistic features and the nature of CRF, our system outperforms state-of-the-art machine-learning-based systems, especially in the recognition of protein names (F=78.5%). In Bio-SRL, first, we construct a proposition bank on top of the popular biomedical GENIA treebank following the PropBank annotation scheme. We only annotate the predicate-argument structures (PAS's) of thirty frequently used biomedical verbs (predicates) and their corresponding arguments. Second, we use our proposition bank to train a biomedical SRL system, which uses a maximum entropy (ME) machine-learning model. Thirdly, we automatically generate argument-type templates, which can be used to improve classification of biomedical argument roles. Our experimental results show that a newswire English SRL system that achieves an F-score of 86.29% in the newswire English domain can maintain an F-score of 64.64% when ported to the biomedical domain. By using our annotated biomedical corpus, we can increase that F-score by 22.9%. Adding automatically generated template features further increases overall F-score by 0.47% and adjunct (AM) F-score by 1.57%, respectively.
This paper explores techniques to take advantage of the fundamental difference in structure between hidden Markov models (HMM) and hierarchical hidden Markov models (HHMM). The HHMM structure allows repeated parts of the model to be merged together. A merged model takes advantage of the recurring patterns within the hierarchy, and the clusters that exist in some sequences of observations, in order to increase the extraction accuracy. This paper also presents a new technique for reconstructing grammar rules automatically. This work builds on the idea of combining a phrase extraction method with HHMM to expose patterns within English text. The reconstruction is then used to simplify the complex structure of an HHMM The models discussed here are evaluated by applying them to natural language tasks based on CoNLL-2004 1 and a sub-corpus of the Lancaster Treebank 2.
This paper reports the results of perception tests administered to speakers of Japanese as part of a cross-language investigation of how voice quality and f 0 combine in the signalling of affect. Three types of synthesised stimuli were presented: (1) 'VQ only' involving variations in voice quality and a neutral f 0; (2) 'f 0 only', with different f 0 contours and modal voice; and (3) combined 'VQ + f 0 ' stimuli, where combinations of (1) and (2) were employed. Overall, stimuli involving voice quality variation (1 and 3) proved to be most consistently associated with affect. In series (2) only stimuli with very high f 0 yielded high affective ratings. Some striking differences emerge in the ratings obtained for Japanese subjects compared to those obtained for speakers of Hiberno-English [7], suggesting that the generation of expressive speech synthesis will need to be sensitive to language specific uses of the voice.
Preface Part I: Background Chapter 1: Factors Influencing the Acquisition and Refinement of Communication Skills. Chapter 2: Communication Access: Overview and Issues Chapter 3: Issues in Assessment and Intervention Part II Chapter 4: Pre-Language Communication Chapter 5: Pre-Language Assessment and Intervention Chapter 6: Later Language Development Chapter 7: Adolescent Language: Assessment and Intervention Appendix A: Resource List for Children with Auditory Processing Disorders (APD) Appendix B: Troubleshooting Technology Appendix C: Technology Resources Appendix D: Recommendations for Comprehensive Assessment: Pragmatics and Semantics. Appendix E: Familiarity and Transparency Ratings for 100 Idioms Appendix F: Familiarity Ratings for 107 Proverbs Appendix G: Pronunciation Skill Inventory Index
This paper evaluates four of the most commonly used, freely available, state-of-the-art parsers on a standard benchmark as well as with respect to a set of data relevant for measuring text cohesion, as one example of a learning technology application that requires fast and accurate syntactic parsing. We outline advantages and disadvantages of existing technologies and make recommendations. Our performance report uses traditional measures based on a gold standard as well as novel dimensions for parsing evaluation. To our knowledge, this is the first attempt to evaluate parsers across genres and grade levels for the implementation in learning technology using both gold standard and directed evaluation methods.
This paper discusses theoretical and methodological problems associated with defining the linguistic norms which emerge from the prescriptive tradition of linguistic research. One acknowledged deficiency of this tradition is the difficulty it has in dealing with linguistic variation and change (section I.). The article will focus on the fundamental distinction between the static concept of explicit standard norms which ignore linguistic variation and change and the dynamic concept of implicit non-standard norms such as those of bilingual varieties which involve such variation and change (section II.). Section III. argues for a specific term of 'code-switching norms' which is able to simultaneously consider the principles of linguistic change and variability as well as that of inherent variability. Finally, it discusses methodological implications for a sampling model.
We introduce Talbanken05, a Swedish treebank based on a syntactically annotated corpus from the 1970s, Talbanken76, converted to modern formats. The treebank is available in three different formats, besides the original one: two versions of phrase structure annotation and one dependency-based annotation, all of which are encoded in XML. In this paper, we describe the conversion process and exemplify the available formats. The treebank is freely available for research and educational purposes. 1.
Abstract Attempting to automatically learn to identify verb complements from natural language corpora without the help of sophisticated linguistic resources like grammars, parsers or treebanks leads to a significant amount of noise in the data. In machine learning terms, where learning from examples is performed using class-labelled feature-value vectors, noise leads to an imbalanced set of vectors: assuming that the class label takes two values (in this work complement/non-complement), one class (complements) is heavily underrepresented in the data in comparison to the other. To overcome the drop in accuracy when predicting instances of the rare class due to this disproportion, we balance the learning data by applying one-sided sampling to the training corpus and thus by reducing the number of non-complement instances. This approach has been used in the past in several domains (image processing, medicine, etc) but not in natural language processing. For identifying the examples that are safe to remove, we use the value difference metric, which proves to be more suitable for nominal attributes like the ones this work deals with, unlike the Euclidean distance, which has been used traditionally in one-sided sampling. We experiment with different learning algorithms which have been widely used and their performance is well known to the machine learning community: Bayesian learners, instance-based learners and decision trees. Additionally we present and test a variation of Bayesian belief networks, the COr-BBN (Class-oriented Bayesian belief network). The performance improves up to 22% after balancing the dataset, reaching 73.7% f-measure for the complement class, having made use only a phrase chunker and basic morphological information for preprocessing.
We report on a series of experiments with probabilistic context-free grammars predicting English and German syllable structure. The treebank-trained grammars are evaluated on a syllabification task. The grammar used by As she evaluates the grammar only for German, we reimplement the grammar and experiment with additional phonotactic features. Using bi-grams within the syllable, we can model the dependency from the previous consonant in the onset and coda. A 10fold cross validation procedure shows that syllabification can be improved by incorporating this type of phonotactic knowledge. Compared to the grammar of Mller (2002), syllable boundary accuracy increases from 95.8% to 97.2% for English, and from 95.9% to 97.2% for German. Moreover, our experiments with different syllable structures point out that there are dependencies between the onset on the nucleus for German but not for English. The analysis of one of our phonotactic grammars shows that interesting phonotactic constraints are learned. For instance, unvoiced consonants are the most likely first consonants and liquids and glides are preferred as second consonants in two-consonantal onsets.
In the present paper, we examined the effects of autobiographically induced mood and music on emotional evaluations of and psychophysiological responses to music in 48 subjects. Participants listened to music after a mood induction. Both music and induction varied on the dimensions of valence (pleasant - unpleasant) and arousal (high - low). During mood induction and listening to music, psychophysiological responses were measured continuously to assess physiological arousal (indexed by electrodermal activity) and the valence of the emotional state (indexed by facial muscle activity) of the participant. After listening to music, participants evaluated the music using pictorial scales for valence and arousal. As expected, subjects were in a more positive emotional state during listening to pleasant than unpleasant music and also evaluated the music more positively after a pleasant compared to an unpleasant pre-existing mood. As also expected, high-arousal music and pre-existing mood generated both higher physiological arousal and higher arousal ratings compared to low-arousal pre-existing mood and music. We found no support for the principle of mood-congruency, which posits that individuals preferentially process emotional stimuli that are congruent in emotional tone with their current mood state.
We exploit the resources in the Arabic Treebank (ATB) and Arabic Gigaword (AG) to determine the best features for the novel task of automatically creating lexical semantic verb classes for Modern Standard Arabic (MSA). The verbs are classified into groups that share semantic elements of meaning as they exhibit similar syntactic behavior. The results of the clustering experiments are compared with a gold standard set of classes, which is approximated by using the noisy English translations provided in the ATB to create Levin-like classes for MSA. The quality of the clusters is found to be sensitive to the inclusion of syntactic frames, LSA vectors, morphological pattern, and subject animacy. The best set of parameters yields an Fβ=1 score of 0.456, compared to a random baseline of an Fβ=1 score of 0.205.
We have introduced a new methodology that maps designs to human perceptions. Perceptions are adjectives/adverbs, phrases or sentences expressed in natural language. We used the lexical database WordNet to compute semantical relationships (or distances) between these perceptions. We partitioned the set of perceptions into k clusters that represent the classes for a further classification task. We have developed a new classifier called “structural hidden Markov model” (SHMM) that combines probability and distances in a seamless way. SHMM enables to learn and predict user perceptions given object designs. We have applied this approach to Kansei engineering in order to map car external contours (shapes) to customer perceptions. The accuracy obtained using the SHMM is 90%. This model has outperformed the neural network and the k-nearest-neighbor classifiers.
In this paper, we construct a biomedical semantic role labeling (SRL) system that can be used to facilitate relation extraction. First, we construct a proposition bank on top of the popular biomedical GENIA treebank following the PropBank annotation scheme. We only annotate the predicate-argument structures (PAS's) of thirty frequently used biomedical predicates and their corresponding arguments. Second, we use our proposition bank to train a biomedical SRL system, which uses a maximum entropy (ME) model. Thirdly, we automatically generate argument-type templates which can be used to improve classification of biomedical argument types. Our experimental results show that a newswire SRL system that achieves an F-score of 86.29% in the newswire domain can maintain an F-score of 64.64% when ported to the biomedical domain. By using our annotated biomedical corpus, we can increase that F-score by 22.9%. Adding automatically generated template features further increases overall F-score by 0.47% and adjunct arguments (AM) F-score by 1.57%, respectively.
Standard techniques used in multilingual terminology management fail to describe legal terminologies as they are bound to different legal systems and terms do not share a common meaning. In the LexALP project, we use a technique defined for general lexical databases to achieve cross language interoperability between lan-guages of the Alpine Convention. In this paper we present the methodology and tools developed for the collection, de-scription and harmonisation of the legal terminology of spatial planning and sus-tainable development in the four lan-guages of the countries of the Alpine Space. 1
This paper discusses a novel probabilistic synchronous TAG formalism, synchronous Tree Substitution Grammar with sister adjunction (TSG+SA). We use it to parse a language for which there is no training data, by leveraging off a second, related language for which there is abundant training data. The grammar for the resource-rich side is automatically extracted from a treebank; the grammar on the resource-poor side and the synchronization are created by handwritten rules. Our approach thus represents a combination of grammar-based and empirical natural language processing. We discuss the approach using the example of Levantine Arabic and Standard Arabic.
Powerful lookup capabilities are included in almost every electronic dictionary. However, improvements are still conceivable in the exploration of the lexical information within the dictionary and in the integration of the dictionary into other applications. In this paper, we illustrate the most salient functions of the Base lexicale dufrançais (BLF), an online lexical database for general French. The major features of this interactive database are the integration of a dictionary (DAFLES), automatically generated exercises (ALFALEX) and corpus applications (CATS) in order to create a powerful learning environment for learners of French. The DAFLES focuses on the learner's needs by providing separate access paths according to receptive or productive needs. ALFALEX trains the language learner on global lexical knowledge as it is described in the DAFLES. The corpus applications enrich the lexical description by providing the (comparative) combinatorial profiles of words as well as examples and sentences for the exercises.
Standard techniques used in multilingual terminology management fail to describe legal terminologies as they are bound to different legal systems and terms do not share a common meaning. In the LexALP project, we use a technique defined for general lexical databases to achieve cross language interoperability between languages of the Alpine Convention. In this paper we present the methodology and tools developed for the collection, description and harmonisation of the legal terminology of spatial planning and sustainable development in the four languages of the countries of the Alpine Space.
This paper describes a parser which generates parse trees with empty elements in which traces and fillers are co-indexed. The parser is an unlexicalized PCFG parser which is guaranteed to return the most probable parse. The grammar is extracted from a version of the PENN treebank which was automatically annotated with features in the style of The annotation includes GPSG-style slash features which link traces and fillers, and other features which improve the general parsing accuracy. In an evaluation on the PENN treebank Its results for the empty category prediction task and the trace-filler coindexation task exceed all previously reported results with 84.1% and 77.4% fscore, respectively.
Based on the analysis of the usage and the syntactic function of Chinese punctuations,this paper proposes a new hierarchical approach to parse the long Chinese sentences.In traditional parsing approaches,the parsing procedure is performed in an one-level way and the punctuation marks are not specially treated.Correspondingly,in our approach,the complex long Chinese sentences are broken into sub-sentences or units(say 'units' hereafter) by using punctuation marks with special functions,so that the original whole sentence is parsed unit by unit.This idea of 'divide-and-conquer' greatly reduces the difficulty in the traditional parsing approaches to recognize the syntactic relationship between the sub-sentences and phrases or inside the sub-sentences or phrases.And also,in our approach,the grammatical rules with punctuation marks and their probabilities are extracted from the large scale treebank,which are very beneficial to the syntactic disambiguation.Our experimental results have shown that comparing with the traditional Chart parsing algorithm,our approach can significantly reduce the time consumption and the numbers of ambiguous edges,and get about 7% of the correct rate and the recall rate increasing while parsing long Chinese sentences.
We built a morphological analyzer, which can be freely used by anyone for research purpose. In order to build a pratical system, a dictionary with reasonable size is necessary. The initial dictionary is built from the Penn Chinese Treebank corpus v4.0 and contains only 33,438 entries. Since the initial dictionary is quite small, unknown word detection methods are applied to a huge raw text in order to extract new words to be added into the system dictionary. We have successfully constructed a dictionary with 120,769 entries. Finally, we propose a two-layer morphological analyzer to cater for two sets of outputs. The first layer produces the minimal segmentation units defined by us, and the second layer transforms the output of the first layer to the original segmentation units defined by Penn Chinese Treebank.
In this paper, a new approach of using temporal information to assist in Mandarin speech recognition is discussed. It incorporates two types of temporal information into the recognition search. One is a statistical syllable duration model which considers the influences of 411 basesyllables, 5 tones, 4 position-in-word factors, and 3 positionin-sentence factors on syllable duration. Another is the timing information of modeling three types of inter-syllable boundary including intra-word, inter-word without punctuation mark (PM), and inter-word with PM. The uses of these two types of temporal information are expected to be useful for improving the segmentation accuracies in both acoustic decoding and linguistic decoding. Experimental results showed that the base-syllable/character/word recognition rates were slightly improved for both MATBN and Treebank datbase.
Incorporating semantic features from the WordNet lexical database is among one of the many approaches that have been tried to improve the predictive performance of text classifica-tion models. The intuition behind this is that keywords in the training set alone may not be extensive enough to enable generation of a universal model for a category, but if we in-corporate the word relationships in WordNet, a more accu-rate model may be possible. Other researchers have previ-ously evaluated the effectiveness of incorporating WordNet synonyms, hypernyms, and hyponyms into text classification models. Generally, they have found that improvements in accuracy using features derived from these relationships are dependent upon the nature of the text corpora from which the document collections are extracted. In this paper, we not only reconsider the role of WordNet synonyms, hypernyms, and hyponyms in text classification models, we also consider the role of WordNet meronyms and holonyms. Incorporating these WordNet relationships into a Coordinate Matching clas-sifier, a Naive Bayes classifier, and a Support Vector Machine classifier, we evaluate our approach on six document collec-tions extracted from the Reuters-21578, USENET, and Digi-Trad text corpora. Experimental results show that none of the WordNet relationships were effective at increasing the accu-racy of the Naive Bayes classifier. Synonyms, hypernyms, and holonyms were effective at increasing the accuracy of the Coordinate Matching classifier, and hypernyms were effec-tive at increasing the accuracy of the SVM classifier.
After parsing an important task is to determine the semantic structure of sentences. In this paper, we attempt to automatically annotate the Penn Chinese Treebank with semantic dependency structure. Initially a small portion of the Penn Chinese Treebank was manually annotated with headword and dependency relations. Two supervised machine learning algorithms with varying features were then used to learn the relations. Finally, a set of rules were created based on features of Chinese to solve some problem patterns that were found in the Penn Chinese Treebank dealing with ambiguous structures. The experimental results show that the algorithms and proposed approach are effective for determining semantic dependency structure automatically.
The Jibiki platform ¡s an online generic environment for writing and querying all kinds of dictionaries: terminological glossaries, bilingual dictionaries, multilingual lexical databases, etc. It has been developed mainly by Mathieu Mangeot (Université de Savoie, France) and Gilles Sérasset (Université de Grenoble 1, France), thanks to research driven by the GETA team of the CLIPS laboratory in Grenoble, France. The platform allows one to lookup all the dictionaries available on the server and to display the results in the same window. The advanced query interface offers a combination of multiple search criteria. The writing of the entries is done directly online on the platform via a web browser. The writing interface is generated automatically from the description of the structure of the entries (an XML schema), thus allowing the edition of (almost) any type of dictionary entry.
Use of structural information and lexicalization are two of the main challenges facing syntactic analysis, and they are investigated in this paper. First, the probabilities of lexical dependencies are obtained by training a large-scale dependency treebank and used to build the lexical model. Second, the governing degree of words is introduced to utilize the structure information. The lexical method overcomes the weakness of POS dependencies in the past work; meanwhile the governing degree of words is helpful to distinguish the syntactic structures so some ill-formed structures are avoided. Finally, the paper shows a good experimental result of around 74% accuracy on the test set that consists of 4000 sentences.
This paper describes methods and tools used for the post-annotation checking of Prague Dependency Treebank 2.0 data. The annotation process was complicated by many factors: for example, the corpus is divided into several layers that must reflect each other; the annotation rules changed and evolved during the annotation process; some parts of the data were annotated separately and in parallel and had to be merged with the data later. The conversion of the data from the old format to a new one was another source of possible problems besides omnipresent human inadvertence. The checking procedures are classified according to several aspects, e.g. their linguistic relevance and their role in the checking process, and prominent examples are given. In the last part of the paper, the methods are compared and scored.
This paper describes an ongoing effort to parse the Hebrew Bible. The parser consults the bracketing information extracted from the cantillation marks of the Masoetic text. We first constructed a cantillation treebank which encodes the prosodic structures of the text. It was found that many of the prosodic boundaries in the cantillation trees correspond, directly or indirectly, to the phrase boundaries of the syntactic trees we are trying to build. All the useful boundary information was then extracted to help the parser make syntactic decisions, either serving as hard constraints in rule application or used probabilistically in tree ranking. This has greatly improved the accuracy and efficiency of the parser and reduced the amount of manual work in building a Hebrew treebank.
Body image has been shown to be influenced by weight loss. Little attention however has been devoted to personal evaluations of physical fitness (fitness evaluation) or the extent to which individuals psychologically invest in improving fitness (fitness orientation) during periods of weight loss. PURPOSE: The purpose of the present study was to examine the relation between fitness evaluation (FE) and fitness orientation (FO) subscales with other body image ratings and physiological measures of fitness during a twelve-week behavioral weight loss program. METHODS: Thirty overweight, sedentary women (age= 42.5 ± 8.5 years, BMI= 29.3 kg/m2 ± 3.2) participated in a twelve-week behavioral weight loss program which reduced energy intake to 1200–1500 kcal/day and dietary fat to <30% of total calories. Subjects were progressively increased to 40 min/day, 5 days/week of home-based walking exercise. Body image was assessed using Multidimensional Body-Self Relations Questionnaire (MBSRQ) subscales. Cardiorespiratory fitness was measured using time to reach 85% of age-predicted maximal heart rate during a submaximal test on a treadmill. RESULTS: Time to reach 85% of age-predicted maximal heart rate significantly increased over the 12 week intervention (10.3+3.0 min vs. 12.5+3.4 min, p< 0.00). Baseline FE and FO were positively associated with baseline cardiorespiratory fitness (r=0.55 and 0.53, respectively). Change in FO was positively associated with baseline FO, changes in appearance evaluation and orientation and change in cardiorespiratory fitness (r=0.40, 0.65, 0.37 and 0.39, respectively). Furthermore, positive change in FO was associated with 12-week scores of appearance orientation, health orientation and overweight preoccupation (r=0.38, 0.37, and 0.51, respectively). CONCLUSION: These findings demonstrate that individuals with higher scores of FO at baseline had the greatest change in FO at 12 weeks, which was positively associated with changes in other perceived personal fitness constructs and fitness improvements. Exercise interventions should target strategies for improving FO and FE as this may have implications for exercise adoption and maintenance during weight loss. Supported by a Research Incentive Grant from the University of Louisville
This paper describes the Chinese NomBank Project, the goal of which is to annotate the predicate-argument structure of nominalized predicates in Chinese. The Chinese Nombank extends the general framework of the English and Chinese Proposition Banks to the annotation of nominalized predicates and adds a layer of semantic annotation to the Chinese Treebank. We first outline the scope of the work by discussing the markability of the nominalized predicates and their arguments. We then attempt to provide a categorization of the distribution of the arguments of nominalized predicates. We also discuss the relevance of the event/result distinction to the annotation of nominalized predicates and the phenomenon of incorporation. Finally we discuss some cross-linguistic differences between English and Chinese. 1.
Summary form only given. We present first some general remarks on challenges faced by modern information technology, notably when a human being is a relevant factor. These challenges are mainly related to inherent difficulties in solving some "meta-problems", in particular broadly perceived decision making. We assume, on the one hand, business intelligence related perspective, augmented with elements of Web intelligence, to fully use all available tools and resources. On the other hand, we assume a human centric computing perspective in the spirit of, for instance, Dertouzos's ideas. First, we present a brief account of modem approaches to real world decision making, emphasize the concept of a decision making process that involves more factors and aspects like: the use of own and external knowledge, involvement of various,actors, aspects, etc., individual habitual domains, non-trivial rationality, different paradigms. As an example we mention Checkland's deliberative decision making (which is an important elements of his soft approach to systems analysis). After an analysis of specifics and difficulties encountered in many real world decision-making situations, we strongly advocate the use of computer based decision support systems. First, we briefly review the history of decision support systems, and then present a popular classification, starting from data driven to Web based and inter-organizational. We indicate that decision support systems should incorporate some sort of "intelligence", and we first briefly mention some views of what intelligence may mean in this concept, and then assume some more pragmatic, though limited, view of intelligent decision support systems. We indicate possible advantages of using elements of fuzzy logic and soft computing, notably, Zadeh's computing with words to be able to somehow merge the ideas presented like: human centric computing, decision making processes, intelligent decision support, etc. Finally, we present an example of implementation in which the above-mentioned ideas have been to some extent implemented. This concern a data and document driven decision support system for a small to medium company in which, first, Zadeh's computing with words and perceptions paradigm is employed via linguistic database summaries, elements of Web intelligence are used to derive additional information, and the ideas of an intelligent decision support and human centric computing are shown to be synergistically combined. We finish with some general remarks emphasizing that fuzzy logic and soft computing, notably as exemplified by Zadeh's computing with words and perceptions may be viewed as providing just the right tools to solve the problems considered