Papers reviewed and determined not to be word norm studies. Use the flag icon to report errors or suggest re-inclusion.
18265 papers
The query language in TIGERSearch is limited due to its lack of universal quantification. This restriction makes it impossible to ask simple queries like „Find sentences that do not include a certain word”. We propose an easy way to formulate such queries. We have implemented this extension to the query language in a tool that allows querying parallel treebanks including their alignment constraints. Our implementation of universal quantification relies on the view of node sets rather than single node unification. Our query tool is freely available.
Recent parsing research has started addressing the questions a) how parsers trained on different syntactic resources differ in their performance and b) how to conduct a meaningful evaluation of the parsing results across such a range of syntactic representations. Two German treebanks, Negra and TüBa-D/Z, constitute an interesting testing ground for such research given that the two treebanks make very different representational choices for this language, which also is of general interest given that German is situated between the extremes of fixed and free word order. We show that previous work comparing PCFG parsing with these two treebanks employed PARSEVAL and grammatical function comparisons which were skewed by differences between the two corpus annotation schemes. Focusing on the grammatical dependency triples as an essential dimension of comparison, we show that the two very distinct corpora result in comparable parsing performance.
The paper analyzes various possible linguistic norms that could govern the feminine forms, which slowly appear in the Polish language, and which correspond to the masculine names of professions. Adopting a basically feminist standpoint leads one to reject those proposals, which would legislate that the masculine forms ought to be applied to men while the feminine forms ought to be applied to women. The article considers in particular the inferential roles of concepts to argue for a gender-neutral rendition of the historically masculine forms.
In this paper, we give a description of the machine translation (MT) system developed at DCU that was used for our third participation in the evaluation campaign of the International Workshop on Spoken Language Translation (IWSLT 2008). In this participation, we focus on various techniques for word and phrase alignment to improve system quality. Specifically, we try out our word packing and syntax-enhanced word alignment techniques for the Chinese–English task and for the English–Chinese task for the first time. For all translation tasks except Arabic–English, we exploit linguistically motivated bilingual phrase pairs extracted from parallel treebanks. We smooth our translation tables with out-of-domain word translations for the Arabic–English and Chinese–English tasks in order to solve the problem of the high number of out of vocabulary items. We also carried out experiments combining both in-domain and out-of-domain data to improve system performance and, finally, we deploy a majority voting procedure combining a language model based method and a translation-based method for case and punctuation restoration. We participated in all the translation \ntasks and translated both the single-best ASR hypotheses and \nthe correct recognition results. The translation results confirm that our new word and phrase alignment techniques are often helpful in improving translation quality, and the data combination method we proposed can significantly improve system performance.
In Brief Objectives: Accurate identification of environmental sounds plays an important role in maintaining listeners’ awareness of their environment, and is a major concern for cochlear implant patients. Although research indicates that decreased spectral resolution has a negative effect on environmental sound identification, little is known about the processes underlying perceptual adaptation to spectrally-degraded input. The goals of this study were (1) to develop a test of environmental sound perception containing a large variety of easily identifiable and familiar sound sources, represented by multiple exemplars, and (2) to examine whether auditory training improves listeners’ identification of spectrally-degraded environmental sounds. Design: In experiment 1, familiarity ratings and identification accuracy were obtained for 21 normal-hearing subjects for 48 environmental sound sources; there were 4 exemplars of each sound source, for a total of 192 stimuli. A second test was developed using a subset of 40 sound sources (4 exemplars each, for a total of 160 stimuli). In experiment 2, seven normal-hearing subjects (who did not participate in experiment 1) were asked to identify spectrally-degraded environmental sounds processed by a four-channel noise-band vocoder. The second stimulus set developed in experiment 1 (40 sound sources, 4 exemplars each) was used in experiment 2. The subjects were tested in a pretest–posttest design with five training sessions between the pretest and the posttest. The training sounds were selected individually for each subject, and comprised one half of the sound sources that were misidentified in the pretest. Each sound source used in training was represented by two exemplars. During training, subjects received trial and block feedback. For each incorrect response, subjects were allowed to replay the stimulus up to five times after being shown the correct response. Results: In experiment 1, listeners’ average identification accuracy was 95% correct, with 178 of all sounds identified with an accuracy of 80% or more. The average identification accuracy of the 160 sounds selected for experiment 2 was 98% correct, and their average familiarity rating was 6.39 (on a 7-point scale). In experiment 2, the average identification accuracy of spectrally-degraded sounds was 33% correct on the pretest. However, after training, average identification accuracy across all sounds improved to 63% correct on the posttest. The largest improvement (86 percentage points) was obtained for the sound exemplars used during training. The identification accuracy for alternative exemplars of the training sounds (that referenced the same sources) improved by 36 percentage points. Finally, the identification of sound sources not included in the training, but perceived with equal difficulty on the pretest, improved by 18 percentage points. Conclusions: These results demonstrate positive effects of training on the identification of spectrally-degraded environmental sounds and suggest that training effects can generalize to other sound exemplars and sources, although with a reduced magnitude of improvement. The findings also indicate a timeline for initial perceptual adaptation to spectrally-degraded environmental sounds, and provide a preliminary basis for incorporating environmental sounds into auditory rehabilitation programs for cochlear implant patients. Environmental sound perception is an important concern for cochlear implant patients. In this study, a large 160-item test of environmental sound perception was developed and used to examine whether auditory training improves listeners’ identification of environmental sounds processed by an acoustically simulated cochlear implant. Seven normal-hearing listeners identified spectrally-degraded stimuli obtained with a four-channel noise-based vocoder before and after five training sessions. Identification performance improved after training, mostly for the sounds included in the training set, but also, to a lesser extent, for untrained sounds. Results provide preliminary basis for incorporating environmental sounds into cochlear implant rehabilitation programs.
This paper introduces the infrastructure and the principles of a semantic framework used for the analysis and classification of verbs, developed with the aim of constructing a lexical database of Mandarin verbal semantics, called the Mandarin VerbNet. Distinct from most existing lexical databases that enumerate word senses without detailed grammatical considerations, the Mandarin VerbNet is designed to provide lexical semantic information based on grammatical descriptions and anchored in linguistic theories. It looks for systematic correlations between syntax and semantics and classifies verbs according to these syntax-to-semantics correspondences. The framework adopts the approach of Frame Semantics (Fillmore & Atkins 1992) in defining verb meanings in a semantic frame and building a frame-based verbal lexicon, but it differs from the structure of the English FrameNet in distinguishing different scopes of frames. Evolved and refined from previous works (Liu, Chiang & Chang 2004, Liu & Wu 2003, Liu 2002), this study summarizes the current model of the analytic framework with a detailed illustration from Mandarin statement verbs. It ultimately seeks to identify a theoretically sound and operationally effective representational scheme that bases its semantic analysis on grammatical behaviors and provides linguistic motivations for its semantic classifications.
Robust spoken language understanding (SLU) is a key component of spoken dialogue systems. Recent statistical approaches to this problem require additional resources (e.g. gazetteers, grammars, syntactic treebanks) which are expensive and time-consuming to produce and maintain. However, simple datasets annotated only with slot-values are commonly used in dialogue systems development, and are easy to collect, automatically annotate, and update. We show that it is possible to reach state-of-the-art performance using minimal additional resources, by using Markov logic networks (MLNs). We also show that performance can be further improved by exploiting long distance dependencies between slot-values. For example, by representing such features in MLNs, but without using a gazetteer, we outperform the hidden vector state (HVS) model of He and Young 2006 (1.26% improvement, a 13% error reduction).
Creation of material, technical and personal prerequisites for lexicographic work by using modern technologies; information on lexicographical bibliographical database.
The effects of emotional content and emotional voice on speech intelligibility in younger and older adults was investigated. Twenty-eight younger adults with good health and clinically normal hearing thresholds in the speech range were tested. The stimuli used were the 200 sentences from the NU6 lists. The stimuli were presented to one group visually as text on paper, and to two groups auditorally, through two loudspeakers in a sound-attenuating booth. Means were obtained for both valence and arousal ratings for all three groups. The SNR threshold data were collected on young adults with normal hearing by Richard Wilson and colleagues using the female voice. Analyses revealed a significant positive correlation between valence and arousal for participants in the visual condition. The results indicated that the emotional arousal of listeners to a particular word can affect intelligibility, depending on the modality of presentation.
Automatic image annotation is very important for image retrieval. Despite continuous efforts in inventing new annotation algorithms, the annotation performance is usually unsatisfactory, and the annotation vocabulary is still limited due to the use of a small scale training set. In this paper, a novel image automatic annotation system based on the WordNet is presented, named WordNet-based image annotation. By using WordNet hierarchical structure, we collect a large image datasets. And each image is loosely labeled with one of the non-abstract nouns in English, as listed in the WordNet lexical database. Then we use PageRank method to delete the wrong images under every word, and make sure that every word covers 100 images. Hence the image database gives a comprehensive coverage of all object categories and scenes. The semantic information from WordNet can be used in conjunction with SVM classifiers to perform object classification over a range of semantic levels minimizing the effects of labeling noise. The system models a real-world situation by including pictures gathered from the Internet and is designed for exploratory large scale image retrieval system based on the internet.
This paper describes the use of two machine learning techniques, naive Bayes and decision trees, to address the task of assigning function tags to nodes in a syntactic parse tree. Function tags are extra functional information, such as logical subject or predicate, that can be added to certain nodes in syntactic parse trees. We model the function tags assignment problem as a classification problem. Each function tag is regarded as a class and the task is to find what class/tag a given node in a parse tree belongs to from a set of predefined classes/tags. The paper offers the first systematic comparison of the two techniques, naive Bayes and decision trees, for the task of function tags assignment. The comparison is based on a standardized data set, the Penn Treebank, a collection of sentences annotated with syntactic information including function tags. We found out that decision trees generally outperform naive Bayes for the task of function tagging. Furthermore, this is the first large scale evaluation of decision trees based solutions to the task of functional tagging.
In this paper, we present the Web-based resource sharing system KnownStyleNoLife, which allows users explicitly annotate fashion-related images. KnownStyleNoLife harnesses human power and collects image metadata such as object locations, object labels, image rating and semantic relationships among images. Metadata of that type are invaluable as it is difficult for ordinary computer software to extract equivalent metadata from the Web. Acquiring image metadata such as the type and location of objects in images requires a computer to use machine learning techniques, needing training with sample images and Web search techniques, with the outcome that only limited image metadata are extracted. Furthermore, the semantic image relationship metadata is based upon peoplepsilas image perception, hence user feedback is required to acquire it. KnownStyleNoLife provides a social value to users in return for their explicit annotations. It is a novel approach to harness human power to acquire relevant image metadata.
To determine how differences in emotion representation and/or inhibitory ability affect adolescents’ responses to emotion words, 13-yr and 16-yr olds, as well as adults, were compared on the processing of emotion-laden and neutral words. Word ratings revealed that 16-yr olds tended towards perceiving all words as more arousing than did adults, irrespective of valence. Also, they rated words more negatively than 13-yr olds. Performance on an Affective Simon task revealed a marked incongruency effect only for 13-yr olds (and then only for negative words) but not for 16-yr olds (who responded fastest) or adults. Performance on a sustained attention task confirmed the expected age-related increase in inhibitory ability and a concomitant increase in response latencies. Our conclusions are two-fold. First, there are age-related differences in lexical representation which appear more marked for 16-yr olds. Second, 16-yr olds are more reactive, irrespective of the emotional content they are processing, yet appear to control its impact as efficiently as adults.
In this paper we present LXGram, a general purpose grammar for the deep linguistic processing of Portuguese that aims at delivering detailed and high precision meaning representations. LXGram is grounded on the linguistic framework of Head-Driven Phrase Structure Grammar (HPSG). HPSG is a declarative formalism resorting to unification and a type system with multiple inheritance. The semantic representations that LXGram associates with linguistic expressions use the Minimal Recursion Semantics (MRS) format, which allows for the underspecification of scope effects. LXGram is developed in the Linguistic Knowledge Builder (LKB) system, a grammar development environment that provides debugging tools and efficient algorithms for parsing and generation. The implementation of LXGram has focused on the structure of Noun Phrases, and LXGram accounts for many NP related phenomena. Its coverage continues to be increased with new phenomena, and there is active work on extending the grammar's lexicon. We have already integrated, or plan to integrate, LXGram in a few applications, namely paraphrasing, treebanking and language variant detection. Grammar coverage has been tested on newspaper text.
This paper proposes an novel approach to annotate function tags for unparsed text. What distinguishes our work from other attempts in such task is that we assign function tags directly basing on lexical information other than on parsed trees. In order to demonstrate the effectiveness and versatility of our method, we investigate two statistical models for automatic annotation, one is log-linear maximum entropy model and the other is margin maximum based support vector machine model, which achieve the best F-score of 82.8 and 86.4 respectively when tested on the text from Penn Chinese Treebank. We also quantity the effect of POS tagger accuracy on system performance. Our results indicate that the function tag types could be determined via flexible and powerful feature representations from words, POS tags and word position indicators, and that, similarly to syntactic parsing, the main difficulty lies in complex constituents with long-distance dependency.
Coordinations in noun phrases often pose the problem that elliptified parts have to be reconstructed for proper semantic interpretation. Unfortunately, the detection of coordinated heads and identification of elliptified elements notoriously lead to ambiguous reconstruction alternatives. While linguistic intuition suggests that semantic criteria might play an important, if not superior, role in disambiguating resolution alternatives, our experiments on the reannotated WSJ part of the Penn Treebank indicate that solely morpho-syntactic criteria are more predictive than solely lexico-semantic ones. We also found that the combination of both criteria does not yield any substantial improvement.
What kinds of lexical resources are helpful for extracting useful information from domain-specific documents? Although domain-specific documents contain much useful knowledge, it is not obvious how to extract such knowledge efficiently from the documents. We need to develop techniques for extracting hidden information from such domain-specific documents. These techniques do not necessarily use state-of-the-art technologies and achieve deep and accurate language understanding, but are based on huge amounts of linguistic resources, such as domain-specific lexical databases. In this paper, we introduce two techniques for extracting informative expressions from documents: the extraction of related words that are not only taxonomically related but also thematically related, and the acquisition of salient terms and phrases. With these techniques we then attempt to automatically and statistically extract domain-specific informative expressions in aviation documents as an example and evaluate the results. 1.
Parsing is important in Linguistics and Natural Language Processing to understand the syntax and semantics of a natural language grammar. Parsing natural language text is challenging because of the problems like ambiguity and inefficiency. Also the interpretation of natural language text depends on context based techniques. A probabilistic component is essential to resolve ambiguity in both syntax and semantics thereby increasing accuracy and efficiency of the parser. Tamil language has some inherent features which are more challenging. In order to obtain the solutions, lexicalized and statistical approach is to be applied in the parsing with the aid of a language model. Statistical models mainly focus on semantics of the language which are suitable for large vocabulary tasks where as structural methods focus on syntax which models small vocabulary tasks. A statistical language model based on Trigram for Tamil language with medium vocabulary of 5000 words has been built. Though statistical parsing gives better performance through tri-gram probabilities and large vocabulary size, it has some disadvantages like focus on semantics rather than syntax, lack of support in free ordering of words and long term relationship. To overcome the disadvantages a structural component is to be incorporated in statistical language models which leads to the implementation of hybrid language models. This paper has attempted to build phrase structured hybrid language model which resolves above mentioned disadvantages. In the development of hybrid language model, new part of speech tag set for Tamil language has been developed with more than 500 tags which have the wider coverage. A phrase structured Treebank has been developed with 326 Tamil sentences which covers more than 5000 words. A hybrid language model has been trained with the phrase structured Treebank using immediate head parsing technique. Lexicalized and statistical parser which employs this hybrid language model and immediate head parsing technique gives better results than pure grammar and trigram based model.
International audience
This paper comments on the speech reconstruction problem and types of disfluencies which are necessary to eliminate in order to get a grammatically and semantically correct text. It summarizes methods dealing with this problem or its specific parts and introduces the current work on Prague Dependency Treebank of Spoken Czech, which is supposed to provide data for future research.
A Treebank is a text corpus in which each sentence has been annotated with its syntactic structure. Although the construction of a treebank is an expensive task, we believe that it is indispensable for the development of real applications in the field of Natural Language Processing (NLP) and also for the development of the Information Society. At a purely linguistic level, the Treebank is an essential database for the study of a language given that it provides analyzed/annotated examples of real language. The linguistic study directly results in an improvement in the quality of several applications, such as Part-Of-Speech (POS) taggers and parsers (Collins 1997, 2000; Charniak 2000), because it provides common training and testing material allowing different algorithms to be compared and improved. In the last few years, treebank corpora such as the Penn Treebank (Marcus et al., 1993) and the Prague Dependency Treebank (Bohmova et al. 2003) have become a crucial resource for building and evaluating natural language processing tools and applications. As Abeille (2003) sets out, there are efforts underway for Czech, German, French, Japanese, Polish, Spanish and Turkish, to name just a few. In Kakkonen (2005) we can find the state of the art of dependency-based treebanks. The Basque Dependency Treebank (BDT) is actually the Reference Corpus for the Processing of Basque (EPEC) annotated at syntactic level. The EPEC is a 300,000 word corpus of standard written texts which aims to be a training corpus for the development and improvement of several NLP tools. It has been manually tagged at different levels: morphology, lemmatization and surface syntax (Aduriz et al. 2006). The next level of tagging —annotation of dependency relations— is currently being carried out in BDT. In this paper, we describe the annotation of noun phrase (henceforth, NP) constructions in detail following the Dependency Grammar theory (Tesniere 1959). For a better understanding of our work it should be noted that for us, NP is a purely descriptive term. We are not concerned with understanding the internal structure of NPs. The syntactic description of Basque NPs has been mainly developed within the generative framework by Goenaga (1980), Eguzkitza (1993), Laka (1993), Artiagoitia (2002),
Recently the LATL has undertaken the development of a multilingual translation system based on a symbolic parsing technology and on a transfer-based translation model. A crucial component of the system is the lexical database, notably the bilingual dictionaries containing the information for the lexical transfer from one language to another. As the number of necessary bilingual dictionaries is a quadratic function of the number of languages considered, we will face the problem of getting a large number of dictionaries. In this paper we discuss a solution to derive a bilingual dictionary by transitivity using existing ones and to check the generated translations in a parallel corpus. Our first experiments concerns the generation of two bilingual dictionaries and the quality of the entries are very promising. The number of generated entries could however be improved and we conclude the paper with the possible ways we plan to explore. 1.
This paper presents a corpus study of parenthetical constructions in two different corpora: the Penn Discourse Treebank (PDTB, (PDTB-Group, 2008)) and the RST Discourse Treebank (Carlson et al., 2001). The motivation for the study is to gain a better understanding of the rhetorical properties of parentheticals in order to enable a natural language generation system to produce parentheticals as part of a rhetorically well-formed output. We argue that there is a correlation between syntactic and rhetorical types of parentheticals and establish two main categories: ELABORATION/EXPANSION-type NP-modifier parentheticals and NON-ELABORATION/EXPANSION-type VP- or S-modifier parentheticals. We show several strategies for extracting these from the two corpora and discuss how the seemingly contradictory results obtained can be reconciled in light of the rhetorical and syntactic properties of parentheticals as well as the decisions taken in the annotation guidelines. 1. Definition Parentheticals are constructions that typically occur embed-ded in the middle of a clause. They are not part of the main predicate-argument structure of the sentence and are marked by special punctuation (e.g. parentheses, dashes, commas) in written texts, or by special intonation in speech.
High doses of ibuprofen have been shown to inhibit muscle protein synthesis after a bout of resistance exercise. We determined the effect of a moderate dose of ibuprofen (400 mg x d(-1)) consumed on a daily basis after resistance training on muscle hypertrophy and strength. Twelve males and 6 females (approximately 24 years of age) trained their right and left biceps on alternate days (6 sets of 4-10 repetitions), 5 d x week(-1), for 6 weeks. In a counter-balanced, double-blind design, they were randomized to receive 400 mg x d(-1) ibuprofen immediately after training their left or right arm, and a placebo after training the opposite arm the following day. Before- and after-training muscle thickness of both biceps was measured using ultrasound and 1 repetition maximum (1 RM) arm curl strength was determined on both arms. Subjects rated their muscle soreness daily. There were time main effects for muscle thickness and strength (p < 0.01). Ibuprofen consumption had no effect on muscle hypertrophy (muscle thickness of biceps for arm receiving ibuprofen: pre 3.63 +/- 0.14, post 3.92 +/- 0.15 cm; and placebo: pre 3.62 +/- 0.15, post 3.90 +/- 0.15 cm) and strength (1 RM of arm receiving ibuprofen: pre 18.6 +/- 2.8, post 23.4 +/- 3.5 kg; and placebo: pre 18.8 +/- 2.8, post 22.8 +/- 3.4 kg). Muscle soreness was elevated during the first week of training only, but was not different between the ibuprofen and placebo arm. We conclude that a moderate dose of ibuprofen ingested after repeated resistance training sessions does not impair muscle hypertrophy or strength and does not affect ratings of muscle soreness.
Abstract A manufacturer of a product that transports, processes, and packages bulk materials with a pneumatic process sued a competing manufacturer that uses a screw process using the latter company's advertising, which compared and evaluated the two methods, and charged that these advertisements constituted a deceptive trade practice. The plaintiff claimed that in these advertisements the defendant not only made false, misleading, and disparaging comments but also failed to reveal the industry data, studies, statistics, and other information that might substantiate its claims. Syntax analysis of these advertisements revealed that in these advertisements the verb tenses indicated that the defendant did not claim that comparisons with other types of conveyors were based on studies or tests. Semantic analysis of the word “ratings” conveys that this word indicates a subjective estimate or comparison, one not requiring research or tests. It also showed that the terms used in the comparisons (best, good, fair, poor, worst) are used regularly to indicate attitudes, beliefs, or dislikes, as opposed to the numerical, statistical measures of qualities that are used in reporting research findings.
Abstract In this paper, I plan to tackle the complex problem of the use and exploitation of bilingual lexical resources available in machine-readable form. The reusability of lexical resources has indeed attracted a lot of attention in the past few years but NLP researchers have tended to concentrate mainly on monolingual English learners’ dictionaries, somewhat neglecting bilingual dictionaries. The less structured format of the magnetic tapes of the latter is probably partly responsible for this lack of interest. The reluctance of publishers to distribute the machine-readable versions of their bilingual dictionaries has also contributed to the concentration of efforts on the exploitation of monolingual (and predominantly English) machinereadable dictionaries (MRDs). However, it has to be admitted, as Atkins and Levin (1991: 255) point out, that “the explicit treatment accorded to restrictions on subjects/objects of verbs in dictionaries for the foreign learner (i.e. bilingual dictionaries) renders such works a valuable source of material for the semi-automatic construction of a lexical database”. In this paper, I wish to describe the construction of a lexical-semantic database from the computerized version of the Collins–Robert English–French dictionary (Atkins and Duval 1978). I will pay special attention to the retrieval programs which make it possible to readily extract the collocational and thesauric information it contains. The emphasis will also be laid on the manual lexicographical work which was required in order to enrich the database with lexical-semantic information based on Mel’čuk’s descriptive apparatus of lexical functions (Mel’čuk et al. 1984/1988/1992). Attention will also be paid to potential exploitations of this bilingual database, ranging from applications in a word sense disambiguation or translation selection perspective to the integration of the lexical database (LDB) into a series of corpus-based tools for collocation extraction, using the dictionary and its lexical-semantic information as an external resource linked to corpus query tools.
ynaga at tkl.iis.u-tokyo.ac.jp This paper proposes a method of con-structing an accurate probabilistic subcat-egorization (SCF) lexicon for a lexicalized grammar extracted from a treebank. We employ a latent variable model to smooth co-occurrence probabilities between verbs and SCF types in the extracted lexicalized grammar. We applied our method to a verb SCF lexicon of an HPSG grammar acquired from the Penn Treebank. Experimental re-sults show that probabilistic SCF lexicons obtained by our model achieved a lower test-set perplexity against ones obtained by a naive smoothing model using twice as large training data. 1
The aim of this presentation is to describe our on-going work on the building of a written Hindi treebank, referred to here as the Uppsala Hindi
It is now generally accepted that phraseological units are an area of great difficulty in foreign language acquisition, even for advanced learners (Howarth, 1998; Schmitt & Carter, 2004; Nesselhauf, 2005; Forsberg, 2006). Accordingly, we analyzed 906 verb-noun combinations – including 703 phraseological units – with two high-frequency verbs, namely (ex.: prendre position) and (ex.: donner a Det occasion de N), taken from a corpus of academic texts produced by advanced learners of French as a foreign language (L2), and from a control corpus of academic texts produced by native speakers of French (L1). The aim of the study was mainly descriptive: we tried to reach a definition of the phraseological competence of the advanced variety of academic French as an L2, and to compare two different groups of advanced learners, i.e. English-speaking learners (185.000 words) and Dutch-speaking learners (90.000 words). First, to establish the phraseological or free status of these 906 verb-noun combinations, we made use of four lexical databases that were available for French as an L1: a lexical database studying predicates and arguments in the way they co-occur in journalistic texts (Fabre & Bourigault, 2006), a list of verbal phraseological units (M. Gross, 1975, 1988), a corpus of written press, and a corpus of spoken French (Francard et al., 2002). The method of analysis that we adopted in this study combines two complementary approaches in the field of phraseology: (i) the functional approach, which depends on linguistic criteria of fixedness (syntax, semantics, lexis); (ii) the statistical approach, which depends on measures of lexical attraction in corpus linguistics (e.a. the mutual information score). Secondly, this analysis allowed us to quantify several external (mother tongue, text type, spoken/written language), and internal factors (structure of the noun phrase, semantic categories) which are likely to influence the acquisition of (erroneous or error-free) phraseological units with high-frequency verbs. Overall results in academic texts showed a significant contrast in the frequency of verb-noun phraseological units with and donner: advanced learners of French as an L2 underuse phraseological units with 'prendre' (ex.: prendre Det exemple de N) but overuse phraseological units with (ex. donner Det instruction). In addition to this collocational analysis, an error analysis on phraseological units allowed us to categorize errors according to: (i) the error type, which can be defined as a formal (ex.: donner l’incentif), contextual (ex.: prendre instead of avoir lieu) or quantitative type of deviation (ex.: donner Det trait a X instead of preter Det trait a X); (ii) the grammatical category, which concerns the place of the deviation (verb + noun, verb, noun, determiner, and modifier); (iii) the degree of error gravity, which was defined here according to the error type, the grammatical category of error, the number of deviations per phraseological unit, the frequency and the collocability of the phraseological unit concerned in L1 French. Thirdly, we investigated 503 phraseological units taken from strictly argumentative essays produced by learners of French as an L2, from the angle of accuracy and complexity, in order to shed some light on the phraseological proficiency level of advanced English-speaking vs Dutch-speaking learners of French. This thesis is rounded off by a conclusion in which a few explanatory and didactic perspectives are suggested, in keeping with the findings of this essentially descriptive study.
Functional Arabic Morphology is a formulation of the Arabic inflectional system seeking the working interface between morphology and syntax. ElixirFM is its high-level implementation that reuses and extends the Functional Morphology library for Haskell. Inflection and derivation are modeled in terms of paradigms, grammatical categories, lexemes and word classes. The computation of analysis or generation is conceptually distinguished from the general-purpose linguistic model. The lexicon of ElixirFM is designed with respect to abstraction, yet is no more complicated than printed dictionaries. It is derived from the open-source Buckwalter lexicon and is enhanced with information sourcing from the syntactic annotations of the Prague Arabic Dependency Treebank. MorphoTrees is the idea of building effective and intuitive hierarchies over the information provided by computational morphological systems. MorphoTrees are implemented for Arabic as an extension to the TrEd annotation environment based on Perl. Encode Arabic libraries for Haskell and Perl serve for processing the non-trivial and multi-purpose ArabTEX notation that encodes Arabic orthographies and phonetic transcriptions in parallel.
Considering the popularity of the Internet, an automatic interactive feedback system for Elearning websites is becoming increasingly desirable. However, computers still have problems understanding natural languages, especially the Chinese language, firstly because the Chinese language has no space to segment lexical entries (its segmentation method is more difficult than that of English) and secondly because of the lack of a complete grammar in the Chinese language, making parsing more difficult and complicated. Building an automated Chinese feedback system for special application domains could solve these problems. This paper proposes an interactive feedback mechanism in a virtual campus that can parse, understand and respond to Chinese sentences. This mechanism utilizes a specific lexical database according to the particular application. In this way, a virtual campus website can implement a special application domain that chooses the proper response in a user friendly, accurate and timely manner.
Traditionally, parsers are evaluated against gold standard test data. This can cause problems if there is a mismatch between the data structures and representations used by the parser and the gold standard. A particular case in point is German, for which two treebanks (TiGer and TüBa-D/Z) are available with highly different
In this paper we describe some technical and theoretical aspects related to a manually aligned bilingual treebank Italian (ITA) – Italian Sign Language (LIS) provided with both constituency and dependency annotation (Siena University Treebank, SUT). We briefly discuss the linguistic rationale behind the feature set and the dependency/constituency structure we adopted. Moreover we discuss the tool we used to annotate, semi-automatically, the treebank that, in the end, will be evaluated qualitatively with respect to a specific Transfer-Based Machine Translation (TB-MT) task.
878 Reviews focus the subject as clearly as possible, the results are unpredictable and heterogene ity isunavoidable. That notwithstanding, thevolume fulfils its simply stated aims of seeing theproblems and possible solutions better incontext. KING'S COLLEGE LONDON D. N. YEANDLE Sprachnormenwandel imgeschriebenenDeutsch an der Schwelle zum 2I. Jahrhundert. By VIT DOVALIL. (Duisburger Arbeiten zur Sprach- und Kulturwissenschaft, 63) Frankfurt a.M.: Peter Lang. 2006. vi+236 pp.?42.50. ISBN 978-3-631 53425-0. The questions surrounding linguistic norms, standardization, and language change have recently become a major focus of research interest, particularly in respect of German, and thiscontribution by a young Czech scholar is a very timelycontribution to thedebates on these issues.Other studies have tended toconcentrate on the spoken language, examining the degree towhich spoken norms may deviate from standard codifications, but thiswork centres on thewritten language, specifically the language of the supraregional press inGermany, and examines the extent towhich thenorms commonly taken to be standard are actually adhered to in practice and accepted by those usually regarded as language authorities. After a short account of the aims of thework in the firstchapter, the author presents inChapter 2 an extensive attempt to define thebasic termswhich are relevant tohis study, i.e. 'norms', 'language variety', 'standard language', 'nonstandard/substandard language', and ' Umgangssprache'. For each of these he takes thedefinitions given in earlier work, summarizes and classifies them,and tries toarrive at a defensible definitionwhich will serve as a basis forhis own study.This isa useful presentation in itsown rightgiven the range ofoften very varied and sometimes contradictory definitions which have been attempted in the literature over the past century (forty-nine are listed here for 'norms'), and if the author does not always finda sureway through theseminefields (very fewwill be capable of this), theattempt is immenselyworthwhile and provides a helpful set of references for those wishing to inform themselves of thedebates surrounding these terms. The main body of the book (Chapters 3 to 7) contains the empirical investigation and a thorough analysis of the findings.The author's method was firstto select ten variable features (e.g. theuse of preposition plus pronoun rather than thepronominal adverb-i.e. zu was? or wozu?-or the use of the genitive or dative case with statt, wdhrend, and wegen), and establish whether the formsconventionally regarded as non standard are used in thenational press. Using principally the corpora of the Institut fur Deutsche Sprache in Mannheim, he finds this is so forall thevariables he selected, with, forexample over a hundred attestations for theuse of preposition plus pronoun where thenorm requires thepronominal adverb. The sheer volume of thedata found and presented here is quite remarkable. In a second stage he conducted an investi gation bymeans of a questionnaire to fifty-three professors ofGerman linguistics in Germany, asking themwhether they considered the specified forms to be standard or non-standard, and acceptable inwriting. The result of this survey is ifanything even more astonishing, in that a significant number of these variants (notably, by a huge majority, the past subjunctive form brduchte, the past participle gewunken, or the use of brauchenwithout a following zu) were considered to be standard German and wholly acceptable by thisgroup of informants. This latter result isparticularly interesting because itparallels findings fromother recent researchwhich has shown thatGerman schoolteachers and German Lektoren in theUK may be quite uncertain of theprescriptive codification incases where usage varies. However, rather than seeing thisas 'destandardization', as theauthor is inclined MLR, I03.3, 20o8 879 to, it isperhaps the case that even inGerman, where formallyprescribed norms have longbeen unchallenged and therehas been a somewhat defensive attitude towards the codified standard, usage norms are beginning to establish themselves as acceptable competitors. But whatever the interpretation which one might wish to put on the development, the author deserves much credit for such enlightening documentation ofwhat iscurrently taking place inGerman. UNIVERSITY OF MANCHESTER MARTIN DURRELL VonMythen undMdren: Mittelalterliche Kulturgeschichte imSpiegel einerWissen schaftler-Biographie. Festschriftfiir Otfrid Ehrismann zum 65. Geburtstag. Ed. by GUDRUN MARCI-BOEHNCKE and JORGRIECKE. Hildesheim: OIms. 2oo6. 68i pp.?84. ISBN 978-3-487-I3179-5 It isa lovely custom togive senior academics a special birthday present in the formof a book towhich colleagues and formerstudents have contributed...
Treebanks have become crucial for the development of data-driven approaches to natural language processing, human language technologies, grammar extraction, and linguistic research in general. Manifold projects aim at compiling representative treebanks for specific languages. Other projects focus on the development of tools for exploration of annotated treebanks, or explore annotation beyond syntactic structure and beyond single languages. The Seventh International Workshop on Treebanks and Linguistic Theories (TLT7) provides a forum for researchers in the field of Computational Linguistics who are experts in the design, creation and exploitation of treebanks and their relation to linguistic theories. A selection of 16 workshop papers is published in these proceedings. Together, they cover a wide range of topics, including the building, querying, exploring, exploiting and evaluating of treebanks.
Abstract. This paper describes the implementation and system details of Klex, a finite-state transducer lexicon for the Korean language, developed using XRCE’s Xerox Finite State Tool (XFST). Klex is essentially a transducer network representing the lexicon of the Korean language with the lexical string on the upper side and the inflected surface string on the lower side. Two major applications for Klex are morphological analysis and generation: given a well-formed inflected lower string, a languageindependent algorithm derives the upper lexical string from the network and vice versa. Klex was written to conform to the part-of-speech tagging standards of the Korean Treebank Project, and is currently operating as the morphological analysis engine for the project. 1
Abstract. The paper addresses a problem of extraction of semantic information from Czech texts from the Web. The method described in this paper exploits existing linguistic tools created originally for a syntactically annotated corpus, Prague Dependency Treebank (PDT 2.0). We are working on development of a system which captures text of web-pages, annotates it linguistically by linguistic tools, extracts data and interprets the extracted data semantically in terms of web ontologies. The proposed extraction method is based on extraction rules – tree queries, which are adopted from the Netgraph application. Semantic interpretation of these rules provides semantics of the extracted data. We present some initial experiments in the domain of reports of traffic accidents.
Data-driven learning based on shift reduce parsing algorithms has emerged dependency parsing and shown excellent performance to many Treebanks. In this paper, we investigate the extension of those methods while considerably improved the runtime and training time efficiency via L2-SVMs. We also present several properties and constraints to enhance the parser completeness in runtime. We further integrate root-level and bottom-level syntactic information by using sequential taggers. The experimental results show the positive effect of the root-level and bottom-level features that improve our parser from 81.17 % to 81.41 % and 81.16 % to 81.57 % labeled attachment scores with modified Yamada’s and Nivre’s method, respectively on the Chinese Treebank. In comparison to well-known parsers, such as Malt-Parser (80.74%) and MSTParser (78.08%), our methods produce not only better accuracy, but also drastically reduced testing time in 0.07 and 0.11, respectively. 1
Abstract Self-deception is an important construct in social, personality, and clinical literatures. Although historical and clinical views of self-deception have regarded it as defensive in nature and operation, modern views of this individual difference variable instead highlight its apparent benefits to subjective mental health. The present four studies reinforce the latter view by showing that self-deception predicts positive priming effects, but not negative priming effects, in reaction time tasks sensitive to individual differences in affective priming. In all studies, individuals higher in self-deception displayed stronger positive priming effects, defined in terms of facilitation with two positive stimuli in a consecutive sequence, but self-deception did not predict negative priming effects in the same tasks. Importantly, these effects occurred both in tasks that called for the retrieval of self-knowledge (Study 1) and those that did not (Studies 2–4). This broad pattern supports substantive views of self-deception rather than views narrowly focused on self-presentation processes. Implications for understanding self-deception are discussed. Acknowledgement The authors acknowledge support from NIMH (MH 068241). Notes 1We also performed parallel analyses on the rating means from all studies. Study 1 involved self-judgements of emotion. We therefore thought it likely that self-deception would predict ratings in the task. Indeed, there were main effects of Self-Deception on emotion ratings, both for positive targets, F(1, 18) = 6.16, p<.05, and for negative targets, F(1, 18) = 6.09, p<.05. That is, individuals high in self-deception reported more intense positive emotions (Ms = 3.23 vs. 3.99) and less intense negative emotions (Ms = 2.70 vs. 2.01), relative to individuals low in self-deception. However, Studies 2–4 used affective rating tasks of low self-relevance, and thus we thought it less likely that self-deception would predict ratings in these tasks. Indeed, there were no relations along these lines, ps>.10. Priming effects on ratings, and possible interactions by self-deception, were weak and inconsistent across studies. 2Contact the first author for a list of stimuli. 3Note that these are relatively long judgement times. What is consistent with Studies 1–3 is that a differentiated rating task was used, which would somewhat necessarily slow processing speed (Fazio, 1990). However, this is viewed as beneficial for present purposes as longer processing times have produced stronger and more robust priming effects (e.g., Hines et al., 1996; Joordens & Becker, 1997). Judgement times were slower in Study 4 than in Studies 1–3 and this is attributed to a more differentiated scale (1–8) as well as elimination of the two-response aspect of the procedures used in the earlier studies. The evaluations assessed here clearly involved time and deliberation. However, priming of the responses, we believe, is reliant on the sorts of spreading activation processes shown to be involved in the continuous priming task (de Mornay Davies, 1998; McNamara & Altarriba, 1988; Shelton & Martin, 1992).
German genitive attributes are usually tagged as such in treebanks. However, it is well known that this information is not sufficient for determining the type of relation between head nouns and attributes, as genitive attributes can express many different semantic relations. Various linguistic classifications have been worked out, but to my knowledge, nobody has so far proposed to apply this linguistic knowledge to a corpus. The challenge here is to come up with a classification that is both easy to verify and sufficiently fine-grained. Using earlier linguistic approaches as guidelines, I propose in this paper a detailed annotation scheme for German genitive attributes based on readily identifiable noun features. First insights from its application to the Smultron Treebank show that it is easy to distinguish between the proposed classes and that my classification of genitive attributes can be related to a more general semantic annotation level.
Modernism constructs a set of linguistic norms to ensure the legitimacy of the express and accept system of significance.Post-modernism while try to subvert discourse-centrism and the disciplined cultural hegemony,and see the discourse and text as the different non-center system,which diminishes the trace of overall accounting with the unlimited,stretching and boundless laissez-faire context.The neo-post-modernism is not satisfied with this purely game of the extreme-expression way,it tries to build a level of trend-following writing and pleasure reading,such a cynical attitude can also lead to the fallen of the text.
As the first holder of the first chair in computational linguistics in Sweden, Anna Sagvall Hein has played a central role in the development of computational linguistics and language technology bo...
This article describes how semantic role resources can be exploited for preposition disambiguation. The main resources include the semantic role annotations provided by the Penn Treebank and FrameNet tagged corpora. The resources also include the assertions contained in the Factotum knowledge base, as well as information from Cyc and Conceptual Graphs. A common inventory is derived from these in support of definition analysis, which is the motivation for this work. The disambiguation concentrates on relations indicated by prepositional phrases, and is framed as word-sense disambiguation for the preposition in question. A new type of feature for word-sense disambiguation is introduced, using WordNet hypernyms as collocations rather than just words. Various experiments over the Penn Treebank and FrameNet data are presented, including prepositions classified separately versus together, and illustrating the effects of filtering. Similar experimentation is done over the Factotum data, including a method for inferring likely preposition usage from corpora, as knowledge bases do not generally indicate how relationships are expressed in English (in contrast to the explicit annotations on this in the Penn Treebank and FrameNet). Other experiments are included with the FrameNet data mapped into the common relation inventory developed for definition analysis, illustrating how preposition disambiguation might be applied in lexical acquisition.
The process of developing, implementing, and refining a registry data validation system is integral to optimal trauma registry operations. Describing registrar skill and proficiency in a manner that was once subjective can be replaced with objective assessment through the use of concrete rating guidelines and examples. The ability to standardize the evaluation of each registry abstract becomes the foundation for analyzing the overall accuracy of registry data. Key to the process is incorporating the validation rating tool as part of the data abstract. If properly implemented, the methodology described becomes a practical means for accuracy reporting, peer benchmarking, orientation and training, and performance management.
Socio-economic decisions are commonly explained by rational cost versus benefit considerations, whereas person variables have not much been considered. The present study aimed at investigating the degree to which dispositional power motivation and affective states predict socio-economic decisions. The power motive was assessed both indirectly and directly using a TAT-like picture test and a power motive self-report, respectively. After 9 months, 62 students completed an affect rating and performed on a money allocation task (social values questionnaire). We hypothesized and confirmed that dispositional power should be associated with a tendency to maximize one’s profit but to care less about another party’s profit. Additionally, positive affect showed effects in the same direction. The results are discussed with respect to a motivational approach explaining socio-economic behaviour.
The suitability of different parsing methods for different languages is an important topic in syntactic parsing. Especially lesser-studied languages, typologically different from the languages for which methods have originally been developed, pose interesting challenges in this respect. This article presents an investigation of data-driven dependency parsing of Turkish, an agglutinative, free constituent order language that can be seen as the representative of a wider class of languages of similar type. Our investigations show that morphological structure plays an essential role in finding syntactic relations in such a language. In particular, we show that employing sublexical units called inflectional groups, rather than word forms, as the basic parsing units improves parsing accuracy. We test our claim on two different parsing methods, one based on a probabilistic model with beam search and the other based on discriminative classifiers and a deterministic parsing strategy, and show that the usefulness of sublexical units holds regardless of the parsing method. We examine the impact of morphological and lexical information in detail and show that, properly used, this kind of information can improve parsing accuracy substantially. Applying the techniques presented in this article, we achieve the highest reported accuracy for parsing the Turkish Treebank.
A number of researchers have recently conducted experiments comparing “deep” hand-crafted wide-coverage with “shallow” treebank- and machine-learning-based parsers at the level of dependencies, using simple and automatic methods to convert tree output generated by the shallow parsers into dependencies. In this article, we revisit such experiments, this time using sophisticated automatic LFG f-structure annotation methodologies with surprising results. We compare various PCFG and history-based parsers to find a baseline parsing system that fits best into our automatic dependency structure annotation technique. This combined system of syntactic parser and dependency structure annotation is compared to two hand-crafted, deep constraint-based parsers, RASP and XLE. We evaluate using dependency-based gold standards and use the Approximate Randomization Test to test the statistical significance of the results. Our experiments show that machine-learning-based shallow grammars augmented with sophisticated automatic dependency annotation technology outperform hand-crafted, deep, wide-coverage constraint grammars. Currently our best system achieves an f-score of 82.73% against the PARC 700 Dependency Bank, a statistically significant improvement of 2.18% over the most recent results of 80.55% for the hand-crafted LFG grammar and XLE parsing system and an f-score of 80.23% against the CBS 500 Dependency Bank, a statistically significant 3.66% improvement over the 76.57% achieved by the hand-crafted RASP grammar and parsing system.
This paper reports on the work carried out developing MedLex+, a medical corpuslexicon workbench for Swedish. This project, which is still under active development, has been going on for some years now within the Department of Swedish language at Goteborg University. At the moment, the workbench incorporates: - an annotated collection of medical texts-including 20 million tokens and 45,000 documents, - a number of language processing software programs, including tools for collocation extraction, compound segmentation and thesaurus-based semantic annotation, and - a lexical database of medical terms-containing 5,000 medical entries. MedLex+ is a multifunctional lexical resource due to a structural design and content which can be easily queried. The medical workbench is intended to support lexicographers compiling lexicons and also lexicon users more or less initiated in the medical domain. MedLex+ can also assist researchers working on either lexical semantics or natural language processing (NLP) applications with focus on medical language. The linguistically and semantically annotated medical texts in combination with a set of smart queries turn the corpora into a rich repository of semasiological and onomasiological knowledge about medical terms and their linguistic, lexical and pragmatic properties. These properties are recorded in the lexical database with a cognitive profile. The MedLex+ workbench seems to offer a constructive help in many different lexical tasks.
Hungarian Academy of ScienceEötvös Loránd UniversityThis paper examines the Afro-Asiatic etymologies of Chadic lexical roots discussed by Olga V. Stolbova in her Chadic Lexical Database, Issue I (2005). The analysis is arranged according to the following sections: (1) Common Chadic reconstructions, (2) Isolated Chadic roots that nevertheless have Afro-Asiatic cognates. The paper represents the third part of my longer series of papers on addenda et corrigenda to Chadic lexical roots.
We present an initial ontology for tactical behaviors conducted by unmanned ground vehicles (UGVs). We focus on activities, which are the denotations of verbs, notably 'move' but also 'look (for)' and several others. These take collective subjects, allowing activities to be attributed to units at various hierarchical levels. The semantics of verbs must consider the denotations of their grammatical complements; that is, we must consider entire verb frames. The thematic relations of the noun-phrase complements are critical, but prepositions also play an important role. FrameNet is an online lexical database of frames derived from text corpora. Our other major resource is Levin's classification of verbs according to how changes in their frames affect their meanings. Although natural languages have a large variety of words for aspects of tactical behaviors, there is motivation to get by with as few basic verbs as possible. A variety of meanings can often be associated with a verb by altering its frame, and we can impose co-reference constraints on combinations of frames to generate structures denoting more complex activities. A simple grammar is developed for the verbs of interest. Protege-Frames ontologies include classes that inherit from linguistically inspired classes but capture domain-specific notions.