Papers reviewed and determined not to be word norm studies. Use the flag icon to report errors or suggest re-inclusion.
18265 papers
The paper deals with the preliminary findings from the morphologically annotated corpus of Lithuanian language (1 million running words). It was compiled and processed at the Center of Computational Linguistics, Vytautas Magnus University. Each annotation for an inflected word form of the corpus contains a lemma and a set of morphological features. The paper presents the strategy for automatic and manual annotation. Automatic annotation was carried out with the help of analyser-lemmatiser. Disambiguation of the homoforms was performed manually. Tag sets and the most prominent features of Lithuanian morphology are discussed in detail. The annotated corpus allowed us to measure the usage of parts of speech and their morphological features in contemporary Lithuanian language. The annotated corpus is of great importance for future development of parsing tools, treebanks and other NLP tools and resources for Lithuanian language.
This paper presents a general method to automatically build large knowledge bases from online lexical resources. While our experiments were limited to generate a knowledge base from WordNet, an online lexical database, the method is applicable to any type of dictionary organized around the elementary structure lexical entry - denition(s). The advantages of using WordNet, or richer online resources such as thesauri, as the source of a knowledge base, are outlined.
This paper presents the results of automatically inducing a Combinatory Categorial Grammar (CCG) lexicon from a Turkish dependency treebank. The fact that Turkish is an agglutinating free wordorder language presents a challenge for language theories. We explored possible ways to obtain a compact lexicon, consistent with CCG principles, from a treebank which is an order of magnitude smaller than Penn WSJ.
Current trends in language technology require treebanks that do not stop at the level of constituent structure, but include deeper and richer levels of analysis, including appropriate meaning structures. Capturing sufficient detail at different levels of linguistic description is too complex a task to be practically achievable by manual annotation or shallow parsing; rather it requires sophisticated tools that help secure the consistency of parallel but different structures. We are constructing a multilevel treebanking tool that incorporates a deep parser and grammar for Norwegian. Thus, we are tightly linking our treebank to grammar development so as to achieve a sound embedding in grammatical theory and yield more useful results for applications.
Traditional Chinese text chunking approach is to identify phrases using only one model and same features. It is shown that one model couldn't comprise each phrase's characteristics, and same features are not suitable to all phrases, data sparseness also appears. Multi-agent strategy uses several model and sensitive features of each phrase to identify different phrases. This paper describes the multi-agent strategy applied in the identification of Chinese phrases whose main features are: 1) easy and quick communication between phrases; 2) avoidance of data sparseness. Through testing on Chinese Penn Treebank, F score of Chinese text chunking using multi-agent strategy achieves to 95.82%, which is higher than the best result that has been reported.
An automatic method for annotating the Penn-II Treebank (Marcus et al., 1994) with high-level Lexical Functional Grammar (Kaplan and Bresnan, 1982; Bresnan, 2001; Dalrymple, 2001) f-structure representations is presented by Burke et al. (2004b). The annotation algorithm is the basis for the automatic acquisition of wide-coverage and robust probabilistic approximations of LFG grammars (Cahill et al., 2004) and for the induction of subcategorisation frames (O’Donovan et al., 2004; O’Donovan et al., 2005). Annotation quality is, therefore, extremely important and to date has been measured against the DCU 105 and the PARC 700 Dependency Bank (King et al., 2003). The annotation algorithm achieves f-scores of 96.73% for complete f-structures and 94.28% for preds-only f-structures against the DCU 105 and 87.07% against the PARC 700 using the feature set of Kaplan et al. (2004). Burke et al. (2004a) provides detailed analysis of these results. \nThis paper presents an evaluation of the annotation algorithm against PropBank (Kingsbury and Palmer, \n2002). PropBank identifies the semantic arguments of each predicate in the Penn-II treebank and annotates their semantic roles. As PropBank was developed independently of any grammar formalism it provides a platform for making more meaningful comparisons between parsing technologies than was previously possible. PropBank also allows a much larger scale evaluation than the smaller DCU 105 and PARC 700 gold standards. In order to perform the evaluation, first, we automatically converted the PropBank annotations \ninto a dependency format. Second, we developed conversion software to produce PropBank-style semantic annotations in dependency format from the f-structures automatically acquired by the annotation algorithm from Penn-II. The evaluation was performed using the evaluation software of Crouch et al. (2002) and Riezler et al. (2002). Using the Penn-II Wall Street Journal Section 24 as the development set, currently we achieve an f-score of 76.58% against PropBank for the Section 23 test set.
We have constructed a large scale and detailed database of lexical types in Japanese from a treebank that includes detailed linguistic information. The database helps treebank annotators and grammar developers to share precise knowledge about the grammatical status of words that constitute the treebank, allowing for consistent large scale treebanking and grammar development. In this paper, we report on the motivation and methodology of the database construction. 1
The PropBank project is creating a corpus of text annotated with information about basic semantic propositions. PropBank I (Kingsbury & Palmer, 2002) added a layer of predicateargument information, or semantic roles, to the syntactic structures of the English Penn Treebank. This paper presents an overview of the second phase of PropBank Annotation, PropBank II, which is being applied to English and Chinese, and includes (Neodavidsonian) eventuality variables, nominal references, sense tagging, and connections to the Penn Discourse Treebank (PDTB), a project for annotating discourse connectives and their arguments. 1
This paper discusses an annotation scheme for Korean null pronouns, which were used in annotating three kinds of Korean text corpora including Penn Korean Treebank. In annotating the corpora, null pronouns and their antecedents were marked up for their type and reference, with coreference relation tracked by numeric identifiers. Based on the annotation scheme, an outline of a potential pronoun resolution strategy is also proposed. The resulting dataset of annotated text is rather small at 11,834 words; we hope the null pronoun classification and annotation scheme proposed in this study will serve as a basis in developing a large-scale annotated corpus in the future.
The syntactically annotated corpora, commonly called ‘treebanks’, play an important role in empirical linguistics as well as in machine learning methods in natural language processing. After a brief summarization of several treebank annotation of different language, we proposed a new annotation scheme for Chinese treebank in this paper. Under this scheme, every Chinese sentence will be annotated with a complete parse tree, where each non terminal constituent is assigned with two tags. One is the syntactic constituent tag, which describes its external functional relation with other constituents in the parse tree. The other is the grammatical relation tag, which describes the internal structural relation of its sub components. These two tag sets consist of 16 and 27 tags respectively. They form an integrated annotation for the syntactic constituent in a parse tree through top down and bottom up descriptions. Based on this scheme, we built a 1,000,000 words Chinese treebank covering a balanced collection of journalistic, literary, academic, and other documents. The annotating experiments on different kinds of complex linguistic phenomena show the availability and compatibility of this annotation scheme.
М. И. ШАПИР... ЭСТЕТИКА НЕБРЕЖНОСТИ В ПОЭЗИИ ПАСТЕРНАКА (Идеология одного идиолекта)1... © 2004 г.... В первой части работы анализируются случаи непроизвольных двусмысленностей в поэзии Пастернака (лексико-фразеологических, грамматических, стилистических); во второй части они рассматриваются в ряду разного рода коллоквиализмов, нарушающих привычные нормы книжного языка и классической ритмики; наконец, в третьей части статьи делается попытка понять психологическую, эстетическую и социальную подоплеку общей установки поэта на опрощение стихотворного языка и его сближение с разговорной речью.... The first section of this article analyzes cases of involuntary ambiguity (lexical, phraseological, grammatical, and stylistic) in the poetry of Pasternak. The second section examines these cases in the context of various colloquialisms which violate the conventional norms of standard literary Russian and Russian classical prosody. The third and the last section attempts to understand the psychological, aesthetic and social background of Pasternak's general tendency towards the simplification of poetic language and its convergence with free colloquial speech.......А ты прекрасна без извилин...... Много лет назад Уильям Эмпсон, незаурядный английский поэт и филолог, расценил семантическую неопределенность (ambiguity) как неотъемлемое свойство поэзии [I]2. В последнее время интерес к поэтической неоднозначности растет и у российских лингвистов: недавно специальное исследование ей посвятил Н. В. Перцов [2] (ср. [3]). В своей книге и в предшествующих статьях я тоже обращался к этой теме (см. [4, с. 12 - 19] и др.). Но до сих пор филологи сосредоточивались главным образом на преднамеренном двоении смыслов; что же касается двусмысленностей непроизвольных (либо кажущихся таковыми), то им должного внимания не уделялось. Это упущение мне бы хотелось восполнить: сначала предметом моего анализа станет такое ветвление у Пастернака, которое, насколько можно судить, не входило в расчеты автора; затем найденные факты я попробую поставить в более широкий лингвистический и наконец - в идеологический контекст. Таким образом, против обыкновения я буду изучать не информацию, а шум, который, однако, на свой лад оказывается весьма информативным.... Размышляя над примерами, постараемся не терять из виду суть проблемы: дело не в том, что какой-то фрагмент текста не допускает верной интерпретации, - дело в том, что он объективно допускает интерпретацию неверную. Именно ощущение неадекватности вторых и третьих смыслов позволяет нам выделять оговорки среди других случаев неоднозначности. Разумеется, это ощущение может сбивать с толку: насчет авторского замысла нам дано лишь строить догадки. Наивно было бы верить, что в поэтическом тексте намеренное всегда надежно отличается от ненамеренного: в душе писателя мы читать не умеем, но попытаться его понять - обязаны3.... Явление, о котором пойдет речь, еще не имеет адекватного терминологического выражения.... 1 Исследование выполнено при поддержке Российского гуманитарного научного фонда (проект 04 - 04 - 00055а). Исправляя и дополняя исходный вариант статьи, автор имел счастливую возможность пользоваться советами и замечаниями М. В. Акимовой, С. Г. Болотова, М. Л. Гаспарова, Ф. Н. Двинятина, В. З. Демьянкова, А. А. Добрицына, И. Г. Добродомова, А. К. Жолковского, Вяч. Вс. Иванова, А. А. Илюшина, Т. М. Левиной, Т. М. Николаевой, А. Б. Пеньковского, И. А. Пильщикова, Н. В. Перцова, О. Ронена, Т. В. Цивьян. Особая признательность Е. Б. Пастернаку и Е. В. Пастернак, помогавшим автору и его поддерживавшим на протяжении всей работы.... 2 Латинское слово ambiguitas 'двусмысленность' соответствует древнегреческому (амфиболия), усвоенному русской научной терминологией.... 3 На пушкинском пленуме Союза писателей (1937) Пастернак заявил: не только намеренных двусмысленностей, но и таких провалов последнего сорта, которые бы давали повод для двусмысленного понимания и в неумышленном плане, - я за собой не помню. Вообще двусмысленности при настоящей любви к искусству немыслимы [5, т. 4, с. 644; 6, с. 401, 404 примеч. 36].... стр. 31... В арсенале испытанных средств филологического метаязыка наиболее подходящим к случаю могло бы стать понятие авторской глухоты, закрепленное в Поэтическом словаре А. П. Квятковского. Это условный термин, предложенный М. Горьким; понимаются под ним явные стилистические и смысловые ошибки..> не замеченные автором. Их можно трактовать по-разному: иногда авторская глухота - результат небрежности или неряшливости, иногда она возникает непроизвольно, когда увлечение главной задачей заслоняет отдельные детали. Явления г, - продолжает Квятковский, - свойственны не только рядовым писателям, но и большим мастерам [7, с. 10]. Он приводит примеры из Пушкина, Лермонтова, Плещеева, Фета, Маяковского, Багрицкого и Уткина. Завершается статья указанием на то, что к А г можно отнести явления сдвига, и ссылками на тематически близкие статьи: Амфиболия, Анаколуф, Солецизм [7, с.
Abstract Advertising copy writers and journalists deserve recognition for the innovative and creative way they use Afrikaans and English in advertising and reportage. An analysis of collected data clearly showed that these texts encompass much more than artistic and intellectual inventive language. More often these word-formations in the media breach standard word-formation norms in a creative way. From a linguistic perspective, the mental and linguistic lexicons consist of simplex and complex words, the latter formed morphologically by means of inflection, deriviation and compounding and combinations of the former. These lexical sets conform to canonical patterns and their operational parameters. These parameters of the canon restrict the kinds of morphemes and words that can appear in the lexicon of the language, which means that the productive patterns of word-formation are rule-governed. A taxonomy of the norm-breaching patterns and systems will be proposed which will make it possible to describe the renewal of the vocabulary in terms of rule-changing creativity against the background of rule-governed productivity. The objective of this article is to distinguish linguistically among concepts in the linguistic lexicon, neologisms and occasional creations. What kinds of information are then needed when we comprehend a (new) word? It is evident that implicit knowledge of all the sub fields of linguistics, namely phonological, morphological, syntactic, semantic and pragmatic information, is needed. Advertensiekopieskrywers en joernaliste verdien erkenning vir die innoverende en kreatiewe wyse waarop hulle Afrikaans en Engels as reklametaal en verslagtaal gebruik. 'n Analise van die versamelde data het getoon dat die innovasies veel meer behels as artistieke en intellektueel vindingryke taalgebruik binne die norme van die standaardtaal, aangesien woordvorming in die media dikwels normwysigend is t.o.v. die kanonieke leksikon. Vanuit 'n taalwetenskaplike perspektief bestaan die mentale en linguistiese leksikons uit simplekse en morfologies gelede woorde. Laasgenoemde is gevorm deur fleksiemorfeme, afleiding en samestelling en kombinasies daarvan. Hierdie versameling woorde is geskep volgens die kanonieke patrone en hulle operasionele parameters. Die parameters beperk die tipe morfeme en woorde wat in die leksikon van die taal kan voorkom, aangesien produktiewe woordboupatrone reëlgehoorsaam is. Met hierdie beskrywing as uitgangspunt sal voorts daaraan aandag gegee word om 'n taksonomie van die normwysigende geledingsreëls in die media te gee wat 'n vergelyking tussen reëlgehoorsame produktiwiteit en reëlwysigende kreatiwiteit moontlik maak. Die artikel het dan ook ten doel om linguisties te onderskei tussen die begrippe linguistiese leksikon, nuutskeppings en geleentheidskeppings. Watter tipe inligting is dus nodig wanneer ons 'n (nuwe) woord begryp? Dit blyk uit die data-analise dat implisiete kennis van die verskillende subdissiplines van die Linguistiek onderliggend is aan ons begrip van veral nuwe woorde. Dit sluit o.a. fonologiese, morfologiese, sintaktiese, semantiese en pragmatiese inligting in.
This paper describes a new, large scale discourse-level annotation project -- the Penn Discourse TreeBank (PDTB). We present an approach to annotating a level of discourse structure that is based on identifying discourse connectives and their arguments. The PDTB is being built directly on top of the Penn TreeBank and Propbank, thus supporting the extraction of useful syntactic and semantic features and providing a richer substrate for the development and evaluation of practical algorithms.
The Penn Discourse TreeBank (PDTB) is a new resource built on top of the Penn Wall Street Journal corpus, in which discourse connectives are annotated along with their arguments. Its use of standoff annotation allows integration with a stand-off version of the Penn TreeBank (syntactic structure) and PropBank (verbs and their arguments), which adds value for both linguistic discovery and discourse modeling. Here we describe the PDTB and some experiments in linguistic discovery based on the PDTB alone, as well as on the linked PTB and PDTB corpora.
The paper presents an unlexicalized probabilistic parsing model for German trained on the Negra treebank. Evaluation is performed with respect to constituency and dependency measures. It is observed that existing models based on Parent Encoding and Markovization optimize for constituency measures at the expense of dependency performance (at least in German). Several linguistically inspired transformation and annotation schemes are proposed which do help with dependency measures. Finally, it is shown that performance compares well with published results for German.
Reviewed by: A rainbow of corpora: Corpus linguistics and the languages of the world ed. by Andrew Wilson, Paul Rayson, and Tony McEnery Heiko Narrog A rainbow of corpora: Corpus linguistics and the languages of the world. Ed. by Andrew Wilson, Paul Rayson, and Tony McEnery. (Linguistics edition 40.) Munich: LINCOM Europa, 2003. Pp. 165. ISBN 3895868728. $73.20 (Hb). This volume is a collection of papers that were originally presented at Corpus Linguistics 2001, a conference held 30 March–2 April 2001 at Lancaster University (UK). The sister volume to the ‘English-oriented’ Corpus linguistics by the Lune: A festschrift for Geoffrey Leech (Frankfurt: Peter Lang, 2003), it includes contributions that are concerned with non-English languages. The languages dealt with in the present volume range from the better-known Indo-European languages to Biblical Hebrew, Korean, and Arabic. Content-wise, the individual articles can be roughly divided into three categories. First, there are three diachronically oriented papers from the workshop ‘Corpus linguistics, ancient languages, and older language periods’: Beatrix Färber on a corpus of Medieval Irish (19–26), Wolf-Dieter Syring on the design and usage of a text database of Biblical Hebrew (141–52), and Matthew Brook O’Donnell, Stanley E. Porter, and Jeffrey T. Reed on the database-assisted discourse analysis of texts from the New Testament (109–21). The other papers, which come from the main session, deal with modern languages. Half of them are primarily concerned with corpora as such that is, their design, mark-up, usage, and so on. Anne Abeillé, Lionel Clément, Alexandra Kinyon, and François Toussenel discuss the PARIS 7 annotated corpus for French (1–10), and Martin Beaudoin and Michel Samard present their work on a large corpus of written Canadian French (11–18). R. Rossini Favretti, F. Tamburini, and C. de Santis deal with the CORIS corpus of written Italian (27–38), and Eva Hajičová and Petr Sgall discuss the Prague Dependency Treebank (39–50). Shereen Khoja, Roger Garside, and Gerry Knowles present a tagset for the morphosyntactic tagging of Arabic (59–72),and Kiril Simov, Gergana Popova, and Petya Osenova introduce a HPSG-based syntactic treebank of Bulgarian (133–40). Apart from language-specific interests, the papers by Abeillé and colleagues and Hajičová and Sgall are particularly impressive. The former shows how a corpus tagged with higher accuracy than previous corpora can completely overturn research results based on corpora tagged with lower accuracy. The latter presents a corpus that is marked up for syntactic features, including topic-focus structure, with a depth that is probably unmatched in any language. The rest of the papers present corpus-based linguistic research. The paper by Beom-mo Kang, Hung-gyu Kim, and Myung-hoe Huh applies Douglas Biber’s multidimensional text analysis to Korean (51–57); Maarten Lemmens investigates the functions and meaning range of posture verbs in Swedish from a typological perspective (73–85); Martina Möllering analyzes the use of the German modal particle eben in a corpus of telephone conversations (87–97); P.-O. Nilsson shows how Swedish texts translated from English exhibit systematically different [End Page 904] lexical and grammatical patterns from those found in texts written originally in Swedish (99–107); Katja Ploog explores the syntax of pronominal subjects in Abidjanee French in contrast to standard French (123–32); and Adriana Vlad, Adrian Mitrea, and Mihai Mitrea investigate Romanian texts from a stochastic perspective (153–65). Many of these contributions combine quantitative corpus methods with qualitative methods of investigation and persuasively demonstrate how corpus-based research can contribute to broader linguistic issues. On the whole, not all papers in this volume are of the same theoretical interest, but each of them is at least informative. In contrast, the editing is rather disappointing. There is a table of contents and a...
This paper presents a Chinese treebank based decision tree approach to identify Chinese BNP. A self-learning mechanism is integrated into our model which includes the following steps: auto-extraction of POS string sequences (BNP rules) and their context information from the corpora and ID3 algorithm based tree training. Experimental results show good performances of our method.
This paper deals with wordnet development tools. It presents a designed and developed system for lexical database editing, which is currently employed in many national wordnet building projects. We discuss basic features of the tool as well as more elaborate functions that facilitate linguistic work in multilingual environment.
this paper I discuss a corpus-based approach to the analysis of some phenomena of lexical semantics using empirical data drawn from a Russian-German paraUel corpus ofDostoevskij's Idiot together with its German translations. This parallel corpus is part of the Austrian Academy Corpus (AAC) at the Austrian Academy of Sciences in Vienna. The subject of investigation is lexical co-occurrences which determine the combinatorial profile ofa word. A corpus-based analysis oflexical co-occurrences contributes to both monolingual and buingual lexicography by providing new and more detailed insights into the contextual behaviour of a word. From the diachronic perspective the semantic change comes about at the periphery of the combinatorial profile of a given word. Some of the peripheral co-occurrences can become so frequent that they drift from periphery to centre while others fall out of use and start to be perceived as norm violations. The comparison ofcombinatorial profiles ofthe same word in the 1860s and in present day Russian proves to be an efficient instrument for defining the combinatorial norms of a given word against the background of its near-synonyms. The comparison of a given word with all possible translation equivalents has a similar function.
This paper describes an incremental parsing approach where parameters are estimated using a variant of the perceptron algorithm. A beam-search algorithm is used during both training and decoding phases of the method. The perceptron approach was implemented with the same feature set as that of an existing generative model (Roark, 2001a), and experimental results show that it gives competitive performance to the generative model on parsing the Penn treebank. We demonstrate that training a perceptron model to combine with the generative model during search provides a 2.1 percent F-measure improvement over the generative model alone, to 88.8 percent.
Tree-based approaches to alignment model translation as a sequence of probabilistic operations transforming the syntactic parse tree of a sentence in one language into that of the other. The trees may be learned directly from parallel corpora (Wu, 1997), or provided by a parser trained on hand-annotated treebanks (Yamada and Knight, 2001). In this paper, we compare these approaches on Chinese-English and French-English datasets, and find that automatically derived trees result in better agreement with human-annotated word-level alignments for unseen test data.
There are three types of sentences that form all existing natural languages: verbal sentences (e.g."I read the book."),copulative sentences (e.g."The book is on the table."),and existential sentences (e.g."There is a book on the table.").Syntactic and semantic recognition of these sentence types are crucially important in computational linguistics although there has not been any significant work towards this end.This thesis, in an attempt to fill this evident gap, is on identifying and assigning semantic categories of Turkish existential sentences in print.Existential sentences in Turkish are minimally characterized by the two existential particles var, meaning there is/are, and yok, meaning there is/are no.In addition to these most basic meanings, other senses of existential particles are possible, which can be categorized into groups such as case existentials and possession existentials.Our system does shallow semantic parsing in defining the predicate-argument relationships in an existential sentence on a word-byword basis, via utilizing Support Vector Machines, after which it proceeds with the semantic categorization of the whole sentence.For both of these tasks, our system produces promising results, in terms of accuracy and precision/recall, respectively.Part of this research contributes to the annotation of the METU-Sabanc Turkish Treebank with semantic information.
The aim of this paper is to present the theoretical principles underlying the making of a lexical database of English collocations of non-specialized words used in scientific language. This project was prompted by the shortage of reference tools providing information about the use and combinatorial properties of general words in specific registers. A case study will illustrate that in scientific texts, words, especially polysemous verbs, have a distinct semantic and combinatorial behaviour. Following the assumption that the meaning and the grammatical and collocational patterns of words are interrelated, we suggest that context-specific information should be included in specialized reference tools to facilitate the written production of scientific texts by nonnative speakers ofEnglish.
In natural language processing a huge amount of structured data is constantly used for the extraction and presentation of grammatical structures in sentences. For example the Chinese Treebank corpus developed at the Institute of Information Science Academia Sinica Taiwan is a semantically annotated corpus that has been used to help parse and study Chinese sentences. In this setting users usually use structured tree patterns instead of keywords to query the corpus.
OBJECTIVE: Transcutaneous electrical nerve stimulation (TENS) is a technique widely used in clinical practice to control pain, although its clinical efficacy remains controversial. Though many mechanisms have been proposed for its analgesic effects, there is a conspicuous lack of experimentally controlled research investigating whether TENS analgesia is related to its effects on the sympathetic nervous system (SNS). METHODS: Using an established psychophysiological paradigm, the present study investigated the effects of high-frequency/low-intensity TENS, low-frequency/high-intensity TENS, and sham TENS on the perception of experimental pain and SNS function in healthy volunteers. Measures of heart rate, digital pulse volume, and skin conductance were recorded during a 20-minute TENS stimulation period and in anticipation of a series of painful electric shocks prior to and following TENS stimulation. Healthy volunteers rated the intensity of the shocks using a 0-10-point verbal pain rating scale. RESULTS: The three TENS conditions failed to differentially effect SNS responses during either the 20-minute TENS treatment period or the shock anticipation periods, and TENS did not affect ratings of pain intensity to the shock stimuli. CONCLUSIONS: While these results may not generalize to acute or chronic pain patients, within the limitations of the present experimental paradigm, no support was found for TENS affecting either SNS function or acute experimental pain perception.
We present a linguistically-motivated algorithm for reconstructing nonlocal dependency in broad-coverage context-free parse trees derived from treebanks. We use an algorithm based on loglinear classifiers to augment and reshape context-free trees so as to reintroduce underlying nonlocal dependencies lost in the context-free approximation. We find that our algorithm compares favorably with prior work on English using an existing evaluation metric, and also introduce and argue for a new dependency-based evaluation metric. By this new evaluation metric our algorithm achieves 60% error reduction on gold-standard input trees and 5% error reduction on state-of-the-art machine-parsed input trees, when compared with the best previous work. We also present the first results on non-local dependency reconstruction for a language other than English, comparing performance on English and German. Our new evaluation metric quantitatively corroborates the intuition that in a language with freer word order, the surface dependencies in context-free parse trees are a poorer approximation to underlying dependency structure.
Statistics from about 17,000 occurrences of the structures “N1 is N1” and “N1 is N7” have proved that (a) there is a functional difference between the two predicative cases and (b) there are strong norms for selecting one of the two cases in communication. The nominative case is a strong norm if the communicative function of the sentence predicate is to (a) identify a sort of denotation in a demonstrative act, (b) identify a sort of denotation in a nominative act, (c) define using qualification or (d) qualify in an expressive manner. The nominative is preferred when stating a person’s profession, in most cases (except for professions such as minister, director, manager etc., especially in a specific sentence structure and discourse function). The statistics show that (a) for 630 different predicate nouns (PN), only the nominative is used and for 270 PN, only the instrumental is used; (b) for 900 PN, one of the two predicative cases is used either exclusively or with a strong preference; (c) for only 89 PN, the use of the two cases is balanced. The corpus statistics for the two predicative cases show that the selection of one of these cases is semantically determined and to a great degree lexically bound.
A program called the Generalized Electronic Interviewing System (GEIS) was developed for conducting interviews, using computer-assisted telephone interview (CATI) and interactive voice response (IVR) modes without the need for a programmed interface. GEIS questionnaires were prepared using a common script syntax in all supported modes. Scripted development allowed for rapid interview development without the need for programming. A GEIS script specified the following: question texts, including variable texts; answer option texts; numeric codes for answers; range check information; logical question-branching information; interview status information; do-loop information; and IVR information, such as key codes and voice messages. GEIS thoroughly checked scripts for logical or syntactical errors. GEIS required SAS Version 8.0, and survey data were accumulated within SAS data sets. An application of GEIS to conduct a survey involving CATI, IVR, and a combined hybrid method is described. The CATI results deviated in the direction expected for sensitive questions, whereas IVR obtained a small sample size, rendering the results unreliable. However, the hybrid method was found to provide more accurate telephone survey data on alcohol consumption than did CATI alone. The program may be downloaded from the Psychonomic Society Web archive atwww.psychonomic.org/archive/.
Function tags are a context-sensitive annotation applied to words and phrases of natural language text, marking their syntactic or semantic role within a larger utterance. As researchers improve results on various other problems in “pure” natural language processing (e.g part-of-speech tagging, parsing), those who work in the more “applied” NLP fields (e.g. question-answering, temporal analysis) are seeking more powerful sorts of linguistic annotation as input for their own systems. Hence, function tags. In the first part of the thesis, I present the problem of function tagging: why it is an interesting problem, who has worked on similar thing, and what exactly I intend to do. I briefly review the function tags of the Penn treebank, and explain the specific metrics by which I will evaluate my work. In the second part of the thesis, I introduce the many features that I will use to train a function tagging system, and then I present some systems that make use of them: one using feature trees, one using decision trees (briefly), and one using perceptron models. For each system, I give a brief historical perspective, an overview of where it has been used before and why I think it will be useful in this task. I will then try a number of feature combinations with interesting properties; and finally, present the best-performing tweaked-out version of that system. Finally, in the third part of the thesis, I bring them all together and discuss the advantages and disadvantages of each system in various situations. More interestingly, I will present an analysis of what features prove to be the most helpful for the different function tagging subtasks. Lastly, I will present a comparison to other systems performing related tasks, and speculate on some interesting future work.
I shall explore the implications for lexical resources of my work on ATT-Meta, a reasoning system designed to work out the signicance of a broad class of metaphorical utterances. This class includes imap-transcendingi utterances, resting on familiar, general conceptual metaphors but go beyond them by including source- domain elements that are not handled by the mappings in those metaphors. The system relies heavily on doing reasoning within the terms of the source domain rather than trying to construct new mapping relationships to handle the unmapped source-domain elements. The approach would therefore favour the use of WordNet-like resources that facilitate rich within-domain reasoning and the retrieval of known cross-domain mappings without being constrained to facilitate the creation of new mappings. The approach also seeks to get by with a small number of very general mappings per conceptual metaphor. The research has also led me to a radical language-user-relative view of metaphor. The question of whether an utterance is metaphorical, what conceptual metaphors it involves, what mappings those metaphors involve, what word-senses are recorded in a lexicon, etc. are all relative to specic language users and shouldn’t be regarded as something we have to make objective decisions about. This favours a practical approach where natural language applications can differ widely on how they handle the same potentially metaphorical utterance because of differences in lexical resources used. The user-relativity is also friendly to a view where the presence of a word-sense in a lexicon has little to do with whether that sense is gurati ve or not. This stance is related to, Patrick Hanks’ view that we should focus on norms and exploitations rather than on gurati vity. The research has furthermore led me to a deep scepticism about the ability to rely in denitions of metaphor on qualitative differences between domains. Scepticism about domains then causes additional difculty in distinguishing between metaphor and metonymy. At the panel I will outline a particular view of the distinction.
This paper describes an algorithm for detecting empty nodes in the Penn Treebank (Marcus et al., 1993), finding their antecedents, and assigning them function tags, without access to lexical information such as valency. Unlike previous approaches to this task, the current method is not corpus-based, but rather makes use of the principles of early Government-Binding theory (Chomsky, 1981), the syntactic theory that underlies the annotation. Using the evaluation metric proposed by Johnson (2002), this approach outperforms previously published approaches on both detection of empty categories and antecedent identification, given either annotated input stripped of empty categories or the output of a parser. Some problems with this evaluation metric are noted and an alternative is proposed along with the results. The paper considers the reasons a principle-based approach to this problem should outperform corpus-based approaches, and speculates on the possibility of a hybrid approach.
Automatic classification of web pages is an effective way to facilitate the process of retrieving information from the Internet. Currently, two major classification methods are used in this area: keyword-based classification and sense-based classification. For keyword-based classification, keywords often have different semantic meanings, and the correct keyword matching is largely based on using exactly the same keywords. Thus, the classification results of keyword-based classification are not always satisfying. Many sense-based classification algorithms and systems have been presented, but they pay little attention to the relationship between senses. In this dissertation, we present a method to automatically classify documents based on the meanings of words and the relationships between groups of meanings or concepts. The classification algorithm builds on the word sense structures provided by a lexical database, which not only arranges words into groups of synonyms, but also arranges these groups of synonyms into hierarchies that represent the relationships between concepts. Another problem with current classification systems is that most of them ignore the conflict between the fixed number of categories and the growing number of documents being added to the system. To address this problem, a category-based clustering method is developed to automatically extract a new category from a category that needs to be split. A category must be divided when the number of documents in the category is larger than a predefined size. Experimental results show that the semantic hierarchy classification algorithm increases the classification accuracy by 13% compared to existing sense-based classification algorithms. The category-based clustering algorithm achieves a higher quality cluster than other existing methods that do not use category information. Combining the automatic classification based on word meanings and the dynamic addition of new categories based on clustering, we develop a new system to meet the current and future needs of a growing Internet.
Schwarz (2001, 2002) proposed the ex-Wald distribution, obtained from the convolution of Wald and exponential random variables, as a model of simple and go/no-go response time. This article provides functions for the S-PLUS package that produce maximum likelihood estimates of the parameters for the ex-Wald, as well as for the shifted Wald and ex-Gaussian, distributions. In a Monte Carlo study, the efficiency and bias of parameter estimates were examined. Results indicated that samples of at least 400 are necessary to obtain adequate estimates of the ex-Wald and that, for some parameter ranges, much larger samples may be required. For shifted Wald estimation, smaller samples of around 100 were adequate, at least when fits identified by the software as having ill-conditioned maximums were excluded. The use of all functions is illustrated using data from Schwarz (2001). The S-PLUS functions and Schwarz’s data may be downloaded from the Psychonomic Society’s Web archive, www. psychonomic.org/archive/.
This paper deals with the relationship between rules and conventionality, as reflected in a case study of the two structures N 1 de N 2 and N 1 du N 2. Most occurrences of these can be explained via two rules which evoke ease of referent identification and are a subset of the principles governing the ± definiteness opposition in French. Via dictionaries and Google searches, however, we detect variation that our rules cannot predict, typically when N 2 is non-countable and abstract. Local semantic contrasts, i.e. dependent on the lexical content of N 1 or N 2, further complicate the picture. We conclude that Coseriu's Norm, alias conventionality, must be recognized as a factor coexisting with the rules.
Despite much research, the distinctive personality characteristics of entrepreneurs are yet to be established and the influence of personality on entrepreneurial behaviour remains unclear. This is particularly evident in our understanding of the personal response of entrepreneurs to business failure. In this thesis the Life Story Model of Identity proposed by McAdams' (1993; McAdams & Pals, 2006) narrative theory of personality formed the main theoretical approach to investigating these two related aspects in the psychological understanding of entrepreneurs. This model overcomes some of the limitations of previous personality research by permitting investigation of personality within the entrepreneurial environment and provides a wholistic and complex view of personality as expressed in the entrepreneurs' own words. The model's qualitative methodology and theoretical emphasis on personal meaning making also rendered it most suitable for exploring entrepreneurs' personal response to business failure. McAdams' (1993; McAdams & Pals, 2006) Life Story Interview was employed to explore the self-narrative identities of 40 highly successful entrepreneurs (39 males, one female). Participants were managing directors of businesses sourced from two lists of the fastest growing small to medium companies in Australia, as compiled by the Australian business magazine, the 'Business Review Weekly'. Participants were the founders of their businesses, and had been pursuing entrepreneurship for at least five years. A broad range of business sectors were represented, including computer services, manufacturing, engineering and communications. Prior to interviews, participants completed the Life Story Interview Questionnaire (LSIQ), which was an adapted form of the Life Story Interview that requested written responses to open-ended questions about the content of participants' life stories. A section requesting affective ratings for key events, derived from Herman’s (Hermans & Hermans-Jansen, 1995) Extended List of Affect Terms, was included to further the exploration of life story themes. A second questionnaire, comprised of measures of personality and a measure of psychological symptoms was also completed. During interviews, participants' responses to the LSIQ were discussed, concentrating on further investigation of the key events that defined their life stories. Findings revealed a prototypical life story of the entrepreneur, highlighting distinctive, commonly shared personality characteristics, with much of their selfnarrative identity grounded in experiences within the entrepreneurial environment. The prototypical life story contained a core theme with an agentic-type emphasis on strengthening the self, and a lesser theme with a communion-type emphasis on valuing relationships. Each of these themes comprised two further themes. The selfstrengthening theme included a redemptive theme of overcoming difficulties in a way that left the protagonist feeling stronger and more able to influence their environment, and a positively toned theme of drawing strength and confidence in one's abilities from achievements and successes. The relational theme included a redemptive theme of responding to private relationship difficulties and losses in one area by strengthening other private relationships, and a negatively toned, sometimes contaminated theme, of experiencing either private or professional relational difficulties and losses as irresolvable. The resulting prototypical life story of the entrepreneur was a story centred upon overcoming adversity and celebrating personal achievement, of confirming and boosting confidence in one’s abilities and a sense of personal power to influence their environment. Running parallel to this main storyline was a less prominent plot involving the importance of relationships, with difficulties and losses sometimes redeemed and sometimes left unresolved. To investigate the impact of business failure, participants were asked to describe their experience of business failure as a key life story event during the Life Story Interview. Additional open-ended questions explored important elements of their critical and retrospective responses. A commonly shared personal response was evidenced. Despite being strongly identified with their business at the time, the failure was evaluated in business rather than personal terms. The causes were most often attributed to a combination of internal and external factors, but business recovery was attributed exclusively to their own actions. The self-narrative meanings given to this event centred upon overcoming the business failure in a self-strengthening way as either: mastering business conflict situations; learning entrepreneurial skills; or affirming entrepreneurial self-confidence. In the midst of the failure, most entrepreneurs remained highly optimistic about their chances of success in the future, based largely upon confidence in their ability to bring about business recovery. Practised coping skills were used to manage negative feelings arising from the failure, and an active problem-solving approach was adopted. When reflecting upon their experience, there was an absence of rumination and regret about the business failure. Instead, it was regarded as an inevitable and even welcome event that provided valuable entrepreneurial learning. In making sense of the failure in relation to the rest of their self-narrative identity, most entrepreneurs were able to integrate its meaning within their larger self-story. This was done by relating the business’ recovery to a story of overcoming obstacles, or by containing the business’ failure within a story of either repeated success or sustained self-confidence in one's ability. It was concluded that these entrepreneurs shared a particular type of selfnarrative identity that was conducive to the pursuit of entrepreneurship; positively influencing their behaviour within the entrepreneurial environment and having particular relevance to how they personally responded to business failure. These findings advance understanding of the personality of entrepreneurs, and begin to inform what constitutes a constructive personal response to the event of business failure.
INTRODUCTION Several authors described cases of dissociated impairment in naming nouns and verbs. There are four accounts of this dissociation: (i) patients may have purely lexical damage, which selectively affects verbs or nouns at a late stage of the linguistic processing (phonological or orthographic lexicons) (Rapp & Caramazza, 2002); (ii) the damage affects a lexical device, either at an ortographic-phonological modality-specific level (the lexeme; Levelt et al., 1999) or at a unitary lexical-syntactic level (the lemma) (Berndt et al., 1997); (iii) N-V dissociation arises from a semantic damage (Bird, Howard & Franklin, 2000); (iv) N-V dissociation is due to syntactic damage (Friedmann & Grodzinsky, 2000). To disentangle imageability and grammatical class effects, a new task was developed allowing to elicit nouns and verbs with identical imageability ratings in a sentence context. The results obtained will permit to address the following three questions: Does imageability play a role in determining N-V dissociation? If so, is imageability the unique cause of dissociation? If there is additional damage, at which level of linguistic processing does it take place? METHODS Twelve Italian aphasic patients and 11 normal controls participated in the study. Nouns and Verbs Retrieval in a Sentence Context (NVR-SC). Forty-five pairs of sentences denoting the same event, either using a noun or the corresponding verb (e.g. the evasion/to evade) were used. The first sentence was presented in complete form, while a gap was left in the second sentence, to be completed with the target word. For each pair of sentences two different conditions were employed, one triggering a verb and one triggering a noun. E.g. V-N: The prisoner was dreaming to evade The prisoner was dreaming the......... N-V: The prisoner was dreaming the evasion The prisoner was dreaming to......... The performance in the NVR-SC task was compared to that obtained on a classic picture naming task (50 nouns and 50 verbs). Statistical methods. Logistic regression analysis (LRA) was applied to the profiles of the patients, making it possible to study the effects of the lexical-semantic variables in univariate and multivariate linear models. RESULTS Picture naming task All patients with predominant verb deficit also have an imageability effect. In eight patients the grammatical class effect was no longer significant after introducing imageability in the statistical design (bivariate LRA). NVR-SC task Two of the verb-impaired patients in the picture naming task maintained the predominant verb deficit also in the NVR-SC task. In eight patients, the difference between nouns and verbs was no longer significant. In two patients, a paradoxical dissociation (V>N) emerged. Group analysis: performance on nouns and verbs across naming tasks (Figure 1) Patients named actions in the NVR-SC task better than in the picture naming task (58% correct versus 37%; p<.001). On the contrary, the naming of objects in the picture naming task was better than in the NVR-SC task (76% versus 61%; p<.05). DISCUSSION As expected, imageability effect is highly associated with noun-superiority. This result may have two explanations. (a) Since nouns are generally more imaginable than verbs, the imageability effect may cause noun-superiority. But, imageability alone cannot completely account for predominant verb impairment. In fact, four patients have predominant verb impairment even when imageability has been partialled out (bivariate LRA) and two patients are still dissociated in the NVR-SC task (where imageability of nouns and verbs was perfectly matched). Therefore, additional damage must be hypothesized. This would account for the fact that nouns are named better in the picture naming task, and verbs in the NVR-SC task. The former result is explained by the reduced imageability ratings; the latter may only be explained by localizing the additional damage at a central, lexical-syntactic level (i.e. the lemma). This explanation would account for the better performance on verbs in the NVR-SC task: the sentence frame and the correspondent noun may provide the patients with those information lost with the lemma damage, i.e. number of arguments and thematic roles. Nonetheless, this may account also for the fact that noun retrieval is not enhanced in the NVR-SC task: the thematic grid is not useful to retrieve nouns, since it is not a crucial aspect of noun lexical representation. In addition, aphasic patients may use a compensatory strategy to face their deficit. Since the thematic grid may be inferred from a mental image, this strategy will probably rely on visual representations of actions. Thus, the effectiveness of the compensatory strategy increases in relation to imageability (Luzzatti & Chierchia, 2002). (b) When left hemisphere language areas are completely damaged, lexical representations located in the right hemisphere emerge. This emergence is limited to high-frequency concrete nouns (Coltheart, 2000).
Idiolects are person-dependent similarities in language use. They imply that texts by one author show more similarities in language use than texts between authors. Sociolects, on the other hand, are group-dependent similarities in language use. They imply that texts by a group of authors, for instance in terms of gender or time period, share more similarities within a group than between groups. Although idiolects and sociolects are commonly used terms in the humanities, they have not been investigated a great deal from corpus and computational linguistic points of view. To test several idiolect and sociolect hypotheses a factorial combination was used of time period (Modernism, Realism), gender of author (male, female) and author (Eliot, Dickens, Woolf, Joyce) totaling 16 corresponding literary texts. In a series of corpus linguistic studies using Boolean and vector models, no conclusive evidence was found for the selected idiolect and sociolect hypotheses. In final analyses testing the semantics within each literary text, this lack of evidence was explained by the low homogeneity within a literary text.
This study investigates the writing stylechange of two Turkish authors, Çetin Altanand Yaşar Kemal, in their old and newworks using respectively their newspapercolumns and novels. The style markers are thefrequencies of word lengths in both text andvocabulary, and the rate of usage of mostfrequent words. For both authors, t-tests andlogistic regressions show that the length ofthe words in new works is significantly longerthan that of the old. The principal componentanalyses graphically illustrate the separationbetween old and new texts. The works arecorrectly categorized as old or new with 75 to100% accuracy and 92% average accuracy usingdiscriminant analysis-based cross validation. The results imply higher time gap may havepositive impact in separation andcategorization. For Altan a regressionanalysis demonstrates a decrease in averageword length as the age of his column increases. One interesting observation is that for oneword each author has similar preference changesover time.
Using role play and verbal‐report data, this study investigates the sequential organization of politeness strategies of 24 learners of Spanish and whether the learners’ ability to negotiate and mitigate a refusal was influenced by length of residence in the target community. Refusal sequences were examined throughout the interaction (head acts, pre‐ and postrefusals) and across conversational turns. Results showed more frequent attempts at negotiation and greater use of lexical and syntactic mitigation among learners who had spent more time in the target community and also revealed a preference for solidarity and indirectness, which approximated native Spanish speaker norms. It is suggested that the variables of proficiency and length of residence should be considered independently. Finally, learners’ perceptions of social status are discussed.
In this paper we introduce a naive algorithm for nondeterminisctic LTAG derivation tree extrac-tion from the Penn Treebank and the Proposi-tion Bank. This algorithm is used in the EM models of LTAG Treebank Induction reported in (Shen and Joshi, 2004). Given the trees in the Penn Treebank with PropBank tags, this algorithm generates shared structures that al-low efficient dynamic programming in the EM models. 1
Recognition and remote memory for odors, faces, and symbols were assessed in patients with pathologically confirmed Lewy body variant of Alzheimer's disease (LBV), patients with pathologically confirmed Alzheimer's disease (AD), and healthy elderly controls. On recognition memory tasks, LBV and AD patients showed significantly lower discriminability (d') than controls, particularly for olfactory stimuli. However no significant differences were found in the bias measure (c). When participants rated familiarity (a proposed measure of remote memory) of olfactory stimuli LBV and AD patients reported significantly lower familiarity than controls. Familiarity ratings were significantly lower in LBV patients than in AD patients for olfactory, but not for visual stimuli. Consistent with prior reports, the LBV patients showed significantly poorer odor thresholds than AD patients. The results suggest that recognition memory for olfactory stimuli is impaired in LBV and AD. However, patients with LBV are more impaired than patients with AD on tasks requiring remote memory for olfactory but not visual stimuli. The findings suggest that odor memory tasks may be useful in the assessment of LBV and AD.
This study analyzes and explains the variation of word and syllable final /s/ in Caracas. In particular, the analysis focuses on the phenomenon of deletion since aspiration is the norm in Caracas Spanish. The results show that phonetic deletion of /s/ in this variety of Spanish is almost non-existent in internal word position, in functional categories, in monosyllables, before non-accented vowels, and in upper and middle social class speakers ‐ contexts in which aspiration/retention is favored. On the other hand, the realization of /s/ is more likely to occur in final position, in lexical categories, in polysyllabic words, before pause, and in low social class speakers. Based on the results, a further exploration of the phenomena of change in progress and stigmatization (using style and age) as well as of the functional restrictions affecting the elision of /s/ is suggested.