Papers reviewed and determined not to be word norm studies. Use the flag icon to report errors or suggest re-inclusion.
18265 papers
In this paper we introduce a naive algorithm for nondeterminisctic LTAG derivation tree extrac-tion from the Penn Treebank and the Proposi-tion Bank. This algorithm is used in the EM models of LTAG Treebank Induction reported in (Shen and Joshi, 2004). Given the trees in the Penn Treebank with PropBank tags, this algorithm generates shared structures that al-low efficient dynamic programming in the EM models. 1
This paper presents the current status of the French treebank developed at Paris 7 (Abeillé et al., 2003a). The corpus comprises 1 million words from the newspaper le Monde, fully annotated and disambiguated for parts of speech, inflectional morphology, compounds and lemmas, and syntactic constituents. It is representative of contemporary normalized written French, and covers a variety of authors and subjects (economy, literature, politics, etc.), with extracts from newspapers ranging from 1989 to 1993. It has been used by computational linguists to train and evaluate taggers, parsers and lemmatizers, as well as by psycholinguists to extract lexical and syntactic preferences (Pynte et al., 2001). It is now being enriched with functional information, and used for parsing evaluation. 1. The French treebank Similarly to the Penn TreeBank, we have annotated both parts of speech and constituents. Differently from the Penn Treebank, we have also annotated compounds, lemmas and inflectional morphology. Our annotation choices are meant to be linguistically motivated and compatible with various linguistic theories. We have chosen surface-based annotations, with no empty categories (Abeillé and Clément,
Systemic functional linguistics offers a grammar that is semantically organised, so that salient grammatical choices are made explicit. This paper describes the explication of these choices through the conversion of the Penn Treebank into a systemic functional grammar corpus. Developing such a resource can help connect work in natural language processing to a significant body of research dealing explicitly with the issue of how lexical and grammatical selections create meaning.
Reviewed by: A rainbow of corpora: Corpus linguistics and the languages of the world ed. by Andrew Wilson, Paul Rayson, and Tony McEnery Heiko Narrog A rainbow of corpora: Corpus linguistics and the languages of the world. Ed. by Andrew Wilson, Paul Rayson, and Tony McEnery. (Linguistics edition 40.) Munich: LINCOM Europa, 2003. Pp. 165. ISBN 3895868728. $73.20 (Hb). This volume is a collection of papers that were originally presented at Corpus Linguistics 2001, a conference held 30 March–2 April 2001 at Lancaster University (UK). The sister volume to the ‘English-oriented’ Corpus linguistics by the Lune: A festschrift for Geoffrey Leech (Frankfurt: Peter Lang, 2003), it includes contributions that are concerned with non-English languages. The languages dealt with in the present volume range from the better-known Indo-European languages to Biblical Hebrew, Korean, and Arabic. Content-wise, the individual articles can be roughly divided into three categories. First, there are three diachronically oriented papers from the workshop ‘Corpus linguistics, ancient languages, and older language periods’: Beatrix Färber on a corpus of Medieval Irish (19–26), Wolf-Dieter Syring on the design and usage of a text database of Biblical Hebrew (141–52), and Matthew Brook O’Donnell, Stanley E. Porter, and Jeffrey T. Reed on the database-assisted discourse analysis of texts from the New Testament (109–21). The other papers, which come from the main session, deal with modern languages. Half of them are primarily concerned with corpora as such that is, their design, mark-up, usage, and so on. Anne Abeillé, Lionel Clément, Alexandra Kinyon, and François Toussenel discuss the PARIS 7 annotated corpus for French (1–10), and Martin Beaudoin and Michel Samard present their work on a large corpus of written Canadian French (11–18). R. Rossini Favretti, F. Tamburini, and C. de Santis deal with the CORIS corpus of written Italian (27–38), and Eva Hajičová and Petr Sgall discuss the Prague Dependency Treebank (39–50). Shereen Khoja, Roger Garside, and Gerry Knowles present a tagset for the morphosyntactic tagging of Arabic (59–72),and Kiril Simov, Gergana Popova, and Petya Osenova introduce a HPSG-based syntactic treebank of Bulgarian (133–40). Apart from language-specific interests, the papers by Abeillé and colleagues and Hajičová and Sgall are particularly impressive. The former shows how a corpus tagged with higher accuracy than previous corpora can completely overturn research results based on corpora tagged with lower accuracy. The latter presents a corpus that is marked up for syntactic features, including topic-focus structure, with a depth that is probably unmatched in any language. The rest of the papers present corpus-based linguistic research. The paper by Beom-mo Kang, Hung-gyu Kim, and Myung-hoe Huh applies Douglas Biber’s multidimensional text analysis to Korean (51–57); Maarten Lemmens investigates the functions and meaning range of posture verbs in Swedish from a typological perspective (73–85); Martina Möllering analyzes the use of the German modal particle eben in a corpus of telephone conversations (87–97); P.-O. Nilsson shows how Swedish texts translated from English exhibit systematically different [End Page 904] lexical and grammatical patterns from those found in texts written originally in Swedish (99–107); Katja Ploog explores the syntax of pronominal subjects in Abidjanee French in contrast to standard French (123–32); and Adriana Vlad, Adrian Mitrea, and Mihai Mitrea investigate Romanian texts from a stochastic perspective (153–65). Many of these contributions combine quantitative corpus methods with qualitative methods of investigation and persuasively demonstrate how corpus-based research can contribute to broader linguistic issues. On the whole, not all papers in this volume are of the same theoretical interest, but each of them is at least informative. In contrast, the editing is rather disappointing. There is a table of contents and a...
This paper deals with wordnet development tools. It presents a designed and developed system for lexical database editing, which is currently employed in many national wordnet building projects. We discuss basic features of the tool as well as more elaborate functions that facilitate linguistic work in multilingual environment.
A part-of-speech (POS) tagged corpus was built on research abstracts in biomedical domain with the Penn Treebank scheme. As consistent annotation was difficult without domain-specific knowledge we made use of the existing term annotation of the GENIA corpus. A list of frequent terms annotated in the GENIA corpus was compiled and the POS of each constituent of those terms were determined with assistance from domain specialists. The POS of the terms in the list are pre-assigned, then a tagger assigns POS to remaining words preserving the pre-assigned POS, whose results are corrected by human annotators. We also modified the PTB scheme slightly. An inter-annotator agreement tested on new 50 abstracts was 98.5%. A POS tagger trained with the annotated abstracts was tested against a gold-standard set made from the interannotator agreement. The untrained tagger had the accuracy of 83.0%. Trained with 2000 annotated abstracts the accuracy rose to 98.2%. The 2000 annotated abstracts are publicly available. 1.
This paper describes Japanese-English-Chinese aligned parallel treebank corpora of newspaper articles. They have been constructed by translating each sentence in the Penn Treebank and the Kyoto University text corpus into a corresponding natural sentence in a target language. Each sentence is translated so as to reflect its contextual information and is annotated with morphological and syntactic structures and phrasal alignment. This paper also describes the possible applications of the parallel corpus and proposes a new framework to aid in translation. In this framework, parallel translations whose source language sentence is similar to a given sentence can be semi-automatically generated. In this paper we show that the framework can be achieved by using our aligned parallel treebank corpus.
A critical issue in the study of speech motor control is the identification of the mechanisms that generate the temporal flow of serially ordered articulatory events. Two staged models of serial ordered events (Lashley, 1951; Lindblom, 1963) claim that time controls events whereas dynamic models predict a relative relation between time and space. Each of these models predicts a different relation between the acoustic measures of formant frequency and segmental duration. The most recent method described herein provides a sensitive index of speech deterioration which is both acoustically robust and phonetically systematic. Both acoustic and magnetic resonance imaging measures were used to describe the speech disturbance in two neurologically distinct groups of cerebellar ataxia: Friedreich’s ataxia and olivo-ponto cerebellar ataxia. The speaking task was designed to elicit six different prosodic conditions and four prosodic contrasts. All subjects read the same syllable embedded in a sentence, under six different prosodic conditions. Pair-wise comparisons derived from the six conditions were used to describe (1) final lengthening, (2) phrasal accent, (3) nuclear accent and (4) syllable reduction. An estimate of speech deterioration as determined by individual and normal subects’ acoustic values of syllable duration, formant and fundamental frequencies was used in correlation analyses with magnetic resonance imaging ratings.
This paper describes an incremental parsing approach where parameters are estimated using a variant of the perceptron algorithm. A beam-search algorithm is used during both training and decoding phases of the method. The perceptron approach was implemented with the same feature set as that of an existing generative model (Roark, 2001a), and experimental results show that it gives competitive performance to the generative model on parsing the Penn treebank. We demonstrate that training a perceptron model to combine with the generative model during search provides a 2.1 percent F-measure improvement over the generative model alone, to 88.8 percent.
This paper presents the preliminary analysis of Kannada WordNet and the set of relevant computational tools. Although the design has been inspired by the famous English WordNet, and to certain extent, by the Hindi WordNet, the unique features of Kannada WordNet are graded antonyms and meronymy relationships, nominal as well as verbal compoundings, complex verb constructions and efficient underlying database design (designed to handle storage and display of Kannada unicode characters). Kannada WordNet would not only add to the sparse collection of machine-readable Kannada dictionaries, but also will give new insights into the Kannada vocabulary. It provides sufficient interface for applications involved in Kannada machine translation, spell checker and semantic analyser.
OBJECTIVE: Transcutaneous electrical nerve stimulation (TENS) is a technique widely used in clinical practice to control pain, although its clinical efficacy remains controversial. Though many mechanisms have been proposed for its analgesic effects, there is a conspicuous lack of experimentally controlled research investigating whether TENS analgesia is related to its effects on the sympathetic nervous system (SNS). METHODS: Using an established psychophysiological paradigm, the present study investigated the effects of high-frequency/low-intensity TENS, low-frequency/high-intensity TENS, and sham TENS on the perception of experimental pain and SNS function in healthy volunteers. Measures of heart rate, digital pulse volume, and skin conductance were recorded during a 20-minute TENS stimulation period and in anticipation of a series of painful electric shocks prior to and following TENS stimulation. Healthy volunteers rated the intensity of the shocks using a 0-10-point verbal pain rating scale. RESULTS: The three TENS conditions failed to differentially effect SNS responses during either the 20-minute TENS treatment period or the shock anticipation periods, and TENS did not affect ratings of pain intensity to the shock stimuli. CONCLUSIONS: While these results may not generalize to acute or chronic pain patients, within the limitations of the present experimental paradigm, no support was found for TENS affecting either SNS function or acute experimental pain perception.
The purpose of this study was to analyze the performance effects of gender and regional dialect on air traffic control statement recall. Sixty-one student volunteers participated in the experiment. Thirty-one participants held a pilot’s license and 30 participants had no flight experience. Each participant listened to one CD with 60 ATC statements each representing a male and female voice and New England, Southern, and General American dialect. Participants were asked to recall exactly what they heard. If the participant could not understand what they heard, they requested a repeat. The participant’s performance was recorded to CD and analyzed. Demographic questionnaires and dialect familiarity ratings were completed and analyzed. Results showed that the best performance was with the male voice compared to the female voice. Results also showed that greater familiarity with a regional dialect will result in better performance when hearing that dialect. Although birth region was not found to have an impact on regional dialect comprehension, the regional dialect a person speaks in helps in comprehension among that dialect. Results also indicate that experience impacts dialect comprehension as the pilot group performed better across all variables than did the novice group.
The aim of this paper is to present the theoretical principles underlying the making of a lexical database of English collocations of non-specialized words used in scientific language. This project was prompted by the shortage of reference tools providing information about the use and combinatorial properties of general words in specific registers. A case study will illustrate that in scientific texts, words, especially polysemous verbs, have a distinct semantic and combinatorial behaviour. Following the assumption that the meaning and the grammatical and collocational patterns of words are interrelated, we suggest that context-specific information should be included in specialized reference tools to facilitate the written production of scientific texts by nonnative speakers ofEnglish.
Tree-based approaches to alignment model translation as a sequence of probabilistic operations transforming the syntactic parse tree of a sentence in one language into that of the other. The trees may be learned directly from parallel corpora (Wu, 1997), or provided by a parser trained on hand-annotated treebanks (Yamada and Knight, 2001). In this paper, we compare these approaches on Chinese-English and French-English datasets, and find that automatically derived trees result in better agreement with human-annotated word-level alignments for unseen test data.
In this paper we present a proposal to extend WordNet-like lexical databases by adding information about the co-occurrence of word meanings in texts. More specifically we propose to add phrasets, i.e. sets of free combinations of words which are recurrently used to express a concept (let's call them Recurrent Free Phrases). Phrasets are a useful source of information for different NLP tasks, and particularly in a multilingual environment to manage lexical gaps. At least a part of recurrent free phrases can also be represented through a new set of syntagmantic (lexical and semantic) WordNet relations.
This paper presents a Chinese treebank based decision tree approach to identify Chinese BNP. A self-learning mechanism is integrated into our model which includes the following steps: auto-extraction of POS string sequences (BNP rules) and their context information from the corpora and ID3 algorithm based tree training. Experimental results show good performances of our method.
Most statistical parsers have used the grammar induction approach, in which a stochastic grammar is induced from a treebank. An alternative approach is to induce a controller for a given parsing automaton. Such controllers may be stochastic; here, we focus on greedy controllers, which result in deterministic parsers. We use decision trees to learn the controllers. The resulting parsers are surprisingly accurate and robust, considering their speed and simplicity. They are almost as fast as current part-ofspeech taggers, and considerably more accurate than a basic unlexicalized PCFG parser. We also describe Markov parsing models, a general framework for parser modeling and control, of which the parsers reported here are a special case.
This paper discusses an annotation scheme for Korean null pronouns, which were used in annotating three kinds of Korean text corpora including Penn Korean Treebank. In annotating the corpora, null pronouns and their antecedents were marked up for their type and reference, with coreference relation tracked by numeric identifiers. Based on the annotation scheme, an outline of a potential pronoun resolution strategy is also proposed. The resulting dataset of annotated text is rather small at 11,834 words; we hope the null pronoun classification and annotation scheme proposed in this study will serve as a basis in developing a large-scale annotated corpus in the future.
Language technology makes extensive use of hierarchically annotated text and speech data.These databases are stored in flat files and manipulated using corpus-specific query tools or special-purpose scripts.While the size of these databases and the range of applications has grown rapidly in recent years, neither method for managing the data has led to reusable, scalable software.The formal properties of the query languages are not well understood.Hence established methods for indexing tree data and optimizing tree queries cannot be employed.We analyze a range of existing linguistic query languages, and adduce a set of requirements for a reusable, scalable linguistic query language.
This study analyzes and explains the variation of word and syllable final /s/ in Caracas. In particular, the analysis focuses on the phenomenon of deletion since aspiration is the norm in Caracas Spanish. The results show that phonetic deletion of /s/ in this variety of Spanish is almost non-existent in internal word position, in functional categories, in monosyllables, before non-accented vowels, and in upper and middle social class speakers ‐ contexts in which aspiration/retention is favored. On the other hand, the realization of /s/ is more likely to occur in final position, in lexical categories, in polysyllabic words, before pause, and in low social class speakers. Based on the results, a further exploration of the phenomena of change in progress and stigmatization (using style and age) as well as of the functional restrictions affecting the elision of /s/ is suggested.
Abstract The activation of arousal components for emotion-laden words in English (e.g. kiss, death) was examined in two groups of participants: English monolinguals and Spanish—English bilinguals. In Experiment 1, emotion-laden words were rated on valence and perceived arousal. These norms were used to construct prime—target word pairs that were used in Experiment 2. Monolingual and bilingual participants performed lexical decisions to English word targets in either high arousal, moderate or unrelated conditions. Results revealed positive priming effects in both arousal conditions for both groups of participants. Interestingly, while the baseline conditions were similar across groups, the arousal conditions produced longer latencies for bilinguals than for monolinguals. These data represent the first demonstration of word—word priming in the domain of perceived arousal in a dominant language for monolingual and bilingual speakers. Results are discussed in terms of the representation of emotion words and the structure of emotion in bilingual memory. Keywords: BILINGUALISMEMOTIONAROUSALAFFECTIVE PRIMINGWORD PRIMINGEMOTION-LADEN
Corpus linguistics prompts a lexicocentric approach to linguistic theory. The theory of norms and exploitations (TNE; Hanks, forthcoming) is such a theory, applying the insights of prototype theory and Sinclairian text analysis to the empirical evidence of large corpora. By studying words in context, we can identify the normal patterns of usage that are associated with each word. A meaning, or meaning potential, can then be associated with each pattern. Thereby, lexical entropy is reduced. A central question in this approach to language analysis concerns metaphors and idioms. In the present paper, conventional metaphors and idioms are classified as ‘norms’ (i.e. conventional uses), while dynamic, ad-hoc metaphors are classified as ‘exploitations’ of norms. However, conventional metaphors can still be distinguished from literal meanings. At least in some cases, conventional metaphors differ from literal senses by their particular syntagmatic patterns. The paper also discusses the importance of text type and domain in achieving a satisfactory interpretation of idiomatic expressions.
The rhetorical forms in Umbuso KaShaka (The Realm of Shaka), the Zulu translation of Nada the Lily, are analysed within the framework of Descriptive Translation Studies. The rhetorical forms investigated are individualisation, stereotyping, validation and structuring. A range of translation strategies is employed by the translator to establish the rhetorical forms. For the rhetorical form, individualisation, the strategies re-lexification, catalysis and idiosyncracy are used. Strategies like repetition, filial address, participatory response and fixed expressions are applied to set up the rhetorical form stereotyping. Lexical and semantic transfer, functional and cultural equivalents, cultural substitution and loanshift establish validation. Structuring is dealt with on grammatical (anastrophe and anacolutha) and textual level. Structuring on textual level includes components like punctuation, paragraphing, addition and omission. It is shown that the early Zulu translations are characterised by slight shifts in rhetorical form to restore their original (authentic) form, function and significance, so that they truly reflect Zulu language and culture. On the one hand, the translator was guided by a set of translator's norms relating to this period, namely to obey certain prescribed rules in order to be regarded as a good translator, i.e.to be faithful to the source text. On the other hand, however, an attempt was made to meet the expectations of the target system.
There is growing interest in the use of semantic collections in order to identify and analyse domain knowledge. This paper describes some technical issues to consider when contemplating research which incorporates small-to-medium domain-specific word sets. The purpose of the corpus construction described was to provide an external word collection which could be transformed to a numeric frequency scale which could take the place of an “expert ” in order to evaluate the lexical content of aircraft Visual Landing Approach concept maps. Although this paper is based on research in the field of aviation education, the underlying principles are more widely applicable. General Corpora Definitions and Uses The study of naturally occurring word frequencies has been a focus of computational linguistics, and in particular of the field of corpus linguistics. A word collection known as a corpus is constructed from some set of texts in order to determine what is characteristic of that text set through the identification of vocabulary patterns that either differ or conform to a norm (Ide & Walker, 1993). In contrast to Chomskyan generative linguistics, which focuses on internal knowledge of language structures (Chomsky, 1957), empirical corpus linguistics seeks to describe
Only very recently have Vietnamese re-searchers begun to be involved in the do-main of Natural Language Processing. As there does not exist any published work in formal linguistics or any recognizable standard for Vietnamese word categories, the fundamental works in Vietnamese text analysis such as part-of-speech tagging, parsing, etc. are very difficult tasks for computer scientists. All necessary linguistic resources have to be built from scratch, and until now almost no re-sources are shared in public research. The aim of our project is to build a common linguistic database that is freely and easily exploitable for the automatic processing of Vietnamese. In this paper, we propose an extensible set of Vietnamese syntactic descriptions that can be used for tagset definition and corpus annotation. These descriptors are established in such a way to be a reference set proposal for Vietnamese in the context of ISO subcommit-tee TC37/SC4 (Language Resource Management).
In this paper we describe the creation, we are carrying out of a special- ized lexicon belonging to the maritime domain (including the technical and commer- cial/maritime transport domain) and the link of this lexicon to the generic one of the ItalWordNet lexical database. The main characteristics of the lexical semantic database and the specic features of the specialized language are described together with the coding performed according to the ItalWordNet semantic relations model and the ap- proach adopted to connect the terminological database to the generic one. Some of the problems encountered and a few expected advantages are also considered.
We present the annotation o f information structure in the MULI project.To learn more about the information structuring means in prosody, syntax and discourse, theoryindependent features were defined for each level.We describe the features and illustrate them on an example sentence.To investigate the interplay o f features, the representation has to allow for inspecting all three layers at the same time.This is realised by a stand-off XM L mark-up with the word as the basic unit.The theory-neutral XM L stand-off annotation allows integrating this resource with other linguistic resources such as the Tiger Treebank for German or the Penn treebank for English.
This study tests the claim that children acquire collections of phonologically similar word forms, namely, dense neighborhoods. Age of acquisition (AoA) norms were obtained from two databases: parent report of infant and toddler production and adult self-ratings of AoA. Neighborhood density, word frequency, word length, Density×Frequency and Density×Length were analyzed as potential predictors of AoA using linear regression. Early acquired words were higher in density, higher in word frequency, and shorter in length than late acquired words. Significant interactions provided evidence that the lexical factors predicting AoA varied, depending on the type of word being learned. The implication of these findings for lexical acquisition and language learning are discussed.
The syntactically annotated corpora, commonly called ‘treebanks’, play an important role in empirical linguistics as well as in machine learning methods in natural language processing. After a brief summarization of several treebank annotation of different language, we proposed a new annotation scheme for Chinese treebank in this paper. Under this scheme, every Chinese sentence will be annotated with a complete parse tree, where each non terminal constituent is assigned with two tags. One is the syntactic constituent tag, which describes its external functional relation with other constituents in the parse tree. The other is the grammatical relation tag, which describes the internal structural relation of its sub components. These two tag sets consist of 16 and 27 tags respectively. They form an integrated annotation for the syntactic constituent in a parse tree through top down and bottom up descriptions. Based on this scheme, we built a 1,000,000 words Chinese treebank covering a balanced collection of journalistic, literary, academic, and other documents. The annotating experiments on different kinds of complex linguistic phenomena show the availability and compatibility of this annotation scheme.
М. И. ШАПИР... ЭСТЕТИКА НЕБРЕЖНОСТИ В ПОЭЗИИ ПАСТЕРНАКА (Идеология одного идиолекта)1... © 2004 г.... В первой части работы анализируются случаи непроизвольных двусмысленностей в поэзии Пастернака (лексико-фразеологических, грамматических, стилистических); во второй части они рассматриваются в ряду разного рода коллоквиализмов, нарушающих привычные нормы книжного языка и классической ритмики; наконец, в третьей части статьи делается попытка понять психологическую, эстетическую и социальную подоплеку общей установки поэта на опрощение стихотворного языка и его сближение с разговорной речью.... The first section of this article analyzes cases of involuntary ambiguity (lexical, phraseological, grammatical, and stylistic) in the poetry of Pasternak. The second section examines these cases in the context of various colloquialisms which violate the conventional norms of standard literary Russian and Russian classical prosody. The third and the last section attempts to understand the psychological, aesthetic and social background of Pasternak's general tendency towards the simplification of poetic language and its convergence with free colloquial speech.......А ты прекрасна без извилин...... Много лет назад Уильям Эмпсон, незаурядный английский поэт и филолог, расценил семантическую неопределенность (ambiguity) как неотъемлемое свойство поэзии [I]2. В последнее время интерес к поэтической неоднозначности растет и у российских лингвистов: недавно специальное исследование ей посвятил Н. В. Перцов [2] (ср. [3]). В своей книге и в предшествующих статьях я тоже обращался к этой теме (см. [4, с. 12 - 19] и др.). Но до сих пор филологи сосредоточивались главным образом на преднамеренном двоении смыслов; что же касается двусмысленностей непроизвольных (либо кажущихся таковыми), то им должного внимания не уделялось. Это упущение мне бы хотелось восполнить: сначала предметом моего анализа станет такое ветвление у Пастернака, которое, насколько можно судить, не входило в расчеты автора; затем найденные факты я попробую поставить в более широкий лингвистический и наконец - в идеологический контекст. Таким образом, против обыкновения я буду изучать не информацию, а шум, который, однако, на свой лад оказывается весьма информативным.... Размышляя над примерами, постараемся не терять из виду суть проблемы: дело не в том, что какой-то фрагмент текста не допускает верной интерпретации, - дело в том, что он объективно допускает интерпретацию неверную. Именно ощущение неадекватности вторых и третьих смыслов позволяет нам выделять оговорки среди других случаев неоднозначности. Разумеется, это ощущение может сбивать с толку: насчет авторского замысла нам дано лишь строить догадки. Наивно было бы верить, что в поэтическом тексте намеренное всегда надежно отличается от ненамеренного: в душе писателя мы читать не умеем, но попытаться его понять - обязаны3.... Явление, о котором пойдет речь, еще не имеет адекватного терминологического выражения.... 1 Исследование выполнено при поддержке Российского гуманитарного научного фонда (проект 04 - 04 - 00055а). Исправляя и дополняя исходный вариант статьи, автор имел счастливую возможность пользоваться советами и замечаниями М. В. Акимовой, С. Г. Болотова, М. Л. Гаспарова, Ф. Н. Двинятина, В. З. Демьянкова, А. А. Добрицына, И. Г. Добродомова, А. К. Жолковского, Вяч. Вс. Иванова, А. А. Илюшина, Т. М. Левиной, Т. М. Николаевой, А. Б. Пеньковского, И. А. Пильщикова, Н. В. Перцова, О. Ронена, Т. В. Цивьян. Особая признательность Е. Б. Пастернаку и Е. В. Пастернак, помогавшим автору и его поддерживавшим на протяжении всей работы.... 2 Латинское слово ambiguitas 'двусмысленность' соответствует древнегреческому (амфиболия), усвоенному русской научной терминологией.... 3 На пушкинском пленуме Союза писателей (1937) Пастернак заявил: не только намеренных двусмысленностей, но и таких провалов последнего сорта, которые бы давали повод для двусмысленного понимания и в неумышленном плане, - я за собой не помню. Вообще двусмысленности при настоящей любви к искусству немыслимы [5, т. 4, с. 644; 6, с. 401, 404 примеч. 36].... стр. 31... В арсенале испытанных средств филологического метаязыка наиболее подходящим к случаю могло бы стать понятие авторской глухоты, закрепленное в Поэтическом словаре А. П. Квятковского. Это условный термин, предложенный М. Горьким; понимаются под ним явные стилистические и смысловые ошибки..> не замеченные автором. Их можно трактовать по-разному: иногда авторская глухота - результат небрежности или неряшливости, иногда она возникает непроизвольно, когда увлечение главной задачей заслоняет отдельные детали. Явления г, - продолжает Квятковский, - свойственны не только рядовым писателям, но и большим мастерам [7, с. 10]. Он приводит примеры из Пушкина, Лермонтова, Плещеева, Фета, Маяковского, Багрицкого и Уткина. Завершается статья указанием на то, что к А г можно отнести явления сдвига, и ссылками на тематически близкие статьи: Амфиболия, Анаколуф, Солецизм [7, с.
This paper investigates the usefulness of sentence-internal prosodic cues in syntactic parsing of transcribed speech. Intuitively, prosodic cues would seem to provide much the same information in speech as punctuation does in text, so we tried to incorporate them into our parser in much the same way as punctuation is. We compared the accuracy of a statistical parser on the LDC Switchboard treebank corpus of transcribed sentence-segmented speech using various combinations of punctuation and sentence-internal prosodic information (duration, pausing, and f0 cues).
Biomimetic design uses ideas from biological phenomena as inspiration in design. To support biomimetic design, biological analogies are identified by finding instances of functional keywords that describe the engineering problem in biological knowledge in natural-language format. Challenges in using this approach include the identification of keywords, and the quantity and quality of results found. WordNet, a lexical database, is used as a language framework to systematically generate alternative keywords to find matches and analyze the results of searches. Troponyms from WordNet were found to provide better and more plentiful keywords than did synonyms. Due to the potentially large number of matches to keywords, matches are analyzed to facilitate extraction of dominant biological phenomena associated with keywords. This analysis found that words that frequently collocated with keywords tend to be objects of the keyword verb or agents that carry out the actions of the keyword. Furthermore, nouns that are inanimate, e.g., substances, tend to be objects, and nouns that are animate e.g., animals, organs, tend to be agents. Distinguishing frequently collocated words and their relationships to keywords can be used to facilitate identification of biological analogies in natural-language format to support design.Copyright © 2004 by ASME
Rating agencies' track record is good in developed countries but poor in emerging economies. Why? Given the almost-monopolistic structure of the industry, we conjecture that agencies might underinvest in information gathering. We propose an indicator quantifying the agencies' effort to gather information and assess whether greater effort affects rating levels. We detect: (i) absolute underinvestment for non-OECD sovereigns (less effort in spite of greater opaqueness); (ii) relative underinvestment for non-OECD firms compared with OECD ones (though the former receive a larger effort, more intense effort boosts firm ratings in non-OECD countries while depressing them in OECD countries).
Investigations of the Canadian quotative system have to this point focused on mainland urban varieties where General Canadian English is considered to be the linguistic norm. The current analysis seeks to expand our understanding of this system by examining quotative usage among young girls in St. John’s, Newfoundland, where the local vernacular differs in significant ways from the national variety. Variationist methodology is employed on a small corpus of St. John’s Youth English (SJYE), revealing notable similarities in the distribution of quotatives as well as the operation of internal constraints across the paradigm between this variety and that of the mainland. At the same time, there is evidence that be like, the most recent of the quotative cohort, has grammaticalized further in SJYE than in General Canadian, raising questions about the routes by which this change is progressing. The results thus situate SJYEwithin the Canadian quotative system while at the same time highlighting the status of Newfoundland English as a unique Canadian variety.
The motivation of the Papillon project is to encourage the development of freely accessible Multilingual Lexical Resources by way of online collaborative work on the Internet. For this, we developed a generic community website originally dedicated to the diffusion and the development of a particular acception based multilingual lexical database.
In author attribution studies function words or lexical measures areoften used to differentiate the authors' textual fingerprints. Thesestudies can be thought of as quantifying the texts, representing thetext with measured variables that stand for specific textual features.The resulting quantifications, while proven useful for statisticallydifferentiating among the texts, bear no resemblance to the understanding a human reader – even an astute one – would develop whilereading the texts. In this paper we present an attribution study that,instead, characterizes the texts according to the representationallanguage choices of the authors, similar to a way we believe close humanreaders come to know a text and distinguish its rhetorical purpose. Fromour automated quantification of The Federalist papers, it isclear why human readers find it impossible to distinguish the authorshipof the disputed papers. Our findings suggest that changes occur in theprocesses of rhetorical invention when undertaken in collaborativesituations. This points to a need to re-evaluate the premise ofautonomous authorship that has informed attribution studies of The Federalist case.
The current research explored the processes that predominate during the anticipation of an emotionally salient event. Experiment 1 (N536), employed three different conditional stimuli followed by pictorial pleasant, unpleasant or neutral unconditioned stimuli. Half the participants were trained with visual CSs, the other half with tactile CSs. In the group trained with visual CSs, startle eyeblinks were larger and faster during CSs that were paired with unpleasant pictures than CSs paired with neutral or pleasant pictures respectively, indicating an affect startle pattern. This linear trend was not found in the group trained with tactile CSs. Experiment 2 (N564) aimed to investigate whether the affective pattern found in the startle data in Experiment 1 could also be found using a behavioural measure of emotion. This time participants’ reaction time during a post-experimental affective priming taskwas used as dependantmeasure to assess the presence of emotional learning. Instead of a simple differential conditioning task, an occasion setting paradigm was employed and participants were trained using either a feature positive or feature negative design with pleasant or unpleasant picture USs. For participants trained with unpleasant USs, valence ratings collected before and after conditioning training suggested the presence of emotional learning, whereas no such pattern was found for participants trained with pleasant USs. These findings were not confirmed in the priming data.
Style is embodied with different meanings in the eyes of different stylists. This paper concentrates on one of the views on style,deviation.Style is deviation of the norm which helps achieve the effect of foregrounding.Some specific examples are to show that foregrounding is prominence that is motivated.