Papers reviewed and determined not to be word norm studies. Use the flag icon to report errors or suggest re-inclusion.
18265 papers
This paper proposes a new inference approach for Chinese probabilisticcontext-free grammar, which implements the EM algorithm based on the bracketmatching schemes. Two characteristics of the algorithm are as follows: 1) To pre-process the training texts with automatic constituent boundary prediction tools,which can provide stronger syntactic restriction upon training texts in lower compu-tational costs; 2) To develop an initial rule set by integrating different knowledgeresources, including a set of basic syntactic rules generated by an automatic gram-mar construction t00l and a set of special rules summarized by linguists or extractedfrom treebanks, and provide a better initialization for the learning process. There-fore, a linguistically-motivated and broad-coverage Chinese PCFG rule set can beeasily generated through this algorithm. Current experimental results prove goodlearning efficiency of this algorithm and high reliability of the generated rule set.
Accurate linguistic annotation is a core requirement of natural language processing systems. The demand for accuracy in the face of rapid prototyping constraints and numerous target languages has led to the employment of machine learning methods for developing linguistic annotation systems. The popularity of applying machine learning methods to computational linguistics problems has given rise to a large supply of trainable natural language processing systems. Most problems of interest have an array of off-the-shelf products or downloadable code implementing solutions using various techniques. In situations where these solutions are developed independently, it is observed that their errors tend to be independently distributed. In this thesis we discuss approaches for capitalizing on this situation in a sample problem domain, Penn Treebank-style parsing. The machine learning community provides us with techniques for combining outputs of classifiers, but parser output is more structured and interdependent than classifications. To overcome this, two novel strategies for combining parsers are used: learning to control a switch between parsers and constructing a hybrid parse from multiple parsers' outputs. In this thesis we give supervised and unsupervised techniques for each of these strategies as well as performance and robustness results from evaluation of the techniques. One shortcoming of combining off-the-shelf parsers is that the parsers are not developed with the intention to perform well on complementary data or to compensate for each others' weaknesses. The individual parsers are globally optimized. We present two techniques for producing an ensemble of parsers in such a way that their outputs can be constructively combined. All of the ensemble members will be created using the same underlying parser induction algorithm, and the method for producing complementary parsers is only loosely coupled to that algorithm.
Word recognition and generation is a fundamental part of the processing of natural language and it requires computationally effective morphological processors, especially for languages with rich morphology such as Modern Greek. Various models have been proposed for developing computerized systems to accomplish the task of recognition of morphosyntactic features of words In the work presented here, the lazy tagging approach was examined, in which taggers are expected to work in the simplest possible way. The model of functional decomposition was extended and adapted for Modern Greek as a target language, following the lazy word-parsing approach, in order to cover a number of morphological phenomena that are encountered in Modern Greek, namely inflection, affixation, and longdistance dependencies. To achieve a more efficient word recognition, several automata of different levels of computing power, based on the original model, were introduced and evaluated according to the criteria of complexity, recognition speed, and accuracy of the results. The proposed system was used for processing a large-scale corpus, and the results are presented and discussed. To accomplish their task, taggers can rely upon large lexical databases, which are expected to be organized in such a way as to provide rapid access to the stored data and efficient memory management. Directed graphs can be used to describe and organize a lexical database of large magnitude in a compact manner. These data structures are named here matrix lexica, where the letters are described as nodes of directed graphs and the lemmata as paths (set of edges). It is expected that matrix lexica will support a tagger efficiently by providing a high speed of resolution, sound mathematical foundation, low memory requirements, and ability to handle distorted input in future developments.
Zahlreiche neuere Arbeiten für das Englische zeigen, daß statistische Analysen großer Korpora und Treebanks gute Heuristiken für die Zuordnung von Präpositionalphrasen liefern können. Entsprechende Untersuchungen für das Deutsche scheitern bisher an den fehlenden Daten. Wir zeigen jedoch, daß durch Einbeziehung weiterer Faktoren auch für das Deutsche mit guten Ergebnissen zu rechnen ist. Betrachtet werden der Einfluß unterschiedlicher Gewichte für Verben und Nomina, die Auswirkungen einer vorgeschalteten lexikalischen Disambiguierung sowie die Kopplung lexikalischer und grammatischer Präferenzen. Recent proposals have shown that statistical analyses of large English corpora and treebanks provide good heuristics for the attachment of prepositional phrases. Similar proposals for German have failed since such resources have not been available. We show that by using some additional factors we can achieve similar results for German. We demonstrate the influence of different weights for verbs and nouns, the influence of lexical disambiguation and the combination of lexical and grammatical preferences.
Determining the attachments of prepositions and subordinate conjunctions is a key problem in parsing natural language. This paper presents a trainable approach to making these attachments through transformation sequences and error-driven learning. Our approach is broad coverage, and accounts for roughly three times the attachment cases that have previously been handled by corpus-based techniques. In addition, our approach is based on a simplified model of syntax that is more consistent with the practice in current state-of-the-art language processing systems. This paper sketches syntactic and algorithmic details, and presents experimental results on data sets derived from the Penn Treebank. We obtain an attachment accuracy of 75.4% for the general case, the first such corpus-based result to be reported. For the restricted cases previously studied with corpusbased methods, our approach yields an accuracy comparable to current work (83.1%).
We argue that the current dominant paradigm in parser evaluation work, which combines use of the Penn Treebank reference corpus and of the Parseval scoring metrics, is not well-suited to the task of general comparative evaluation of diverse parsing systems. We propose an alternative approach which has two key components. Firstly, we propose parsed corpora for testing that are much flatter than those currently used, whose "gold standard" parses encode only those grammatical constituents upon which there is broad agreement across a range of grammatical theories. Secondly, we propose modified evaluation metrics that require parser outputs to be `faithful to', rather than mimic, the broadly agreed structure encoded in the flatter gold standard analyses. 1. Introduction Interest in the evaluation of language technology has grown immensely in the past few years. This interest varies depending on the perspective one has on the technology: users and suppliers want to know how accurate, usabl...
A series of four experiments were conducted to examine viewer perceptions of three sets of five nonrepresentational paintings. Increased complexity was embedded in the hierarchical structure of each set by carefully selecting colors and ordering them in each successive painting according to certain rules of transformation which created hierarchies. Experiment 1 supported the hypothesis that subjects would discern the hierarchical complexity underlying the sets of paintings. In Experiment 2 viewers rated the paintings on collative (complexity, disorder) and affective (pleasing, interesting, tension, and power) scales, and a factor analysis revealed that affective ratings were tied to complexity (Factor 1) but not to disorder (Factor 2). In Experiment 3, a measure of exploratory activity (free looking time) was correlated with complexity (Factor 1) but not with disorder (Factor 2). Multidimensional scaling was used in Experiment 4 to examine perceptions of the paintings seen in pairs. Dimension 1 contrasted Soft with Hard-Edged paintings, while Dimension 2 reflected the relative separation of figure from ground in these paintings. Together these results show that untrained viewers can discern hierarchical complexity in paintings and that this quality stimulates affective responses and exploratory activity.
In comparison to conventional displays, 3D stereoscopic displays convey additional information about the 3D structure of a scene by providing information that can be used to extract depth. In the present study we evaluated the psychovisual impact of stereoscopic images on viewers. Thirty-three non-expert viewers rated sensation of depth, perceived sharpness, subjective image quality, and relative preference for stereoscopic over non-stereoscopic images. Rating methods were based on procedures described in ITU- Rec. 500. Viewers also rated sequences in which the left- and right-eye images were processed independently, using a generic MPEG-2 codec, at bit-rates of 6, 3, and 1 Mbits/s. The main finding was that viewers preferred the stereoscopic version over the non-stereoscopic version of the sequences, provided that the sequence did not contain noticeable stereo artifacts, such as exaggerated disparity. Perceived depth was rated greater for stereoscopic than for non-stereoscopic sequences, and perceived sharpness of stereoscopic sequences was rated the same or lower compared to non-stereoscopic sequences. Subjective image quality was influenced primarily by apparent sharpness of the video sequences, and less so by perceived depth.
Historically, research on the construct of body image has focused on its stability. Many researchers are beginning to reexamine whether the body image construct is stable, and they have shown that the construct is subject to change after experimentally induced situations and after major life events. This study attempted to determine whether minor life events and mood had a significant relationship to body image ratings and whether a change in minor life events and mood over the course of one month would predict body image ratings. For men, it was found that minor life events were not significantly related to body image ratings, though higher mood scores were significantly related to lower ratings of physical appearance. For women, a greater number of positive minor life events was significantly related to engaging in more behaviors to keep oneself physically attractive, and higher mood scores were significantly related to lower ratings of physical appearance. For men, changes in minor life events or mood over the course of one month did not predict change in body image ratings. For women, an increase in positive minor life events predicted an increase in behaviors associated with keeping oneself physically attractive. A post-hoc analysis was conducted to determine whether individuals whose mood worsened over the course of one month would show greater changes in body image ratings. However, this post-hoc hypothesis was not supported. The main hypotheses were reanalyzed with the subsample stratified into younger and older adult men and women. Though the sample size was small, there appeared to be differences between older and younger adults, with younger adults more susceptible to body image fluctuations than older adults. In the overall sample, body image ratings changed little over the course of one month, though this discovery fits well within an overall personality contruct model proposed by Mischel (1968). Effect sizes for this study were small, and the sample size was too small to confidently find significant relationships or make predictions. Other limitations of this study as well as future directions in research are discussed.
BOOK NOTICES 207 sure's concept ofmotivation should not be associated with words but rather with cotext and context. Everything is relative in language, and the systematic character of language does not lie in separate phonemic, lexical, grammatical, and textual systems. In fact the systematic and universal features of language are made up by the human capabilities of thinking and experiencing. Speakers are able to use a limited number of signs to express highly complicated ideas, and they can decode expressions in very complex situations. To do this requires applying the basic principle of language and language description—the idea of economy. This elementary principle of human behavior can be seen when a speaker tries to avoid unnecessary redundancy and when simpler ways of pronunciation are preferred to more difficult ones. In addition language economy can be seen on the deeper level of reinterpreting linguistic units new to a certain speaker. This gives sense to an utterance since under the principle ofeconomy, a speaker must assume that no text is uttered without meaning. As a result of new expressions and new interpretations produced by the principle ofeconomic use, language might change over time. Considering language change again, D emphasizes how the individual reflects about language. For the individual the main goal of language and speaking is to impart and to decode sense or meaning. D rejects the idea of the 'invisible hand phenomenon' as well as the concept of teleology in language change. Again he links the systematic character of language to the speaker's purposeful acts rather than to the whole speech community or to language as an abstract system. In short, D questions the traditional structuralist conception of system in language. He emphasizes the systematic cognitive behavior of each individual that uses the relative means of language. In order to support his opinion D illustrates his ideas with many detailed examples mainly taken from German, English, and French. [Dieter Aichele, Fachhochschule Neubrandenburg.] The Oxford English-Hebrew Dictionary. Ed. by N. S. Doniach and A. Kahane. Oxford: Oxford University Press, 1996. Pp. xxiii, 1091. The late N. S. Doniach (d. 16 April 1994), the chief editor ofthis dictionary, is well known to Semitologists for editing the excellent Oxford EnglishArabic dictionary ofcurrent usage (1972). The introduction by Professor A. Kahane explains D's goal of using various styles of modern Hebrew in this volume, including colloquial language and slang. From abacus to Zulu, this dictionary, happily, has it all! It even has the f-word with many of its most common idiomatic usages, such as '__ up' and '__ off!' (351). The tome's particularly noteworthy features include up-to-date terminology ofall sorts, such as that dealing with computers. However, some inconsistencies can be found. 'Software' is written as toxna with a vav (883), but it is spelled without a vav (using a kamats katan) under 'hardware' (400). And curiously, the name of the vowel kamats is written kamatz, yet the vowel chataf kamats is spelled differently on the very next line (ix). AU Hebrew words are given in their fully vocalized or pointed forms, including dagesh, which is said to have the phonetic value 'stress mark', a puzzling statement (ix). Phonological matters on the whole, however, have been handled well. It was a wise decision for the editors to give preference to a word's modern pronunciation if it differs from its traditional pointing (xxiii). Along these lines, I wanted to check the vocalization and pronunciation of the irregular Classical Hebrew plural for bayit 'house'—battiim or (bottiim); to my surprise, however, it was not given (426). British English has been chosen as the norm of this dictionary, a reasonable and expected decision by Oxford University Press. An American has little difficulty getting used to British spellings such as 'programme'; however, 'program' is also listed, but with the stipulation, 'US and Comput.' (724). However, it may be difficult at first for an American to appreciate 'farther' and 'father' transcribed exactly the same. There are a few discrepancies to report. American English 'buggy' is given as eglat-tinok (114), but 'pram' lists only eglat-tinokot, its plural (708), whereas the more formal 'perambulator' (called 'formal ' by the editors) lists...
The kinds of tree representations used in a treebank corpus can have a dramatic effect on performance of a parser based on the PCFG estimated from that corpus, causing the estimated likelihood of a...
Forty-nine young adults (M age = 22 years) and 30 elderly adults (M age = 69 years) rated the 60 pictorial stimuli from the Boston Naming Test (BNT) on familiarity, providing the first such normative data for these stimuli along this dimension. Participants also made speeded lexical decisions about the word item representations of each BNT picture. B NT word frequency values were also examined in relation to BNT familiarity and speeded lexical decision performance. For both young and elderly adults, lexical decision reaction times to the word representations of BNT stimuli were negatively related to word frequency and familiarity of the BNT pictures. These patterns suggest that increases in word frequency and picture familiarity facilitate (i. e., speed up) the processing of BNT word representations. Furthermore, speed of processing appears to be a relevant dimension of BNT performance, at least when young and elderly adults free from clinical aphasia are involved.
`Linguistic annotation' is a term covering any transcription, translation or annotation of textual data or recorded linguistic signals. While there are several ongoing efforts to provide formats and tools for such annotations and to publish annotated linguistic databases, the lack of widely accepted standards is becoming a critical problem. Proposed standards, to the extent they exist, have focussed on file formats. This paper focuses instead on the logical structure of linguistic annotations. We survey a wide variety of annotation formats and demonstrate a common conceptual core. This provides the foundation for an algebraic framework which encompasses the representation, archiving and query of linguistic annotations, while remaining consistent with many alternative file formats. 1. INTRODUCTION `Linguistic annotation' is a cover term for any orthographic, phonetic or prosodic transcription; any speech, part-of-speech, disfluency or gestural annotation; and any free or word-level tr...
We investigated the feasibility of a computer-graphics-based method of assessing stereomotion thresholds (Silicon Graphics Stereoview stereoscopic system). Stereomotion thresholds for a rectangle oscillating in depth were determined with the use of a dual randomly interleaved staircase design. In a group of 31 naive observers, the average thresholds of 5.97′ of arc forcrossed stereomotion and 6.00′ of arc foruncrossed stereomotion were comparable to those assessed in earlier work done with optics-based techniques. By assessing the thresholds for a rectangle that was defined either by lateral motion or by changing size, in a group of experienced observers, we were able to show that any potential residual translational motion present in the display would not have influenced the stereomotion thresholds. Our findings suggest that this computer-graphics-based technique may be a reasonable alternative to optics-based methods of assessing stereomotion thresholds.
This paper is about two aspects of subcategorisation in NLP. First, it is about the automatic extraction of subcategorisation information from corpora. More specifically, we are concerned with unsupervised learning of subcategorisation information from tagged text by means of hierarchical clustering. The second aspect of the paper is the usage of this subcategorisation information for parsing, especially for the distinction between complements and adjuncts. We show that the information learned by unsupervised clustering can be exploited by a memory-based learner, to improve upon the complement-adjunct distinction. We compare the improvement gained by the use of this unsupervised information (1%) to that of different representations of subcategorisation information extracted from the tree-bank annotation (maximum 1.5%). The unsupervised information thus achieves two thirds of the improvement that can be obtained from the hand-crafted treebank information. 1 1 Introduction Subcategoris...
194 LANGUAGE, VOLUME 74, NUMBER 1 (1998) erature in dialects and minority languages: Max and Moritz in Scots' (220-45). Many of the articles display the concerns of applied linguistics or language planning. For example, one goal proposed in the first article (37) is the description of regional English dialects and creóles so that teachers may be able to discern errors in second language (or dialect) acquisition ofthe standard variety from local usage. The third article is a detailed examination of word-formation processes in a range of English varieties. Article 5 is a fairly exhaustive report of dictionaries of world Englishes. Article 6 is a good look at 135 Irish-derived lexical items in ten dictionaries of English (mostly American and British); the author also makes an excellent case for the creation ofa dictionary ofIrish English (186-90). The final article brings up many relevant points regarding interlanguage and dialect translations of texts. There are relatively few typographical errors in the book. However, several points of confusion still manage to arise. For example, one cannot speak of the USA as such until after 1776; before this date, this North American territory was one of several English colonies (cf. p. 15). Also, I am unable to understand how UsE (United States English) (57) and AmE (American English) (59) are different from each other; that is, these two varieties, which are referenced within the same article, appear to have the same referent. Regarding usage, the author employs the terms 'Irishmen' (188) and 'Englishmen' (226) as if they were synonyms for all speakers of a language or for persons associated with a specific ethnicity or geographical territory. Furthermore, I am unclear why some varieties of English are labeled 'deficient' (124). Also, what does 'dictionary-worthy ' (51) mean precisely? And why are innovations in Tok Pisin labeled as 'clumsy paraphrases' (fn. 19, p. 53)? It seems to me that this book would have been better served if such expressions had been removed or at least made clear. The collection is indexed according to name and topic (269-76), which is a welcome feature. The references are divided between two subsections, called 'Dictionaries' (246-52) and 'General' (253-68). However, the usefulness of such a distinction is lost on me (except that the former subsection is an excellent resource ofpublished dictionaries ofEnglish varieties ). For example, I had to look up many references twice because I was uncertain under which heading a particular entry fell. G writes (33), the 'linguist is certainly not some kind of language referee who can make grammatical deviance from a foreign norm acceptable, i.e. turn the "mistakes" of a prescriptive tradition into permissible alternatives on the grounds that they are the consequences of necessary adaptation'. I know that as a researcher gathering data on an English-derived creóle variety which native speakers themselves often referred to as 'di bad inglish', I sometimes felt this role of referee had been uncomfortably thrust upon me by the speech community. However, the differences between standard varieties ofEnglish and regional varieties (creóle or otherwise) may perhaps best be seen as adaptations and innovations within the context oflanguage change. Thus, from this point of 'mistakes' or adaptations, speakers may begin sowing the seeds of natural and common diachronic change. [Michael Aceto, University of Puerto Rico.] Prague linguistic circle papers. Travaux du cercle linguistique de Prague, new series. Vol. 1. Ed. by Eva Hajicová, Miroslav Cervenka, Oldrich Leska, and Petr Sgall. Amsterdam & Philadelphia: John Benjamins, 1995. Pp. x, 336. Vol. 2. Ed. by Eva Hajicová, Oldrich Leska, Petr Sgall, and Zdena Skoumalov á. Amsterdam & Philadelphia: John Benjamins, 1996. Pp. viii, 346. These are the two initial volumes of the new, third series of the Prague Travaux. The first series, eight volumes of Travaux du cercle linguistique de Prague (1929-1939), was brought to an end by World War II; the second, Travaux linguistiques de Prague (1964-1971), encompassed only four volumes when it was strangled by the political authorities. The list of contributors to this new series is quite international: Besides 21 from the Czech Republic, there are 16 from other countries—Austria, France, Germany, Israel, Macedonia, the Netherlands, Slovenia, Switzerland, Russia, United...
Machine learning techniques can be used to make lexicons adaptive. The main problems in adaptation are the addition of lexical material to an existing lexical database, and the recomputation of sublanguage-dependent lexical information when porting the lexicon to a new domain or application. Inductive lexicons combine available lexical information and corpus data to alleviate these tasks. In this paper, we introduce the general methodology for the construction of inductive lexicons, and discuss empirical results on a case study using the approach: prediction of the gender of nouns in Dutch. 1. Introduction In computational lexicography, lexicons of language engineering applications should come with acceptable lexical coverage, and with the information necessary for the intended applications. They should also come equipped with methods for the automatic extension and adaptation of the lexicon with new or modified lexical entries. Computational lexicology should therefore try to solve t...
Studied implementational choices in pronunciation by analogy (PbA) for English to assess its ultimate suitability - both as a model of the human process of reading aloud, and as a component of a text-to-speech (TTS) system. The variables studied were the specific lexical database used as the basis of the analogy process, the way of ranking/scoring candidate pronunciations, and the effect of manual vs automatic alignment of letters and phonemes. When tested with short (monosyllabic) pseudowords, the lowest error rate achieved was 14.3%. This suggests that the current PbA systems are at best poor models of pseudoword pronunciation by humans. When tested with lexical words temporarily removed from the dictionary, the best performance obtained was 93.5% phonemes correct for a 16,280-word dictionary. This was superior to the 25.7% words correct obtained using a set of popular letter-to-sound rules, indicating considerable scope for analogy methods to be exploited in future TTS systems. (PsycINFO Database Record (c) 2016 APA, all rights reserved)
Nonparametric regression techniques, which estimate functions directly from noisy data rather than relying on specific parametric models, now play a central role in statistical analysis. We can improve the efficiency and other aspects of a nonparametric curve estimate by using prior knowledge about general features of the curve in the smoothing process. Spline smoothing is extended in this paper to express this prior knowledge in the form of a linear differential operator that annihilates a specified parametric model for the data. Roughness in the fitted function is defined in terms of the integrated square of this operator applied to the fitted function. A fastO(n) algorithm is outlined for this smart smoothing process. Illustrations are provided of where this technique proves useful.
Automatic text categorization is a complex and useful task for many natural language processing applications. Recent approaches to text categorization focus more on algorithms than on resources involved in this operation. In contrast to this trend, we present an approach based on the integration of widely available resources as lexical databases and training collections to overcome current limitations of the task. Our approach makes use of WordNet synonymy information to increase evidence for bad trained categories. When testing a direct categorization, a WordNet based one, a training algorithm, and our integrated approach, the latter exhibits a better perfomance than any of the others. Incidentally, WordNet based approach perfomance is comparable with the training approach one.
After presenting a novel O(n^3) parsing algorithm for dependency grammar, we develop three contrasting ways to stochasticize it. We propose (a) a lexical affinity model where words struggle to modify each other, (b) a sense tagging model where words fluctuate randomly in their selectional preferences, and (c) a generative model where the speaker fleshes out each word's syntactic and conceptual structure without regard to the implications for the hearer. We also give preliminary empirical results from evaluating the three models' parsing performance on annotated Wall Street Journal training text (derived from the Penn Treebank). In these results, the generative (i.e., top-down) model performs significantly better than the others, and does about equally well at assigning part-of-speech tags.
Whether we approve or disapprove of the style or the content of what writes in A Funerall Elegye has nothing to do with the issue of its authorship. What we believe authorship to be, however, has everything to do with it. Attribution research asks for an understanding of how an author creates, what an author leaves of himself in the work, and how that differs from the linguistic system belonging to the period. The study of authoring and authorial idiolects is now interdisciplinary and empirical. It rests on the testimony of authors and, over the past half century, on repeated experiments by neuroscientists, cognitive psychologists, linguists, and other disinterested observers on the process of uttering sentences. Once we know how the mind shapes language into speech or writing, we will be in a good position to understand how Shakespeare did so. Attribution evidence takes three forms. It can be external, found on the title page or the author's preface, and here resting in the historical events described in the poem. It can be interpretive, stemming from a reading of the meaning of the text or of the author's style. Or it can be linguistic, extracting substylistic characteristics of the writing in the hope they may distinguish the writer from any other writer: that is, a fingerprint. The external evidence for Shakespeare as is good but not incontrovertible, despite best efforts by Foster, Abrams, and others in the debate. Someone with Shakespeare's initials published an elegy in 1612 with the printer who published his sonnets in 1609. The poet in his preface says that he personally knew the subject of the poem, an Oxford graduate murdered in Exeter. Because several poets with these initials are active at this time, and none has a very strong link to the deceased young man, the identification of author remains open. Even if Shakespeare's full name were on the title page, doubt would dog this attribution. Anyone can claim authorship of a work falsely; and anyone can incorrectly attribute a work and publish it. Only the author knows for sure, and Shakespeare never left a list of his works in his own hand. Printers sold books supposedly by Shakespeare that are not. Even his friends could innocently pass off, as his, many scenes by another playwright. Henry VIII belongs to both Shakespeare and Fletcher, not just to Shakespeare, as Heminges and Condell lead us to believe. To assign authorship on the basis only of testimony by others is to judge on circumstancial evidence, which routinely leaves juries deeply worried. Abrams's reading of the poem plausibly argues that several passages self-identify the poet as an actor-playwright. Other passages contain word clusters from Shakespeare's known works, one of them then unpublished. Because the only practicing playwright in 1612 known to have the initials is Shakespeare, Abrams reasonably concludes that the authorship question can be answered. On the other hand, Foster argues that Thorpe the printer misidentified Shakespeare as W. H. in the preface to his sonnets. Critics who disagree with Foster and Abrams may say that W. S. was reversed or, less plausibly, that someone misread f for long s in the manuscript. Further, verbal parallels between Shakespeare's work and that of others are well known in scholarship. They appear even in arguments that Shakespeare did not write the works known certainly as his. Source studies use such parallel passages. The main difficulty with them relates to our fragmentary knowledge of the state of common English idiom in the period. How rare is the fixed phrase or collocation (word pair in variable order) or word cluster (combination of fixed phrase and collocation)? Foster uses a large Renaissance lexical database, Shaxicon, to establish that the phrase Court opinion appears only in A Funerall Elegye and in a late redaction of Shakespeare's lost play, Cardenio. Abrams cites a choice parallel between the poem and Richard II, but he does not say that he has searched the literature of the period to see if the cluster is elsewhere. …
The increasing availability of corpora annotated for linguistic structure prompts the question: if we have the same texts, annotated for phrase structure under two different schemes, to what extent do the annotations agree on structuring within the text? We suggest the term tree alignment to indicate the situation where two markup schemes choose to bracket off the same text elements. We propose a general method for determining agreement between two analyses. We then describe an efficient implementation, which is also modular in that the core of the implementation can be reused regardless of the format of markup used in the corpora. The output of the implementation on the Susanne and Penn treebank corpora is discussed.
Large tree databases as knowledge repositories become more and more important; a prominent example are the treebanks in computational linguistics: text corpora consisting of up to five million words tagged with syntactic information. Consequently, these large amounts of structured data pose the problem of fast tree retrieval: Given a database T of labeled multiway trees and a query tree q, find efficiently all trees t ∈ T that contain q as subtree. This paper presents a generalization of the classical n-gram indexing technique for supporting fast retrieval of multiway tree structures: Treegram indexing covers database trees with subtrees of fixed height; each entry of the resulting index represents such a subtree together with the database trees that contain this subtree. The evaluation of a given query q preselects those database trees that contain all of q ’s cover trees and, in turn, tests these candidates rigorously for containment of q. As an application of treegram indexing, we describe the VENONA retrieval system, which handles the BH t treebank containing 508,650 phrase structure trees found in the morphosyntactical analysis of The Old Testament with altogether 3.3 million wordforms—results of a computational-linguistics project at the Ludwig-Maximilian’s University of Munich.
In a study that used Hermans's (1987, 1988) valuation procedure, 40 participants each provided a highly valued experience of four types: science, religion, interpersonal and intrapersonal conflict between science and religion. They then rated these valuations on 30 affect terms, some of rich were later organized into categories: Positive, Negative, Self, and Other. Participants also filled out questionnaires, that were used to categorize them as low or high in scientific and religious orientation. Typical valuations of participants in these four science and religion categories are presented as qualitative idiographic information. In addition, quantitative analyses of affect ratings are presented as nomothetic information. Generally, affect ratings of scienfific experiences were more Self-directed while religious experience valuations involved equally high levels of Self and Other affect. Both scientific and religious experiences were evaluated as having Positive but not Negative affect. Interpersonal and intropersonal conflict were experienced as more Self-directed than Other-oriented. While interpersonal conflicts displayed more Negative affect than Positive, intrapersonal conflict was evaluated equally on these two measures.
We present an efficient algorithm for retrieving from a database of trees, all trees that differ from a given query tree by a small number additional or missing leaves, or leaf label changes. It has natural language processing applications in searching for matches in example-based translation systems, and retrieval from lexical databases containing entries of complex feature structures. For large randomly generated synthetic tree databases (some having tens of thousands of trees), and on databases constructed from Wall Street Journal treebank, it can retrieve for trees with a small error, in a matter of tenths of a second to about a second.
This paper reports on an empirically based system that automatically resolves VP ellipsis in the 644 examples identified in the parsed Penn Treebank. The results reported here represent the first s...
This study examined signs of mania on the Rorschach, specifically whether manic inpatients (n = 24) produce different thematic content and thought disorder than comparison groups of paranoid schizophrenic (n = 27) and schizoaffective (n = 25) inpatients. Rorschach protocols were scored by a trained rater for the Thought Disorder Index and the Schizoid-Affective Rating Scale. Results indicated that all 3 groups had moderate levels of thought disorder, but the manic inpatients produced significantly more combinatory thinking and affective content responses than the other 2 groups. The paranoid schizophrenic and schizoaffective patients did not produce significantly more schizoid content and were not different on any other types of thought disorder than the manic patients. These findings are discussed in terms of the contribution of thought disorder and affective thematic content in making the diagnosis of mania on the Rorschach.
In two studies, pedestrians in Old and New Delhi (India) and Dhaka (Bangladesh) were asked about their reactions to three stressors common to rapidly growing urban areas in South Asia: noise, air pollution, and crowding. Results from the first study, a survey of men in Old Delhi, indicated that respondents who were more upset by noise and by crowding also reported more physical symptoms and less perceived control. In the second study, male and female pedestrians were interviewed in New Delhi and Dhaka. Results revealed consistent gender, country, and gender by country effects on measures of general affect, ratings of stressors, and coping responses. In addition, results from an experimental manipulation in Study 2 indicated that in both countries, telling pedestrians about the effects of air pollution or crowding made them feel significantly worse than they would have felt had they not been given any information.
Mexican Spanish is often described as a conservative variety whose distinguishing features go back only to the 19th century. This dissertation proposes that the origins of Mexican Spanish, at least with regard to verbal paradigm reduction, may be traced to the first century of the colonial period. My analysis is based on the collection of documents compiled in Documentos linguisticos de la Nueva Espana: Altiplano Central by Concepcion Company. The documents in this book are non-literary sources written during the colonial period. The following subjects were studied in this dissertation: language policy and Castilianization during the colonial period in Mexico, forms of address and the elimination of the second person plural, the -ra form and its function as an imperfect subjunctive and, finally, a historical analysis of the temporal values of the present perfect. It was found that the Mexican Spanish verbal paradigm, since the first century of the colonial period, has been characterized by the reduction of plural forms and a tendency to prefer the ending -ra form for past subjunctive, and a temporal value of an open past for the present perfect. However, there were two specific periods of time when this tendency was broken, at the beginning of the 17th and of the 19th century. At the same time, these two periods were characterized by a reassessment of the European values and imposition of its linguistic norms. This study confirms that, during the colonial period Mexican Spanish was caught between two strong opposing tendencies. One of these tendencies is the process of linguistic simplification, common in colonial territories; and the other is the process of monocentric standardization, by means of which the linguistic norm of Spain exerted pressure in the colonies.
Although past research suggests that the spatial diffusion of linguistic features across a landscape is a simple and clear-cut process, our research in Oklahoma suggests otherwise. We collected data for this study in a statewide, multifaceted investigation of grammatical, lexical, and phonological variation in Oklahoma and analyzed it using several cartographic and statistical procedures. Our purpose was to uncover some of the diffusion processes tied to language that are at work in Oklahoma. We used the General Linear Model (GLM), a multivariate statistical procedure, to identify barriers and amplifiers that influence the geographic distributions of linguistic features. The results suggest that linguistic diffusion in Oklahoma happens in a hierarchical pattern in some cases; in others, the spread is contra-hierarchical with innovations expanding up, rather than down, the urban hierarchy. A correlation of diffusion patterns with social factors that serve as barriers to, or amplifiers of, the diffusional process suggests that different patterns of diffusion may be tied to the different social meanings that linguistic features carry. In the data examined, those innovations that diffuse hierarchically represent the encroachment of external norms into an area, while those features that diffuse in a contra- hierarchical fashion represent the revitalization of traditional norms. ©1997 Oklahoma Academy of Science
The Berber lexicon today is affected by two major factors. First, there is the historically strong influence of Arabic; and second, there is irregular individual neologism, in the sense that there exist many newly coined words or phrases that usually disobey the Berber linguistic norms and have not yet received general public acceptance. The desire to purify the language by eliminating the foreign loans and supplanting them by newly coined erratic terms may do more harm than good to the Berber lexicon. The main cause of this wide-ranging borrowing is Moroccan Arabic/ Berber bilingualism, which characterizes the speech of Berberophones. Moroccan Arabic is spoken or at least understood by most Berber Speakers across Central Morocco. Moroccan Arabic, however, is used by Berberophones out of necessity, especially in formal settings or when the Speakers are in contact with local authorities of public administration personnel This sociolinguistic Situation is due to the large linguistic and cultural impact of urban centres over rural areas, and to the fact that literacy is achieved through Arabic in schools. This kind of bilingualism has engendered a regression of Berber as a language of communication, but it does not imply that Berber isfaced with imminent death.
A lexical modeling methodology was employed to examine how the distribution of phonemic patterns in the lexicon constrains lexical equivalence under conditions of reduced phonetic distinctiveness experienced by speechreaders. The technique involved selection of a phonemically transcribed machine-readable lexical database; definition of transcription rules based on measures of phonetic similarity; application of the transcription rules to a lexical database and formation of lexical equivalence classes; and computation of 3 metrics to examine the transcribed lexicon. The metric percent words unique demonstrated that distribution of words in the language preserves lexical uniqueness across a wide range in the number of potentially available phonemic distinctions. Expected class size demonstrated that if at least 12 phonemic equivalence classes were available, any given word would be highly similar to only a few other words. Percent information extracted provided evidence that high-frequency words tend not to reside in the same lexical equivalence classes as other high-frequency words. The steepness of the functions obtained for each metric shows that small increments in the number of visually perceptible phonemic distinctions can result in substantial changes in lexical uniqueness. (PsycINFO Database Record (c) 2016 APA, all rights reserved)
Queer theory offers insights for political economy on how humans induce categories and conflate traits in ways psychologists call "illusory correlations." A Bayesian simulation is constructed of people interacting and using probits to compare their rankings of alternatives to estimate the subjective probability i.) that others have the same tastes as them, and ii.), for each alternative, that this alternative is their best choice. These simulations are found to, at least simplistically, resemble a type of illusory correlation which gained increased prominence in the United States from 1930 to 1960, and earlier in the United Kingdom when queer panics conflated "predatory" and "traitorous" with lesbian/gay. This modeling of the social articulation of preferences leads to conjectures on the role of Michel Foucault's épistémè (1972), Barbara Ponse's principle of consistency (1978), Jeffrey Escoffier's master code (1985), Sandra Bem's schema (1981), John R. Searle's Background (1990, 1992, 1995), and Judith Butler's linguistic norms (1993). Here these are called cognitive dispositions or codes, and are seen as social structures which grow as individuals try to form homopreference networks to process information in parallel, collectively. The concept of identity or ideology entrepreneurs is used to establish the importance of institutional analysis for political economy. Such an incorporation of desire on a par with logic-"rationality"-is called post/modern and is used to overcome the silence in both neoMarxian and neoclassical political economy on queer theory and queer issues.
The statement, ’’Results of most non-traditional authorship attribution studies are not universally accepted as definitive,'' is explicated. A variety of problems in these studies are listed and discussed: studies governed by expediency; a lack of competent research; flawed statistical techniques; corrupted primary data; lack of expertise in allied fields; a dilettantish approach; inadequate treatment of errors. Various solutions are suggested: construct a correct and complete experimental design; educate the practitioners; study style in its totality; identify and educate the gatekeepers; develop a complete theoretical framework; form an association of practitioners.
The increasing availability of corpora annotated for linguistic structure prompts the question: if we have the same texts, annotated for phrase structure under two different schemes, to what extent do the annotations agree on structuring within the text? We suggest the term tree alignment to indicate the situation where two markup schemes choose to bracket off the same text elements. We propose a general method for determining agreement between two analyses. We then describe an efficient implementation, which is also modular in that the core of the implementation can be reused regardless of the format of markup used in the corpora. The output of the implementation on the Susanne and Penn treebank corpora is discussed.
Introduction to the special issue on computational linguistics using large corpora, Kenneth W. Church and Robert L. Mercer generalized probabilistic LR parsing of natural language (corpora) with unification-based grammars, Ted Briscoe and John Carroll accurate methods for the statistics of surprise and coincidence, Ted Dunning a program for aligning sentences in bilingual corpora, William A. Gale and Kenneth W. Church structural ambiguity and lexical relations, Donald Hindle and Mats Rooth text-translation alignment, Martin Kay and Martin Roescheisen retrieving collocations from text - Xtract, Frank Smadja using register-diversified corpora for general language studies, Douglas Biber from grammar to lexicon - unsupervised learning of lexical syntax, Michael R. Brent the mathematics of statistical machine translation - parameter estimation, Peter F. Brown et al building a large annotated corpus of English - the Penn treebank, Mitchell P. Marcus et al lexical semantic techniques for corpus analysis, James Pustejovsky et al coping with ambiguity and unknown words through probabilistic models, Ralph Weischedel et al.
A large collection of texts may be reached through the Internet and this provides a powerful platform from which common-sense knowledge may be gathered. This paper presents a system that contains a core knowledge base structured around WordNet, a lexical database, capable of extracting contextual information from a given input text. Such context information is then used to retrieve other texts from the Internet that relate to that context. When processed by the system, these new texts bring more information that represents an enhanced domain context for the initial text. This is an incremental method for text processing that acquires domain knowledge from other texts. The paper describes the system architecture, its core knowledge base and inference engine, and the acquisition of new knowledge from corpora.
This technical report is an appendix to Eisner (1996): it gives superior experimental results that were reported only in the talk version of that paper. Eisner (1996) trained three probability models on a small set of about 4,000 conjunction-free, dependency-grammar parses derived from the Wall Street Journal section of the Penn Treebank, and then evaluated the models on a held-out test set, using a novel O(n^3) parsing algorithm. The present paper describes some details of the experiments and repeats them with a larger training set of 25,000 sentences. As reported at the talk, the more extensive training yields greatly improved performance. Nearly half the sentences are parsed with no misattachments; two-thirds are parsed with at most one misattachment. Of the models described in the original written paper, the best score is still obtained with the generative (top-down) "model C." However, slightly better models are also explored, in particular, two variants on the comprehension (bottom-up) "model B." The better of these has an attachment accuracy of 90%, and (unlike model C) tags words more accurately than the comparable trigram tagger. Differences are statistically significant. If tags are roughly known in advance, search error is all but eliminated and the new model attains an attachment accuracy of 93%. We find that the parser of Collins (1996), when combined with a highly-trained tagger, also achieves 93% when trained and tested on the same sentences. Similarities and differences are discussed.
This paper describes the multilingual text editor MtScript developed in the framework of the MULTEXT project.MtScript enables the use of many differentwriting systems in the same document (Latin, Arabic,Cyrillic, Hebrew, Chinese, Japanese, etc.). Editingfunctions enable the insertion or deletion of textzones even if they have opposite writing directions.In addition, the languages in the text can be marked,customized keyboard input rules can be associated witheach language and different character coding systems(one or two bytes) can be combined. MtScript isbased on a portable environment (Tcl/Tk). MtScript.1.1version has been developed underUnix/X-Windows (Solaris, Linux systems) and otherversions are planned to be ported to the Windows andMacintosh environments. The current 1.1 versionpresents several limits that will be fixed in futureversions, such as the justification of bi-directionaltexts, printing support, and text import/exportsupport. Future versions will use SGML and TEI norms,which offer ways of encoding multilingual texts andare to a large extent meant for interchange.
Abstract Previous studies examining the association between social comparison processes and body image dissatisfaction have yielded inconsistent findings. This study examined whether such discrepancies are due to either the use of identical comparison targets for all subjects or variability in body mass. Specifically, 216 subjects were randomly assigned to one of three experimental conditions: self-generated upward comparison group, self-generated downward comparison group, or control group. Dependent variables were measures of body image. Results indicated that increasing body mass and trait comparison tendencies were associated with increased body dissatisfaction. However, the experimental manipulation did not affect body image ratings. Results suggest that social comparison processes may operate similarly over a range of body mass index (BMI) values.
We show how a treebank can be used to cluster words on the basis of their syntactic behavior. The resulting clusters represent distinct types of behavior with much more precision than parts of speech. As an example we show how prepositions can be automatically subdivided by their syntactic behavior and discuss the appropriateness of such a subdivision. Applications of this work are also discussed. 1 Introduction The construction of classes of words, or calculation of distances between words, has frequently drawn the interest of researchers in natural language processing. Many of these studies aimed at finding classes based on co-occurrences, often combined with the aim of establishing semantic similarity between words (McMahon and Smith, 1996; Brown et al., 1992; Dagan, Markus, and Markovitch, 1993; Dagan, Pereira, and Lee, 1994; Pereira and Tishby, 1992; Grefenstette, 1992). We suggest a method for clustering words purely on the basis of syntactic behavior. We show how the necessary...
Computer self-efficacy and outcome expectancy scales were developed using 306 responses to a questionnaire distributed by a national mail survey to end users of computer systems in a variety of functional business areas. Confirmatory factor analysis using a structural equations approach was used to develop three scales. The scales were found to demonstrate satisfactory psychometric properties. The reliability coefficients for these scales were as follows: .85 for computer self-efficacy; .88 for work-related outcome expectancy; and .89 for personal outcome expectancy. The scales provide a strong foundation from which to refine the measurement of computer self-efficacy and outcome expectancy. From these refinements, empirical models that include self-efficacy and outcome expectancy as determinants of information technology acceptance at the individual level of analysis can be improved.