Papers reviewed and determined not to be word norm studies. Use the flag icon to report errors or suggest re-inclusion.
16504 papers
Some recent studies in computational linguistics have aimed to take advantage of various cues presented by punctuation marks. This short survey is intended to summarise these research efforts and additionally, to outline a current perspective for the usage and functions of punctuation marks. We conclude by presenting an information-based framework for punctuation, influenced by treatments of several related phenomena in computational linguistics.
The Limits of Irony Berel Lang (bio) Irony... disciplines and punishes. Kierkegaard If my title is not ironic—and it isn't—it must be tendentious, since it implies that irony has limits and so invites only the question of what or where they are. But the basis for this assumption is not self-evident, and indeed the emphasis in modern theories of irony runs directly counter to it. This more common view finds irony limit less, with its characteristic starting point in a contradiction or reversal that then evokes an indefinite number of further displacements. On this account, one ironic turn opens the way to another and that one opens on still another, with no nonironic end in sight except for the ironic consciousness itself with its denial of any stopping point that might interrupt its own continuing reflection. On this view, irony initiates (more precisely, takes place in) an infinite movement—perhaps regress, perhaps progress—that violates the supposed boundaries of every context in which it appears: it is irony that irony affirms, not just this or that single turn. Thus, irony discloses its purpose not in its individual assertions or truths but in the common manner of their rendering; in this sense irony is an adverb, not a noun—insatiable in its appetite in a way that "persons, places, or things" could not be even if they imagined or willed that state. Kierkegaard's view at the beginning of his doctoral dissertation, The Concept of Irony, of irony's "infinite absolute negativity" 1 may seem too burdensome for merely human practice, and indeed that view later turns out to be too weighty even for him. But master-ironist that he was, his opening description of irony's apparently limitless reach warrants attention, and not only because in it he echoes, at least for this one time, his antagonist of choice, Hegel. I shall later be disputing this view of irony and propose an alternative to it, but we do well to look first at the process leading to its historical dominance. This scrutiny is the more pertinent because irony does not openly admit its expansionist design. More typically, it avoids speaking of itself at all—and just this reticence motivated the critic and aesthetician Jean-Paul's proposal for an "irony mark" to serve, like the question mark [End Page 571] or the exclamation point, as a means of identification: readers could then be confident of recognizing an ironic text when they met one. But Jean-Paul's idea never caught on, evidently because it would be self-defeating. The avowal by ironic statements of their "point" or even of their identity as ironic could only undermine the effect they strive for: to have the reader himself, as the writer's "double," infer and then inscribe the ironic conclusion in its reversal of the assertions offered at face value (at one-face value, that is) by the writer. We thus look to the ironic text for what it omits at least as much as for what it acknowledges; its ironic claims, fittingly, are the more present for being the more absent. Admittedly, other literary tropes or figures—metaphor, allegory—typically also avoid announcing themselves; for them, too, blatancy is self-defeating or, more basically, antifigurative. (Figurative language that cited itself as figurative might turn that reference too into a figure of speech—perhaps under the title of "truth" or "candor." 2 ) Even allowing for this reticence, however, individual figures often do provide clues of figurative design—all of them, for example, involving deviation from a linguistic norm. None, however, displays irony's combination of contradiction and then subordination of what is first "literally" affirmed. The characteristic reversal in oxymoron, for example, emphasizes a specific quality (the "silence" of "loud silence"), not contradiction as such. Similarly for catachresis—assigning a figurative term where no literal term exists (the "leg" of a chair)—and for aporia, which in so many words places its subject beyond words, as ineffable: although perhaps hinting beyond their immediate referents, these make no claim to represent a mode of consciousness. By contrast, the reversals of irony, in addition to turning sentences or texts back...
Difficulties in the acquisition of the reading process. This paper tackles the learning problems of reading starting from the visual habilities which are the basis of control and recognition of written information. As the processes of recognition are completed with those of comprehension, we introduce strategies for their training from a lexical point of view and from structural strategies of content, this last aspect incorporates a set of basic norms for their realization. Finally, we mention that this theoretical framework of diagnosis and intervention is in process of investigation with a representative group of subjects through wich we try to confirm our hypothe-
Our sample consisted of 420 children who were making normal progress in learning to read (35 girls and 35 boys at each of the ages from 7 to 12 years inclusive, from a variety of schools in Sydney). They were given a set of 30 exception words, 30 nonwords, and 30 regular words to read aloud. As predicted by a dual-route account of learning to read, the correlations between regular word and exception word reading accuracy and between regular word and nonword reading accuracy were higher than the correlation between exception word and nonword reading accuracy; also as predicted by this account, regular word reading accuracy was higher than exception word and nonword reading accuracy.We present our data as age-related norms which can be used in conjunction with our materials to assess how well children in this age range who appear to have reading difficulties are acquiring the lexical and nonlexical reading procedures as they learn to read.
A procedure that processes a corpus of text and produces numeric vectors containing information about its meanings for each word is presented. This procedure is applied to a large corpus of natural language text taken from Usenet, and the resulting vectors are examined to determine what information is contained within them. These vectors provide the coordinates in a high-dimensional space in which word relationships can be analyzed. Analyses of both vector similarity and multidimensional scaling demonstrate that there is significant semantic information carried in the vectors. A comparison of vector similarity with human reaction times in a single-word priming experiment is presented. These vectors provide the basis for a representational model of semantic memory, hyperspace analogue to language (HAL).
Reviews251 tion (152), one ofthe interesting points ofwhich is an important passage on the problem of 'possible features', i.e., features that are not minimally distinctive, hence theoretically unnecessary, but still highly informative. As far as the terminological and terminographic products are concerned, printed versions are analyzed separately from terminological databanks. The section on die latter topic contains highly useful subsections on type and quality of data, and on programs and software for retrieval and interrogation. The last sections on the various national and international institutions and their regulatory activities (158) bring us to die next to last chapter, die one dealing widi the legal status of standardization. The translation again occasionally renders comprehension somewhat difficult. For example, in order to understand a sentence (179) asserting diat a legally normative text very often "is constitutional by its generality but only amounts to a declaration of intent," whereas a simple order or decree may have more immediate impact, one must realize that the term 'constitutional' is not intended in the American English meaning, i.e., = 'valid', because it is opposed to 'unconstitutional, hence not valid', but rather in the meaning of 'constitutional ', i.e., 'basic', not casuistically normative. This chapter also contains a decisive discussion ofthe notion of linguistic norm (167). The discussion is based on French attitudes toward norm and normative activities. This is one of the best descriptions in English of die French attitudes. Chapter 1 1, the last in the book, offers a case study, namely of the terminology in Le Grand Robert, which is a general-language dictionary edited by the author. The summary of his experience is most instructive. In sum, the book is most valuable and interesting. Manual of Specialised Lexicography: The Preparation of Specialised Dictionaries. 1995. Henning Bergenholtz and Sven Tarp. With contributions by Grete Duvâ, Anne-Lise Laursen, Sandro Nielsen, Ole NordingChristensen, and Jette Pedersen. Amsterdam/Philadelphia, John Benjamins. 254 pages. $49.95 hardcover. The book contains a step-by-step description of the ways of compiling a dictionary of a specialized register. At each of these steps, the authors discuss all the options open at that stage, so that the treatment is methodologically rich. The major portion of this book was written by Bergenholtz and Tarp, and at some distance, by Nielsen. The authorship of each section is clearly indicated in the Preface; in any case, one of die achievements of the primary authors is diat the final printed text is seamless, completely unified in style, and devoid of repetitions. The Manual, comprising 13 chapters, deals first with the basic notions; as the first step, a delineation between terminology (and terminography) on the one hand, and specialized lexicography on the other, is sought, motivated by die notion that terminology is necessarily prescriptive. What is meant by this is that 252Reviews terminology must necessarily be based on die normative authority of a board, committee, or institute that has decided which terms should be used and what meanings diese terms are to carry. This certainly is, in general, true, but only in areas where such normative institutions actually exist; if one examines a dictionary of, say, linguistic terminology (such as Akhmanova 1966, Knobloch 1961, or Springhetti 1962), one will find descriptive explanations of synonymic and homonymie terms; diere is notiiing else that can be done in die absense of any normative authority, particularly when die lexicographer is faced widi contradictory understandings of the same terms by competent schools of linguistic thought. In any case, by 'specialised lexicography' die authors mean, broadly speaking, a lexicography diat deals widi special languages and registers. How much professionaljargon and even general-language vocabulary should, or may, also be included in such specialized dictionaries is one of the points on which individual lexicographers differ. These preliminary clarifications lead to a discussion of the basic problems of specialized lexicography, such as dictionary functions and die like. Particularly important and well-conceived is section 3.5, 'die use ofcomputers in specialised dictionary making'. It is a detailed discussion of methods for building corpora, concordances, and databases, and for editing dictionary entries. The discussion is highly concrete and replete with examples, and with indications and comparative discussions of software programs; it will...
Yvonne Touchard - Corrections of the Annales du Brevet: Ready-made thinking, or reflexions on how to argue a point in school. The study analyzes texts submitted by various publishers of the Annales du Brevet as "correct versions" of argumentative texts. Written by adults (teachers), these texts imitate the way pupils are supposed to write. The enunciative, discursive, syntaxic and lexical characteristics are studied. From the point of view of enunciation, they are polyphonic, since the writer adds his/her own adult voice to that of the pupil he is supposed to be imitating. They apply the rules usually taught in middle schools to argue a point, showing an expertise that makes these texts into "models", thus demonstrating what the ideal discursive norms are taken to be. The points chosen for the argument also inform us as to the values propounded by the school system.
Literature instructors are using hypertext to enhance their teaching in a broad variety of ways that includes putting course materials on the WWW; creating online tutorials; using annotated hypertexts in addition to or in lieu of print texts; having students write hypertexts; examining the medium of hypertext as a literary and cultural theme; and studying hypertext fiction in the context of traditional literature classes. The article describes examples of each of these uses of hypertext in teaching literature and provides sources of further examples of and information on using hypertext as a teaching tool in literature classes.
Abstract OUR objective in this book is to trace and illustrate the main changes that have taken place in the Russian language since the beginning of the twentieth century, and particularly those between 1917 and the late 1980s, a time which can be identified as the Soviet period, by now a reasonably complete stage in the history of Russia and the former Soviet Union. The period addressed in this book witnessed the unprecedented expansion of the standard Russian language, both stylistically and geographically. Around 1900, the range of functions of standard Russian increased, in some instances replacing Church Slavonic, in some instances with the emergence of new phenomena. Russian replaced Church Slavonic as the language of sermons (parallel to the maintenance of Church Slavonic as the language of liturgy); standard Russian emerged as the language of the courts (especially following the judicial reform of the 1860s), of political debate in the Russian Parliament (Duma), of growing business and trade, and as the language of poetry and prose read to large audiences from the stage of concert-halls and literary cabarets. Theatre, the most popular type of entertainment in Russia of the second half of the nineteenth century and at the beginning of the twentieth century, played a very important role in promoting a uniform language norm, especially in pronunciation; the role of theatre as the bearer of such a norm became clear by the mid 1Soos (Panov 1990: 94). In the twentieth century, radio, cinema, and television, gradually taking over from the theatre, gave the standard language the ability to travel, unknown in earlier periods; mass media, together with compulsory mass education, became much more effective than writing alone in acquiring new speakers for the emerging standard. These new speakers of the standard, in tum, had a profound effect on the language itself: never before had it borne the fingerprints of so many diverse speakers, many of whom were, following M. V. Panov’s characteriza tion, ‘new recruits to culture’ (Panov 1990: 18, 22).
Small groups are called upon to make important policy decisions under a wide variety of procedural constraints. ACPE is a flexible, computerized system for conducting small-group voting experiments. It permits researchers to examine the impact of electronic communication on group deliberation and choice. The system runs under a variety of different personal computer networks and is designed to permit the specification of voting rules, communication, and group sizes. The system also facilitates the study of group process by tracking all messages sent and votes taken. An experiment in which the system was used is briefly described.
Eric Brill introduced transformation-based learning and showed that it can do part-of-speech tagging with fairly high accuracy. The same method can be applied at a higher level of textual interpretation for locating chunks in the tagged text, including non-recursive ``baseNP'' chunks. For this purpose, it is convenient to view chunking as a tagging problem by encoding the chunk structure in new tags attached to each word. In automatic tests using Treebank-derived data, this technique achieved recall and precision rates of roughly 92% for baseNP chunks and 88% for somewhat more complex chunks that partition the sentence. Some interesting adaptations to the transformation-based learning approach are also suggested by this application.
The article gives a general introduction to the form and function of the TEI header, points out some of the reasoning of the Text Documentation Committee that went into its design, and discusses some of its limitations. The TEI header's major strength is that it gives encoders the ability to document the electronic text itself, its source, its encoding principles, revisions, and characteristics of the text in an interchange format. Its bibliographical descriptions can be loaded into standard remote bibliographic databases, which should make electronic texts as easy to find for researchers as texts in other media, including print. Its major weakness is that it does not yet provide the ability for retrieval across texts in a networked environment, which users may want now or in the future.
A computerized multimedia instructional system has been developed that adheres to behavioral systems principles for presenting both adaptive programmed instructional materials and laboratory simulations. This instructional software system, called MediaMatrix, is both an authoring environment and a presentation vehicle that adapts the complexity of presentations in real time to changes in a student’s current rate of progress. It incorporates an automated knowledge-generation system that tracks all interactions between the user and any instructional objects within the system. Such knowledge is used both for a research database and for an artificial-intelligence engine that constructs an estimate ofconcept association networks (Verplanck, 1992a). Such networks reflect a learner’s developing knowledge and skill base, and may be used for tutorial advising during student use.
espanolEn castellano, la silaba parece actuar como un elemento prelexico-fonologico de relacion con el nivel lexico. Su mayor o menor frecuencia determina la cantidad de palabras que se activaran en el nivel lexico. Esta cualidad de restriccion lexica nos ha sugerido la necesidad de elaborar un estudio normativo en el cual los sujetos evocaban palabras de 2 y 3 silabas a partir de una inicial dada que despues pueden ser utilizados como base para diversos estudios experimentales. Se produjeron asi 130 conjuntos de candidatos competidores lexicos (ccl), que proporcionan informacion sobre dos aspectos fundamentales: su tamano (numero de candidatos lexicos) y la accesibilidad relativa de cada una de las salidas lexicas que las componen. EnglishThe syllable in Spanish could operate as a phonological and prelexical unit with relation to lexical level. The syllable frequency determines the number of words that will be activated at lexical level. This quality of accessibility constriction suggests the need to produce some candidate set norms. In this normative study, the subjects recover two and three syllable words from one initial syllable given. In this way, we obtained 130 sets of lexical competitor candidates (CCL), which provide information about two basic aspects: the set size (i.e. number of lexical candidates), and relative accessibility of each lexical output composing the set.
This paper focusses on the types of questions that are raised in the encoding of historical documents. Using the example of a 17th century Scottish Sasine, the authors show how TEI-based encoding can produce a text which will be of major value to a variety of future historical researchers. Firstly, they show how to produce a machine-readable transcription which would be comprehensible to a word-processor as a text stream filled with print and formatting instructions; to a text analysis package as compilation of named text segments of some known structure; and to a statistical package as a set of observations each of which comprises a number of defined and named variables. Secondly, they make provision for a machine-readable transcription where the encoder's research agenda and assumptions are reversible or alterable by secondary analysts who will have access to a maximum amount of information contained in the original source.
This exploratory field study evaluated a bilingual computerized speech-recognition cellular telephone prototype of the Center for Epidemiological Studies—Depression scale (CES-D). Thirty Spanish and 22 English speakers completed both computer-telephone and face-to-face CES-D methods and an oral depression checklist in counterbalanced order. Both language groups reported high positive ratings for the computer-telephone method, with the English sample preferring the computer-telephone over the face-to-face method. In both samples, the computer-telephone method yielded high internal consistency estimates, strong alternate form reliabilities, and similar high correlations to the depression checklist. Both groups reported significantly elevated scores with the computer-telephone method, but total score variances for both methods did not differ. Computer-telephone limitations included occasional misrecognitions and template training constraints.
In this paper, we concentrate on justifying the decisions we made in developing the TEI recommendations for feature structure markup. The first four sections of this paper present the justification for the recommended treatment of feature structures, of features and their values, and of combinations of features or values and of alternations and negations of features and their values. Section 5 departs briefly from the linguistic focus to argue that the markup scheme developed for feature structures is in fact a general-purpose mechanism that can be used for a wide range of applications. Section 6 describes an auxiliary document called a “feature system declaration” that is used to document and validate a system of feature-structure markup. The seventh and final section illustrates the use of the recommended markup scheme with two examples, lexical tagging and interlinear text analysis.
We set out to develop a computer-assisted finger-tapping task (the T3) that would measure motor speed much like the Reitan test, but that would also measure endurance. Data were collected for a convenience sample on both the T3 and the Reitan finger-tapping test. Moderate and significant correlations were obtained between the T3 and the Reitan test for both hands. Mean scores for the first 50 sec of the T3 were approximately 0.15 taps greater than the mean Reitan score for both the preferred and the nonpreferred hands, while the mean scores for the full 2 min of the T3 were 1.52 taps less than those of the Reitan test for the preferred hand, and 1.32 taps less for the nonpreferred hand. The mean for the last 40 sec with the preferred hand averaged 3.93 taps (7.62%) slower than for the first 40 sec, whereas for the nonpreferred hand, the difference was 5.12 taps (11.15%). These results are consistent with our intent to develop measures of (1) relatively pure motor speed (the first 50 sec of the T3); (2) motor speed combined with endurance (the full 2 min of the T3); and (3) finger endurance (the first 40 sec compared with the last 40 sec of the T3).
focuses on certain differences between Standard English (SE) and the African-American Vernacular English (AAVE) spoken in many inner-city and rural American communities / presents 2 points that are crucial for the treatment of languages that differ much more widely / 1st, vernacular languages—the ways that ordinary people ordinarily talk—are not 'ungrammatical' or otherwise imperfect approximations to standard or literary linguistic norms / 2nd, small changes in an abstract grammatical system may produce complex patterns of change on the surface of the language, magnifying the apparent differences (PsycINFO Database Record (c) 2016 APA, all rights reserved), no part of cognitive science illustrates the problem of abstract inference better than the interpretation of linguistic zeroes: the absence of the very behavior that we have come to observe / [discuss] the interpretation of such linguistic zeroes / engage a particular problem that has been the center of much linguistic research: t)
There is a great deal of variation in the encoding of spoken texts in electronic form, both with respect to the types of features represented and the way particular features are rendered. This paper surveys problems in the electronic representation of speech and presents the solutions proposed by the Text Encoding Initiative. The special tags needed for the encoding of spoken texts are discussed, including a mechanism for temporal alignment. Further work is needed on phonological aspects, parallel representation, and on the development of software which connects the systematic underlying representation with a workable format for input and display.
Many languages make use of word-formation devices to allow speakers or writers to create new words when the existing vocabulary proves inadequate. In this paper we consider how these devices can be expressed formally, allowing them to be used in word- and sentence-generation, for dictionary expansion, and the like. The paper begins with some typical word-formation rules drawn mostly from French. Attention is drawn to some features of these rules which must be captured in any formal representation. The formal representation of a basic lexical transformation is presented in some detail, along with a number of examples. A computer implementation of the transformation system is described, together with a range of applications. A discussion of static and dynamic generation leads to the concept of an inverted transformation.
This paper traces the history of the Text Encoding Initiative, through the Vassar Conference and the Poughkeepsie Principles to the publication, in May 1994, of theGuidelines for the Electronic Text Encoding and Interchange. The authors explain the types of questions that were raised, the attempts made to resolve them, the TEI project's aims, the general organization of the TEI committees, and they discuss the project's future.
In this paper, a method for indexing cross-language databases for conceptual query matching is presented. Two languages (Greek and English) are combined by appending a small portion of documents from one language to the identical documents in the other language. The proposed merging strategy duplicates less than 7% of the entire database (made up of different translations of the Gospels). Previous strategies duplicated up to 34% of the initial database in order to perform the merger. The proposed method retrieves a larger number of relevant documents for both languages with higher cosine rankings when Latent Semantic Indexing (LSI) is employed. Using the proposed merge strategies, LSI is shown to be effective in retrieving documents from either language (Greek or English) without requiring any translation of a user's query. An effective Bible search product needs to allow the use of natural language for searching (queries). LSI enables the user to form queries with using natural expressions in the user's own native language. The merging strategy proposed in this study enables LSI to retrieve relevant documents effectively using a minimum of the database in a foreign language.
ABSTRACT: This paper analyzes the data from three questionnaires administered to Pakistani male and female journalists, teachers, and university students in Islamabad, Karachi, and Lahore during a period from 1987 to 1992. The first questionnaire deals with respondents’(320) choice of a model of English (British, American, or Pakistani). The second and third questionnaires measure the acceptability of selected Pakistani English lexical and grammatical items (150 respondents) and complementation types (165 respondents). Results show that while an exonormative model of English (British) still has considerable influence in the former colony (in both ‘ideal’ as well as in reported ‘actual’ usage), a Pakistani norm is also beginning to emerge. This trend is most evident in respondents’ acceptance of typically Pakistani features of English such as Urdu borrowings, Urdu‐English hybrids, and local morphological and syntactic innovations.
This paper discusses the basic design of the encoding scheme described by the Text Encoding Initiative'sGuidelines for Electronic Text Encoding and Interchange (TEI document number TEI P3, hereafter simplyP3 orthe Guidelines). It first reviews the basic design goals of the TEI project and their development during the course of the project. Next, it outlines some basic notions relevant for the design of any markup language and uses those notions to describe the basic structure of the TEI encoding scheme. It also describes briefly the “core” tag set defined in chapter 6 of P3, and the “default text structure” defined in chapter 7 of that work. The final section of the paper attempts an evaluation of P3 in the light of its original design goals, and outlines areas in which further work is still needed.
In studies of author attribution, measurement of differential use of function words is the most common procedure, though lexical statistics are often used. Content analysis has seldom been employed. We compare the success of lexical statistics, content analysis, and function words in classifying the 12 disputedFederalist papers. Of course, Mosteller and Wallace (1964) have presented overwhelming evidence that all 12 were by James Madison rather than by Alexander Hamilton. Our purpose is not to challenge these attributions but rather to useThe Federalist as a test case. We found lexical statistics to be of no use in classifying the disputed papers. Using both classical canonical discriminant analysis and a neural-network approach, content analytic measures — the Harvard III Psychosociological Dictionary and semantic differential indices — were found to be successful at attributing most of the disputed papers to Madison. However, a function-word approach is more successful. We argue that content analysis can be useful in cases where the function-word approach does not yield compelling conclusions and, perhaps, in preliminary screening in cases where there are a large number of possible authors.
This paper chronicles the work of the TEI textual criticism working groups through several phases, documenting how and why the design goals were shaped by the requirements of several distinct user communities and by the nature of the textual evidence itself. Encoding schemes for the representation of physical details of textual witnesses were unified with encoding schemes for critical editing practices when it was observed that the two phenomena were inextricably layered and linked within real texts. Rationale is offered for the development teams' adherence to exceedingly general design principles: (a) the requirement that the encoding notations be neutral in text-theoretic terms; (b) the need to accommodate dramatically different text-transmission phenomena and research goals within diverse text-critical arenas; (c) the need for commensurability of the text-critical markup with encoding notations used in closely related text-analytic research. The paper also assesses the results of the effort in terms of the encoding scheme's adequacy for several scholarly purposes: suggestions are made concerning the need for programmatic testing, for refinement, and for extension of the encoding model to support a broader range of text-transmission phenomena and research objectives.
For a decade or so, Liszt thrilled and astounded audiences at a time when virtuosity (often as an end in itself) was the norm and the piano had rapidly evolved into a form recognisable as a close relative of the instrument we know today. During this period Liszt frequently performed hisGrandes Etudes (1838), which he had developed from his boyhoodEtude en 12 exercices (1826) and which he later revised and technically simplified asEtudes d'Exécution transcendante (1851). Although Liszt's own performances cannot be recreated, procedures for generating electronic realizations, which contain nuances of balance and tempo, are described. All three versions of the eighth of Liszt's set of 12 studies are used for illustration. Contrary to received opinion, it is argued that the 1838 version is more satisfying than the 1851 revision and that this is due to its formal structure.
This paper traces a progression of four computer-based methods for studying and fostering both the structure and the on-line development of knowledge. Each empirical technique employs ECHO, a connectionist model that instantiates the theory of explanatory coherence (TEC). First, verbal protocols of subjects’ reasonings were modeled post hoc. Next, ECHO predicted, a priori, subjects’ text-based believability ratings. Later, the bifurcation/bootstrapping method was developed to elicit and account for individuals’ background knowledge, while assessing intercoder reliability regarding ECHO simulations. Finally,Convince Me, our “reasoner’s workbench,” automated the explication both of subjects’ knowledge bases and of their belief assessments; theConvince Me software permits contrasts between the model’s predictions and subjects’ proposition-wise evaluations. These experimental systems enhance our understanding of the relationships among—and determinant features regarding—hypotheses, evidence, and the arguments that incorporate them.
Natural language processing will grow into a vital industrial technology in the next five to 10 years. But this growth depends on the development of large linguistic databases that capture natural language phenomena [1, 2]. Another important theme for future work is development of large knowledge bases that are shared widely by different groups. One promising approach to such knowledge bases draws on natural language processing and linguistic knowledge. This article describes the EDR Electronic Dictionary [3], which seeks to provide a foundation for linguistic databases, and explains the relation of electronic dictionaries to very large knowledge bases.
The prevalence of the use of teams in a variety of occupations and environments has increased the importance of investigating the processes involved in their performance. However, in the past, there have been few methodologies available for the investigation of team performance. The present manuscript attempts to contribute to this area of research by describing the rationale underlying the use of computer-based simulations in research on team performance. This is followed by a review of the networked simulations that are currently being used in team-performance research. This review emphasizes the capabilities provided by the networks and the types of research concerns for which they are effective. Finally, the application of this technology to the broader study of group performance is discussed.
Many aspects of the guidelines of the Text Encoding Initiative (TEI) are applicable to corpora and text collections, and to the texts that these contain. As the first large corpus developed using mark-up conforming to the guidelines, the British National Corpus (BNC) is a test-bed for many TEI-developed mechanisms. This is particularly true in the case of the TEI header, which has three intended applications — to describe a corpus, to describe an individual text, and as a free-standing bibliographic record — all of them used by the BNC. This paper describes the application of the TEI header to the BNC. It is intended that this information should, through a description of experience on a practical project, serve as a guide for those wishing to use TEI headers in the documentation and management of other corpora and collections of texts.
Another current major issue in lexical semantics is the definition and the construction of real-size lexical databases that will be used by parsers and generators in conjunction with a grammatical system. Word meaning, terminological knowledge representation and extraction of knowledge in machine readable dictionaries are the main topics addressed. They really represent the backbone of a lexical semantics knowledge base construction.
There are many ways in which to estimate thresholds from psychometric functions. However, almost nothing is known about the relationships between these estimates. In the present experiment, Monte Carlo techniques were used to compare psychometric thresholds obtained using six methods. Three psychometric functions were simulated using Naka-Rushton and Weibull functions and a probit/logit function combination. Thresholds were estimated using probit, logit, and normit analyses and least-squares regressions of untransformed orz-score and logit-transformed probabilities versus stimulus strength. Histograms were derived from 100 thresholds using each of the six methods for various sampling strategies of each psychometric function. Thresholds from probit, logit, and normit analyses were remarkably similar. Thresholds fromz-score- and logit-transformed regressions were more variable, and linear regression produced biased threshold estimates under some circumstances. Considering the similarity of thresholds, the speed of computation, and the ease of implementation, logit and normit analyses provide effective alternatives to the current “gold standard”—probit analysis—for the estimation of psychometric thresholds.
Parsing is often seen as a combinatorial problem. It is not due to the properties of the natural languages, but due to the parsing strategies. This paper investigates a Constrained Grammar extracted from a Treebank and applies it in a non-combinatorial partial parser. This parser is a simpler version of a chunking-and-raising parser. The chunking and raising actions can be done in linear time. The short-term goal of this research is to help the development of a partially bracketed corpus, i.e., a simpler version of a treebank. The long-term goal is to provide high level linguistic constraints for many natural language applications. 1
The article contains a stylistic and linguistic analysis of Kasprowicz’s poems published in the period of Modernism. The following elements were selected from the texts of the poems: 1. deviations from the linguistic norm of the turn of the XIX and XX century, as well as potential variants of this norm; 2. phenomena remaining within the norm, but, e.g., typical of the period in question or of a given text. The selection comprised the following elements: phenomena considered then as rare, archaisms, dialectal phrases, localismus and neologismus. They were all described in chapters on phonetics, morphology, syntax and vocabulary. The kind of the elements selected and the extent to which they permeat the texts point out to moderate dialectizing and slight archaizing of poetry. Moderate permeation with innovations is also present. All these devices were, in accordance with the tradition, introduced primarily by means of lexical phenomena. Most of the linguistic means used can be classified as traditional and known in literature. The quality of these means and the way they are used make it possible for Kasprowicz to be counted among poets using interesting language.
In this paper we show, for the first time, how Radial Basis Function (RBF) network techniques can be used to explore questions surrounding authorship of historic documents. The paper illustrates the technical and practical aspects of RBF's, using data extracted from works written in the early 17th century by William Shakespeare and his contemporary John Fletcher. We also present benchmark comparisons with other standard techniques for contrast and comparison.
As a result of an analysis of the language of Jerzy Żuławski’s poetry, some devices have been selected which go beyond the linguistic norms common and universally accepted on the turn of the XIX and XX centuries. The selected linguistic devices can be divided into traditional, known in literature, and idiosyncratic, which make the language of the poems somewhat original. All these devices have been described and evaluated in chapters on phonetics, inflexion, syntax and lexis. Phonetic idiosyncracies are scarce and not many of them are typical of artistic language. Inflexion reflects the commonly accepted norms and patterns of style. Trying to make the poems original and their language individual, the poet used syntactic means more frequently than lexical ones, which is not typical of methods of stylization. The language of Żuławski’s poetry is „classical"; moderately permeated with traditional stylistic devices; slightly „poeticized”, archaized and individual. It lacks dialectal phrases and expressive devices characteristic of Modernist poets.
This article describes the major problems in devising a TEI encoding format for dictionaries, which, because of their high degree of structuring and compression of information, are among the most complex text types treated in the TEI. The major problems for this task were (1) the tension between generality of the description, in order to be widely applicable across dictionaries, and descriptive power, that is, the ability to describe with precision the particular structure of any given dictionary; and (2) the need to accommodate different views and uses of the encoded dictionary, for example, as printed object and as a database of information.
There are currently two philosophies for building grammars and parsers -- Statistically induced grammars and Wide-coverage grammars. One way to combine the strengths of both approaches is to have a wide-coverage grammar with a heuristic component which is domain independent but whose contribution is tuned to particular domains. In this paper, we discuss a three-stage approach to disambiguation in the context of a lexicalized grammar, using a variety of domain independent heuristic techniques. We present a training algorithm which uses hand-bracketed treebank parses to set the weights of these heuristics. We compare the performance of our grammar against the performance of the IBM statistical grammar, using both untrained and trained weights for the heuristics.
An extragrammatical sentence is what a normal parser fails to analyze. It is important to recover it using only syntactic information although results of recovery are better if semantic factors are considered. A general algorithm for least-errors recognition, which is based only on syntactic information, was proposed by G. Lyon to deal with the extragrammaticality. We extended this algorithm to recover extragrammatical sentence into grammatical one in running text. Our robust parser with recovery mechanism -- extended general algorithm for least-errors recognition -- can be easily scaled up and modified because it utilize only syntactic information. To upgrade this robust parser we proposed heuristics through the analysis on the Penn treebank corpus. The experimental result shows 68% ¸ 77% accuracy in error recovery. 1 Introduction Extragrammatical sentences include patently ungrammatical constructions as well as utterances that may be grammatically acceptable but are beyond the synta...
Adult children of alcoholics' (n = 68) perceptions of their relationships with parents were compared with those of a control sample (n = 37) to examine independent and joint influences of interpersonal status and affect on family dynamics. Visual metaphors for relationships using circle drawings and a status-affect rating scale from the Grasha-Ichiyama Psychological Size and Distance Scale were employed. Compared with the control group, adult children of alcoholics drew smaller circles to represent themselves, i.e., indicating less interpersonal status, only when assessing their relationships with their fathers. Analyses of status-affect ratings showed that the drawings of smaller circles reflected feeling less competent, i.e., having less personal knowledge and expertise, rather than perceptions of being submissive in the relationship. The distance drawn between the circles of adult children of alcoholics and their parents, i.e., psychological distance, was much larger than that of the control group. Ratings showed that perceptions of a negative emotional climate and submissiveness together accounted for 25% of the unique variance in predicting psychological distance. Perceptions of being submissive, however, were not associated with perceptions of psychological distance among adult children of nonalcoholic parents.
Research based on a treebank is active for many natural language applications. However, the work to build a large scale treebank is laborious and tedious. This paper proposes a probabilistic chunker to help the development of a partially bracketed corpus. The chunker partitions the part-of-speech sequence into segments called chunks. Rather than using a treebank as our training corpus, a corpus which is tagged with part-of-speech information only is used. The experimental results show the probabilistic chunker has more than 92% correct rate in outside test. The well-formed partially bracketed corpus is a milestone in the development of a treebank. Besides, the simple but effective chunker can also be applied to many natural language applications.