Papers reviewed and determined not to be word norm studies. Use the flag icon to report errors or suggest re-inclusion.
16504 papers
Humanities computing (HC) has failed to integrate into its practices many of the key theoretical elements of contemporary text and discourse theory. This has in turn contributed to the marginalization of HC in research and teaching. Outdated theoretical models must be abandoned in order to develop a critical discourse based on the insights of HC. HC projects remain far too attached to micro-analyses and have not developed the theoretical and methodological tools necessary to undertake systemic macro-analyses on the level of discourse. Given that texts are a mixture of determinate and dynamic systems, recent developments in chaos theory may be of help in modelling the interrelationship of these elements at discourse level.
I present an entropy measure for evaluating parser performance. The measure is fine-grained, and permits us to evaluate performance at the level of individual phrases. The parsing problem is characterized as statistically approximating the Penn Treebank annotations. I consider a series of models to "calibrate" the measure by determining what scores can be achieved using the most obvious kinds of information. I also relate the entropy measure to measures of recall/precision and grammar coverage.
For the analysis of continuous discourse in a wide range of corpora, it is essential both to model and to expand whole-language lexical resources (e.g.,Roget’s International Thesaurus), in order to make such whole-language lexical resources adaptable to differentiated-discourse domains by means of rapid extensibility. Thus, rapidly extensible lexicons are of interest as special-domain extensions to a whole-language lexicon. My presentation argues for the validity of this approach, with specific reference to a viable conceptual, whole-language, foundational lexicon,Roget’s International Thesaurus (1962).
Efficient exploratory data analysis (EDA) may be aided by succinct, but informative, graphical representations (e.g., Tukey plots) that convey information about central tendency, variability, and shape of distributions, and that permit detection of outliers. Using research strategies adapted from studies of cross-modal perceptual equivalence, we show how auditory analogies of such displays may offer an effective alternative to visual plots for EDA.
The MR Report Assistant is designed to perform the types of routine tasks that computers do well, while freeing the skilled psychologist to concentrate on tasks that are best done by knowledgeable humans. The Assistant is intended for reporting the results of psychological evaluation of known or suspected mental retardation patients using a standard battery consisting of the WAIS-R, the Vineland, and an interview. The program encourages the inclusion of additional information by writing to a disk file suitable for editing with a word processor, rather than directly to a printer. Research and training are encouraged by making the program available to qualified persons at no charge.
The elements of desirable research design for the evaluation of educational technology are discussed with reference to the context of existing research. Sources of internal invalidity, type of compared educational activity, and outcome measures are considered. Finally, recommendations regarding the direction of evaluation research are made. Research designs that take into account the characteristics of the learner, the software, and the teacher preferably within the framework of a model of the learning process should be adopted.
ABSTRACT: Despite the current extensive research on many variables associated with English language testing, little attention has been focused on the inherent variability in the linguistic ‘norms’ for Standard English that are generally tested. This paper examines data from domains of Standard English in several native‐speaker (e.g. British) and non‐native (e.g. Malaysian) varieties of English, focusing on linguistic and attitudinal factors leading to changes in these norms. It then discusses implications of these changes for test construction aimed at assessing proficiency in English as a global language, particularly through such standardized instruments as the Test of English for International Communication (TOEIC).
Describes a synthesis-by-analogy system which is a model of novel-word pronunciation by humans. It uses analogy in both orthographic and phonological domains and is applied to the pronunciation of novel words in British English and German. A major part of this cross-language study concerned the impact of implementational choices on performance, where this was defined as the ability of the system to produce pronunciations in line with those given by humans. The size and content of the lexical database on which any analogy system must be based were also considered. The better performing implementations produced useful results for both British English and German. However, best results for each of the 2 languages were obtained from different implementations. The system described is also a psychological model of reading aloud. (German & French abstracts) (PsycINFO Database Record (c) 2018 APA, all rights reserved)
Although many scholars in literature currently seem mainly interested in theory, the focus on literary texts is what defines literature studies. Computer technology and the statistical methods it fosters are applicable to both the theoretical and to the interpretative issues which scholars of literature habitually address. Genette's distinction between the homodiegetic and the autodiegetic perspective in first-person narrative can be confirmed statistically. Roquentin's loneliness inLa nausée can be shown to be a formal characteristic of the type of novel he narrates, thus validating his commentary on his society. The computer can be used to deal with standard literary questions in a principled fashion, and a new orientation of literature studies on a cultural history model, which Mark Olsen recommends, is not necessary.
Attention is drawn to the need for controlling (during encoding) and checking (after encoding) the quality or accuracy of musical data. Some large databases of melodies are now becoming available, and methods of control and checking are presented which are specially suited to these. Two applications are discussed in detail: to Gregorian Chant and to German folksong. An effective method in tonal and modal music is found to be the investigation of melodic progressions which remain unusual even after amalgamation by transposition to a central register.
The study of signs is divided between those scholars who use the Saussurian binary sign (semiology) and those who prefer Charles Peirce's tripartite sign (semiotics). The common view of the opposition between the two types of signs does not take into consideration the methodological conditions of applicability of these two types of signs. This is particularly important in the field of literary studies and hence for the preparation of electronic programs for text analysis. The Peircian sign explicitly entails the discovery of a truth of meaning that claims to be universal and not reducible to a collection of opinions based on fragmented information; it also imposes the task of elucidating a transhistorical and universal significantion encoded in a text. Contrary to Peirce's view of the sign, our use of computer programs for text analysis, however, demonstrates that we implicitly treat every literary text as a set of linguistic data (letters, phonemes, syntagmatic segments, etc.) which are reducible to units that can be treated separately. A brief comparison of the results obtained from computer analyses of the French poet Stéphane Mallarmé's text, “Le Cygne,” with those obtained from two Peircian analyses (by Riffaterre and Champigny) of the same text demonstrates that our current methods of computer textual analysis are based on a Saussurian semiology, which is unidimensional and limited, and that these methods are still quite unable to produce a semiotic interpretation based on a totalizing hierarchy of the text's various discursive components.
The Kay Elemetrics nasometer measures nasalance, a parameter of speech that reflects the proportion of total acoustic energy that is emitted nasally, making it possible to infer velopharyngeal (VP) function noninvasively. Nasometric evaluation is potentially widely applicable in the clinical assessment of suspected VF impairment. For clinical use, a patient’s mean nasalance on a passage of known phonetic composition must be compared to age-appropriate population norms. Most potential clinical subjects are young; many are preliterate. Passages for which norms have been established (Zoo, Rainbow, NasalSentences) are syntactically, semantically, and lexically complex, phonetically heterogeneous, phonologically mature, and long. Individual child subjects’ nasalance scores, if obtained at all, are therefore likely to be contaminated by artifacts created through hesitation noises, filled pauses, phonetic deviance from normed target, age differences, and measurement errors induced by coaching procedures. Differences in phonetic content between normed and actual utterances are almost inevitable; they lead to uninterpretable results. This study reports a technique for obtaining clinically useful nasalance scores from young, preliterate subjects, even those evidencing phonological deficits or noncompliant behavior. Large-n norms for preschool and primary children are presented.
Recent articles have noted that humanities computing techniques and methodologies remain marginal to mainstream literary scholarship. Mark Olsen's paper discusses this phenomenon and argues for large scale analyses of text databases that would incorporate a shift in theoretical orientation to include greater stress on intertextuality and sign theory. Part of Olsen's argument revolves on the need to move away from the syntactic and overt grammatical elements of textual language to more subtle semantics and meaning systems. While provocative and important, Olsen's stance remains rooted in literary theoretical constructs. Another level of language, the cognitive, offers equally interesting challenges for humanities computing, though the paradigms for this type of computer-based exploration are derived from disciplines traditionally removed from the humanities. The riddle, a nearly universal genre, offers a window onto some of the cognitive processes involved in deep level language function. By analyzing the riddling process, different methods of computational modelling can be inferred, suggesting new avenues for computing in the humanities.
The Learning Systems Department at Siemens Corporate Research is investigating the use of concept spaces to increase retrieval effectiveness. Similar to a semantic net, a concept space is a construct that defines the semantic relationships among ideas. The current focus of our research is to exploit the information in such a structure to ameliorate known shortcomings of statistical retrieval methods while maintaining the statistical methods' robustness. Our initial concept space is extracted from WordNet, a manually-constructed lexical database developed at Princeton University.
This paper describes the use of correspondence analysis to create the “space” of a book, constructs that of Kierkegaard'sFear and Trembling as an illustration, and distinguishes three separate contexts of some of its most important words: thespatial context (where the search word lies in that named and ordered space); theoverall context (the x words closest to the search word in multi-dimensional space); and the “role/sense” context (the words associated with the search word in each of its most important roles, some of which may represent new senses.) It describes the identification of these contexts, discusses their importance and concludes by noting certain respects in which the procedure might perhaps be improved.
We should follow Mark Olsen's lead and think with maximum ambition of the role of the computer in supporting literary research of the highest order. Thus the computer enables us to answer one of the great questions of literary criticism: how does a given writer contribute to the changing language? We can now chart the influence of given writers by correlating their words and phrasing with computerized dictionaries so as to produce profiles and histories of the way words have entered the language.
Could a troupe of monkeys really produce Shakespeare if allowed to bang away at the word processor long enough? This age old question, commonly referred to as the Eddington problem, relates to fundamental issues of probability, and is examined in this article in a new light. Based on earlier research by the physicist William Bennett, Jr., the author describes a data structure which enables the computer to simulate the hypothetical monkeys. Exploiting principles of cryptology, the computer leads the simulated monkeys closer to their goal. Though the intent of the article is to encourage humanistic speculation, the final result proves to be quite practical and may come as a surprise to computer scientists and humanists alike.
Optical Character Recognition is shown to be significantly more expensive than keyboarding, using off-shore contractors, for entry of large amounts of text where high accuracy is required. Using large test samples in French and English, the paper indicates that OCR applications which require significant post-scan editing are labor intensive projects that can be accomplished more efficiently by keyboarding. Most OCR systems are still not capable of entering large amounts of text accurately enough to avoid an expensive editing step.
This article uses recent work on the computer-aided analysis of texts by the French writer Céline as a framework to discuss Olsen's paper on the current state of computer-aided literary analysis. Drawing on analysis of syntactic structures, lexical creativity and use of proper names, it makes two points: (1) given a rich theoretical framework and sufficiently precise models, even simple computer tools such as text editors and concordances can make a valuable contribution to literary scholarship; (2) it is important to view the computer not as a device for finding what we as readers have failed to notice, but rather as a means of focussing more closely on what we have already felt as readers, and of verifying hypotheses we have produced as researchers.
Benoît de Cornulier's writings on French poetry concentrate on metrical boundaries, or caesura; however, the the criteria upon which he bases his analyses are useful in studying rhythm, or the relationship between syllables within the alexandrine's twohémistiches. This study focuses on three aspects of rhythm in French poetry: the definition of rhythm following Cornulier; the development of a method using the computer to detect rhythmic patterns in traditional isometrical alexandrines; the results of such a study when applied to three classical seventeenth-century plays which are composed of isometrical alexandrines (Corneille'sPolyeucte, Racine'sPhèdre, and Molière'sLe Tartuffe).
The effects of insulting political campaign rhetoric were examined in a laboratory setting. In Study 1, subjects categorized as for, against, or undecided about the issue of French language rights in English Canada read a political debate transcript focusing on this issue in which one candidate insulted or did not insult the other candidate. Backlash was evident on trait ratings: the insult source was rated more negatively in the insult condition than in the control condition, whereas ratings of the target were unaffected. Affective ratings of both candidates were lower in the insult condition than in the control condition. Subjects’ attitudes about French language rights also influenced their impressions of the candidates, but did not interact with the insult manipulation. In Study 2, subjects read debate transcripts embedded with insults that attacked either controllable target attributes, uncontrollable target attributes, or that contained no insults. The major findings of Study 1 were replicated, with the additional finding that insults directed at uncontrollable traits of the target produced effects in the same direction as but more extreme than those effects noted for controllable‐trait insults. These findings are discussed in terms of political campaign tactics and the application of attribution theory to political person perception.
: Attempts to improve the performance of handwriting recognition systems have often involved the exploitation of linguistic constraints such as syntax or semantics. In either case, successful implementation requires the creation of a lexical database containing the relevant information. However, to create a database of semantic information from scratch for a realistically sized vocabulary is an enormous task - which is a major reason why so many semantic theories fail to "scale up" from the small, artificial domains in which they were developed. A better approach is to use existing sources of semantic information, such as machine-readable dictionaries (from which definitions may be extracted) and text corpora (from which collocations may be derived). This paper describes the development of techniques that use such resources to improve the performance of handwriting recognition systems. KEY WORDS: handwriting recognition, semantic information, machine-readable dictionaries, text corpora...
The present study examined the effects of observational focus and performance cues on rating accuracy. It was hypothesized that these two factors would affect ratings independently: a) subjects given a good performance cue would rate the target more positively than subjects given a poor performance cue and, b) subjects using an event focus would rate the target person more accurately than subjects using a person focus. One hundred twenty undergraduates viewed the same videotape and subsequently stated which of a set of 48 behaviors were exhibited by the target person. The results supported the hypothesized performance cue effect. Observational focus did not have the hypothesized main effect on rating accuracy but was involved in a four‐way interaction. The interaction indicated that, as the focus of the subject's attention broadened to the entire event, the errors associated with the performance cue effect were lessened. The subjects using an event focus were likely to generate ratings characterized by a positivity bias, whereas those using a person focus were more likely to generate ratings biased by performance cues.
ABSTRACT: Until the late nineteenth century French was the dominant international language of modern Western Europe, and, with the spread of empire, many other areas around the globe. Now French itself feels threatened by the spread of English. The protection of the French language has both ‘negative’ and ‘positive’ aspects. On the negative, defensive side, purists rail against the specific qualities of English, and the social values these are said to represent. The rejection of English and specifically American influence on the French language is related to the rejection of modernity, and of the nation‐state based on shared political principles rather than shared culture. On the positive, offensive side, international French‐language organizations promote French as the language of francophone brotherhood. This co‐operative effort, however, conflicts with the traditional formulation and role of linguistic norms in French society. With the changing composition of French society, the definition of the nation‐state in France, and the conception of linguistic norms of French, may be in the process of changing.
During the early 1980s, Itek Optical systems and the United States Air Force developed and published an Image Interpretability Rating Scale based on a set of image interpretability criteria specific to military operations. There appears to be a similar need for a commonly understood and used scale to describe the relative resolution of aerial photographs in fields of environmental concern. Using our extensive backgrounds in image acquisition and use and in assessing problems common to natural resource managers, we set out to develop an image rating scale. Our objective was to more clearly relate image characteristics (such as scale, spatial resolution, spectral resolution, and format) to specific examples of image interpretation tasks which are performed by environmental scientists. To do this, we created rating scale categories and established the minimum resolution requirements in terms of ground resolved distance for a variety of typical environmental features. We found that our needs for image quality closely parallel needs in military intelligence. This information will provide designers of image acquisition systems specific criteria to consider in addition to more abstract resolution measures such as lines per millimeter. We suggest that this IIRS will compliment the earlier Itek scale and like it, will become a standard for discussions of image quality and resolution.
PURPOSE: This article summarizes demographic characteristics of Korean Americans and reviews health issues in this population. METHODS: The authors reviewed census data, monographs, books and medical literature published in the English language. FINDINGS: Korean Americans are one of the fastest growing Asian American groups in the United States. They are a heterogeneous population, differing in their cultural, religious and linguistic norms. Early data suggest that Korean Americans have lower overall mortality rates than the general United States population. However, they have special problems with respect to stomach cancer, liver cancer, hepatitis, mental health, access to health care and other issues. CONCLUSIONS: The health issues of Korean Americans have been generally overlooked until the present time. Examination of these emerging problems should contribute to the future of American health. RELEVANCE TO ASIAN PACIFIC ISLANDER AMERICAN POPULATIONS: This paper is particularly relevant to Korean Americans. KEY WORDS: Korean Americans, health education, hepatitis B, stomach cancer and tuberculosis
Eighty pianists each listened to 21 trials of solo piano music. Trials consisted of two different performances of the same excerpt, and the same music was played on all trials for any given subject. Over the 21 trials, seven different interpretations were presented in all possible pair-wise combinations. Forty subjects listened to a slow excerpt from Liszt's Totentanz, and the other forty listened to a fast excerpt from the same piece. The subjects' task was to select which of the two performances on each trial, if either, they preferred. Subjects were assigned randomly to one of four conditions: preference only, preference plus use of the musical score, preference plus use of rating scales, and preference plus use of both the musical score and rating scales ( n = 10 per treatment per excerpt). Results indicated that the use of the musical score and of rating scales both separately and in combination with each other did not improve consistency as compared to their nonuse. Subjects who responded to rating scales, however, were more consistent when they did not also use the musical score than when they did use it. Subjects were less consistent for the slow excerpt than they were for the fast excerpt, and consistency was unaffected by their piano experience. Presentation order within trials (first versus second excerpt) did not affect ratings; however, ratings on Trials 12-21 were slightly and significantly higher than ratings on Trials 1-10. Also, preference for a performance was not affected by preference for the performance immediately preceding it.
Resume Dans cette etude, nous avons voulu distinguer deux aspects de la familiarite des concepts: la familiarite avec le sens des mots et la familiarite avec leur referent. La familiarite avec le sens a ete mesuree en demandant aux sujets d'evaluer la facilite avec laquelle ils pourraient identifier une definition des mots presentes. La familiarite avec le referent a ete mesuree en demandant aux sujets d'evaluer la facilite avec laquelle ils pourraient identifier une photographie des objets designes. Les sujets se sont en general declares plus confiants de reconnaitre une photographie qu'une definition des items presentes. Cependant, l'inverse s'est produit pour les items les moins familiers, lesquels etaient generalement des termes designant des categories subordonnees. Malgre ces differences, les jugements de familiarite relatifs aux photographies et aux definitions se sont averes tres fortement correles. Des correlations plus faibles ont ete obtenues entre nos mesures de familiarite d'une part, et la frequence des mots ou la typicite des concepts, d'autre part.Abstract In the study reported, we have tried to distinguish two aspects of concept familiarity namely, familiarity with the meaning or sense of words and familiarity with their referent. Meaning familiarity was measured by asking subjects to evaluate the ease with which they would recognize a definition of the words presented. Familiarity with the referent was measured by asking subjects to evaluate the ease with which they would recognize a picture of the objects corresponding to the words presented. In general, the subjects believed that they would more likely recognize a picture than a definition of the items. However, the opposite pattern was obtained for the least familiar items which, interestingly, consisted mainly of subordinate category terms. Despite these differences, the picture and definition familiarity ratings were highly correlated. Less pronounced correlations were obtained between our measures of familiarity and measures of word frequency and concept typicality.Cette recherche a ete motivee par deux buts principaux, l'un de nature methodologique et l'autre de nature plus theorique. Au plan methodologique, il s'agissait essentiellement de colliger des donnees normatives concernant le sentiment de familiarite qu'evoquent differents concepts. Au plan theorique, nous avons voulu verifier s'il etait possible de distinguer la familiarite avec le sens des mots de la familiarite avec leur referent. En d'autres termes, nous avons voulu determiner si, dans leurs jugements subjectifs de familiarite, les sujets pouvaient distinguer les proprietes qui definissent les mots d'une part, des objets qu'ils designent, d'autre part.La question de la nature des concepts est absolument centrale en psychologie cognitive et en psycholinguistique. Au cours des deux dernieres decennies, nombre de recherches ont ete effectues dans le but d'identifier les facteurs qui influent sur la representation et l'utilisation des connaissances conceptuelles. Parmi les facteurs consideres, la typicite(f.1) est un de ceux qui a recu le plus d'attention suite aux travaux effectuees par Rosch et coll. (Rosch, 1973, 1975; Rosch & Mervis, 1975) et par Smith et coll. (Rips, Shoben & Smith, 1973; Smith, Shoben & Rips, 1974) sur la representation mentale des categories taxonomiques.Ces travaux ont montre que les differents membres d'une meme categorie ne sont pas tous juges egalement representatifs de la categorie. Ainsi, par exemple, lorsque Rosch a demande a des sujets de coter dans quelle mesure differentes especes d'oiseaux correspondaient a leur idee d'un oiseau, les pinsons ont ete juges plus representatifs ou plus typiques que les autruches. Les memes travaux ont en outre montre que les membres plus typiques sont generalement categorises plus rapidement que les membres moins typiques. Enfin, les membres les plus typiques d'une categorie ont aussi tendance a etre mentionnes plus frequemment lorsque la tache des sujets est de nommer les membres d'une categorie qui leur viennent spontanement a l'esprit. …
<h3>Objective.</h3> —To assess the feasibility and measurement characteristics of ratings completed by professional associates to evaluate the performance of practicing physicians. <h3>Design.</h3> —The clinical performance of physicians was evaluated using written questionnaires mailed to professional associates (physicians and nurses). Physician-associates were randomly selected from lists provided by both the subjects and medical supervisors, and detailed information was collected concerning the professional and social relationships between the associate and the subject. Responses were analyzed to determine factors that affect ratings and measurement characteristics of peer ratings. <h3>Setting and Participants.</h3> —Physician-subjects were selected from among practicing internists in New York, New Jersey, and Pennsylvania who received American Board of Internal Medicine certification 5 to 15 years previously. <h3>Main Outcome Measure.</h3> —Physician performance as assessed by peers. <h3>Results.</h3> —Peer ratings are not biased substantially by the method of selection of the peers or the relationship between the rater and the subject. Factor analyses suggest a two-dimensional conceptualization of clinical skills: one factor represents cognitive and clinical management skills and the other factor represents humanistic qualities and management of psychosocial aspects of illness. Ratings from 11 peer physicians are needed to provide a reliable assessment in these two areas. <h3>Conclusions.</h3> —These findings suggest that it is feasible to obtain assessments from professional associates of practicing physicians in areas such as clinical skills, humanistic qualities, and communication skills. Using a shorter version of the questionnaire used in this study, peer ratings provide a practical method to assess clinical performance in areas such as humanistic qualities and communication skills that are difficult to assess with other measures. (<i>JAMA</i>. 1993;269:1655-1660)
This research examined how training and experience, family roles, and gender of observed family leadership affect ratings of family and individual parent functioning. Observer perceptions of video-taped family interviews were compared to determine the relative influence of male and female leadership. The influence of training was determined by comparing the perceptions of trained and naive observers. One hundred forty (70 naive, 70 experienced) adults from a large university and midwestern city participated in the experiment. Participants were randomly assigned to one of two family interview conditions: (a) matriarchical, and (b) patriarchical. After viewing the video-taped family interviews, participants completed forms assessing the parents and families. The assessments included ratings of family communication, family conflict negotiation, family support and nurturance, parent helpfulness to family functioning, and parent individual adjustment. For the three family measures, a multivariate analysis of variance was conducted. Results indicate experienced observers are vulnerable to bias against female family leadership whereas naive observers show no comparable tendency. For the two measures of parent functioning, a mixed model multivariate analysis of variance was conducted. Results indicate that although experienced observers offer more critical ratings than their naive counterparts, vulnerability to bias against women as family leaders is not significantly influenced by training and experience. Evaluations of parent and family functioning appear the result of the combined influence of several factors. As a consequence, vulnerability to bias in family evaluations likely occurs in a complicated and potentially discrete fashion. This complexity promotes serious concern with respect to the ability of training to reduce the propensity for gender discriminating practice.
Young, young-old, and old adults were examined in immediate and delayed episodic recognition of common odors. Items were presented in 3 different formats: name-only, odor-only, or odor-name. Ss made familiarity ratings for all items at study. In the delayed recognition test, Ss were asked to name the odors. Young Ss outperformed the 2 older age groups in both recognition tests, although the 2 older groups did not differ. Performance was higher in the odor-name condition than in the single-format conditions. Both familiarity and naming were related to recognition in all age groups. Most important, when naming was statistically controlled, age differences in odor recognition disappeared, suggesting that access to verbal labels largely determine age differences in recognition of common odors. Finally, the finding that recognition was enhanced in both young and older Ss in the odor-name condition suggests that odor memory may involve a similar degree of plasticity as other varieties of episodic memory.
In this paper we propose to define selectional preference and semantic similarity as information-theoretic relationships involving conceptual classes, and we demonstrate the applicability of these definitions to the resolution of syntactic ambiguity. The space of classes is defined using WordNet [8], and conceptual relationships are determined by means of statistical analysis using parsed text in the Penn Treebank.
We describe a generative probabilistic model of natural language, which we call HBG, that takes advantage of detailed linguistic information to resolve ambiguity. HBG incorporates lexical, syntactic, semantic, and structural information from the parse tree into the disambiguation process in a novel way. We use a corpus of bracketed sentences, called a Treebank, in combination with decision tree building to tease out the relevant aspects of a parse tree that will determine the correct parse of a sentence. This stands in contrast to the usual approach of further grammar tailoring via the usual linguistic introspection in the hope of generating the correct parse. In head-to-head tests against one of the best existing robust probabilistic parsing models, which we call P-CFG, the HBG model significantly outperforms P-CFG, increasing the parsing accuracy rate from 60% to 75%, a 37% reduction in error.
Linguistic Exploitation of Syntactic Databases: The Use of the Nijmegen Linguistic Database Program. Hans van Halteren and Theo van den Heuvel. Amsterdam: Rodopi, 1990. Pp. ix + 207. $33.50. - Volume 15 Issue 1
It is shown that grammatical inference is applicable to natural language processing. Given the wide and complex range of structures appearing in an unrestricted natural language like English, full grammatical inference, yielding a comprehensive syntactic and semantic definition of English, is too much to hope for at present. Instead, the authors focus on techniques for dealing with ambiguity resolution by probabilistic ranking; this does not require a full formal Chomskyan grammar. They give a short overview of the different levels and methods being investigated at CCALAS for probabilistic ranking of candidates in ambiguous English input.
The lexicons for Knowledge-Based Machine Translation systems require knowledge intensive morphological, syntactic and semantic information. This information is often used in different ways and usually formatted for a specific NLP system. This tends to make both the acquisition and maintenance of lexical databases cumbersome, inefficient and error-prone. In order to solve these problems, we have developed a program called COOL which automates the acquisition and maintenance processes and allows us to standardize and centralize the databases. This system is currently being used in the ESTRATO machine translation project at the Center for Machine Translation.
Lack of a critical mass of scholars involved with the computer-assisted analysis of texts (CAAT), coupled with insufficient communication among various sectors of the literary and linguistic disciplines, has led to a skewed notion of computing humanists' work among their colleagues. This paper highlights the gap through examples of misunderstood humanist needs and achievements drawn from both recent media reports and humanities conferences. It suggests that networking and less modesty in manuscript submission can be at least partial solutions. The author cites some of his own published work and work-in-progress on Stendhal and Gobineau in refuting Mark Olsen's thesis that the dominance of single- or dual-author studies must be the cause of CAAT's “failure” to make significant inroads in mainstream literary journals. The author builds a case for the use of both diachronic and synchronic lexico-statistical data in carrying out such studies successfully. He recommends a new “Synthetic Criticism” where relevant quantitative methods would not be absent.
This article addresses the methodological problem of the non-linear representation of philosophical systems in a computerized knowledge base. It is a problem of knowledge representation as defined in the field of artificial intelligence. Instead of a purely theoretical discussion of the issue, we present selected results of a practical experiment which has in itself some theoretical significance. We show how one can represent different philosophies using CODE, a knowledge engineering system developed by artificial intelligence researchers. The hypothesis is that such a computer based representation of philosophical systems can give insight into their conceptual structure. We argue that computer aided text analysis can apply knowledge representation tools and techniques developed in artificial intelligence and we estimate how philosophers as well as knowledge engineers could gain from this cross-fertilization. This paper should be considered as an experiment report on the use of knowledge representation techniques in computer aided text analysis. It is part of a much broader project on the representation of conceptual structures in an expert system. However, we intentionally avoided technical issues related to either Computer Science or History of Philosophy to focus on the benefit to enhance traditional humanistic studies with tools and methods developed in AI on the one hand and the need to develop more appropriate tools on the other.
Computer-aided literature studies have failed to have a significant impact on the field as a whole. This failure is traced to a concentration on how a text achieves its literary effect by the examination of subtle semantic or grammatical structures in single texts or the works of individual authors. Computer systems have proven to be very poorly suited to such refined analysis of complex language. Adopting such traditional objects of study has tended to discourage researchers from using the tool to ask questions to which it is better adapted, the examination of large amounts of simple linguistic features. Theoreticians such as Barthes, Foucault and Halliday show the importance of determining the linguistic and semantic characteristics of the language used by the author and her/his audience. Current technology, and databases like the TLG or ARTFL, facilitate such wide-spectrum analyses. Computer-aided methods are thus capable of opening up new areas of study, which can potentially transform the way in which literature is studied.
The difficulties inherent in the evaluation of educational software are described in terms of the tradeoffs between internal, external, and ecological validity. Larger issues in evaluation research design and computer-based instruction are highlighted by primary and metaanalytic studies designed to reveal the effects of computer simulations in psychology classrooms and laboratories. The effectiveness of classroom and laboratory computer activities depends on how the inclusion of software, as well as the evaluation process itself, changes the entire instructional process.
It has long been known that the pupil dilates as a consequence of attentional effort. But the function that relates attentional input to pupillary output has never been the subject of quantitative analysis. We present a system analysis of the pupillary response to attentional input. Attentional input is modeled as a string ofattentional pulses. We show that the system is linear; the effects of input pulses on the pupillary response are additive. The impulse response has essentially a gamma distribution with two free parameters. These parameters are estimated; they are fairly constant over tasks and subjects. The paper presents a method of estimating the string of attentional input pulses, given some average pupillary output. The method involves the technique of deconvolution; it can be implemented with a public-domain software package, Pupil.
Commercial database programs such as dBase and Paradox, although developed originally for business applications, are versatile and powerful tools that can be used for an academic purpose such as evaluating student performance. They can be used to write and store test questions, assemble and print classroom or on-line laboratory tests, and calculate grades, test statistics, and so forth. Databases are flexible, unlike textbook “ancillary” test bank programs that are inextricably bound to the strictly linear format and brief shelf life of specific textbook editions. A prototypical relational database program is described, with which an instructor can produce tests based on generic terms adapted from Boneau’s (1990) study of psychological literacy, as well as on behavioral learning objectives adapted from Bloom’s (1956) taxonomy of educational objectives. As a relational database, the program integrates terms, objectives, questions, tests, and test scores, and avoids unnecessary data duplication and waste of computer storage space.
Although Columbus'Diary of the first voyage to America as we know it is largely a transcription of the original diary carried out by Bartolomé de las Casas, commentators and readers often treat it as if it were Columbus' work alone. Editions published to date do not separate the explorer's narrative from that of his transcriber or editor. Since style can influence readers' perceptions of a writer's personality, it is important to determine characteristics of writing attributed to Columbus that may pertain instead to his trascriber. This study employs the computer to explore the style of Las Casas and that of Columbus. Differences in the writing of each “author” emerge with computer assistance by isolating Columbus' words from those of his transcriber and analyzing selected features of vocabulary, sentence length, and syntax.1
This paper concurs with Mark Olsen's premise that computer-aided literature studies should take a different direction, one that is more suited to the computer's strength in analyzing large corpora of texts. However, the authors take issue with his conclusion that a reorientation of the notions of textual analysis is necessary in order to exploit the computer's capabilities. Contemporary medieval studies already provides us with models of textual analysis which are well suited to computer development. Though they stem from the particularities of medieval textual production, these models can perhaps be useful in the study of modern literatures.