Papers reviewed and determined not to be word norm studies. Use the flag icon to report errors or suggest re-inclusion.
16504 papers
Introduction / B. Boguraev and E. Briscoe -- Placing the dictionary on-line / Alshawi, B. Boguraev, and D. Carter -- An independent analysis of the LDOCE grammar coding system / E. Akkerman -- Utilising the LDOCE grammar codes / B Boguraev and E. Briscoe -- The derivation of a large computational lexicon o English from LDOCE / J. Carroll and G. Grover -- LDOCE and speech recognitio D. Carter -- Analysing the dictionary definitions / H. Alshawi -- Meaning an structure in dictionary definitions / P. Vossen, W. Meijs, and M. den Broede -- A tractable machine dictionary as a resource for computational semantics Y. Wilks... [et al.] -- Conclusion / B. Boguraev and E. Briscoe -- Appendic: Lexical database-user guide -- Semantic types of LDOCE verbs -- Dative alternations -- Lexicon development environment-user guide -- The Longman semantic codes -- The Longman grammar coding system.
The purpose of this study is to examin the relationships between the following three data sets:(1) evaluative ratings (5-point scale) on 10 most advancedscientific technologies (i. e., artificial intelligence, bio-technology, nuclear power generation, space technology, linear motor car, tube baby, 5th generation computer, super conductivity, organ transplant, and high-speed reactor), (2) familiarity ratings (5-point scale) on the same technologies, and (3) ratings on the subject's own personality (Y-G Personality Inventory). A hierarchical component analysis technique for the multiset data (Murakami, 1989) was applied. Three first order components for the image data sets were interpreted as indicating dangerous and harmful, useful and development, and personally beneficial dimensions, while the two first order components for the familiarity, as industrial technology and medical technology. Second order components showed several intersting relationships between image and familiarity which were found to be specific to particular scientific technologies. For example, industrial technologies high in the familiarity demension tended to have an image of usefulness and development.
The teaching of literature through CAI raises problems of both a linguistic and instructional nature; student involvement and creativity in studying literature, and especially poetry, is difficult to build into a computer-based lesson. We have confronted these difficulties in the lessonPoetry I, which introduces undergraduates to basic concepts of poetic verse in a design using screen display, speech synthesis, and verse processing to maximize interactivity and student involvement.
A cooperative team of researchers from various Italian universities are collaborating on a project for Latin lexicography. This article describes the linguistic work being done on the Latin language. The various problems — and their solutions — are discussed: allographs, homographs, source material and classification methods, word searching, etc.
The IDG was originally founded to carry out research into the collection and processing of documentation relating to Italian legislation, case law and legal authority. The Institute has since concentrated on automated documentation and legal informatics, as well as the application of artificial intelligence to the law. This article describes the many projects undertaken at the Institute.
Abstract Subjects estimated how many Germans drink vodka or beer, or estimated the caloric content of these drinks. The former judgment, but not the latter, produced contrast effects on subsequent ratings of how ‘typically German’ various drinks are. Thus, highly accessible extreme stimuli did only affect ratings if the first judgment pertained to the same underlying dimension.
This manuscript is a description of the research activities of the Lessico Intellettuale Europeo. The Centre works on the lexicographical analysis of philosophical and scientific texts, mainly of the seventeenth and eighteenth centuries. An important project is the Philosophical Dictionary of the Seventeenth and Eighteenth Centuries. The LIE series of publications include a number of indices, concordances and lexicons. Another project at the Centre is the Thesaurus Mediae et Recentioris Latinitatis. The Centre also publishes the proceedings of the three-yearly Colloqui internazionali.
This article describes a second aspect of the Project for Latin Lexicography (see previous article). We here concentrate on two aspects of the project. First, we describe the morphological analyzer, which comprises a base dictionary, a table of suffixes, a table of endings and a table of postfixes. Second, we describe the lemmatization module, which operates by reference to a series of grammatical codes or information given for the base, and reference codes.
A research paradigm is suggested that combines the perspectives of the humanistic scholar and the behavioral scientist: After differentiating the popularity of actual aesthetic products using archival indices and then subjecting these compositions to objective computer content analyses, further statistical treatment may divulge the intrinsic properties responsible for differences in impact. This approach is illustrated by an analysis of the 154 sonnets attributed to William Shakespeare. Each sonnet was partitioned into four consecutive units (three quatrains and a couplet), and then a computer gauged how the number of words, different words, unique words, primary process imagery, and secondary process imagery changed within each sonnet. Taking advantage of a previous objective measure of the relative aesthetic merit of the sonnets, and implementing a statistical search for interaction effects, it was demonstrated that Shakespeare's lexical choices adopt a discernible pattern in the highly popular creations that is not found in the more obscure poems. Perhaps the most fascinating aspect of this pattern shift is the distinct manner in which the poet modifies his vocabulary when composing the concluding couplet in his best sonnets.
Ever since the pioneering work of Roberto Busa, computers have had an important work in humanistic research. The Istituto per le Scienze Religiose has been involved in such research since the 1970s. This article describes the Pope John XXIII project whose aim is to produce a computer-aided index of all the Pope's writings in both Italian and Latin; as well as future projects currently under consideration.
Several standard solutions have been developed for problems encountered when monitoring laboratory response sensors with microcomputers. These problems include voltage-level translation, power-supply isolation, contact debouncing, and temporary data storage. Data storage has typically been accomplished through the use of edge-sensitive devices, which introduce difficulties when a researcher wishes to monitor the duration of a response. A circuit is described that includes an optoisolator with a silicon-controlled rectifier at its output stage. The circuit acts as a latch for the detection of brief responses, and, because it is not edge-sensitive, it can also be used to record the duration of sustained responses.
A concordance of the early Italian poetic language is being compiled at the University of Florence, based on the need to recognize the special nature of the language in earlier times. The corpus consists of some forty-five manuscripts, that is, all that remains of book production from the origins to the end of the thirteenth century. Once the work of entering the single word-tokens is done, a complete concordance results, with accompanying grammatical connotations, in which not only are Tuscan and dialectal words grouped under separate standard headwords, but also homographs are clearly distinguished.
Techniques are described for using an IBM PC equipped with a mouse to investigate the reading of character sequences of maximally one display line. Solutions are given to problems that arise and derive from the PC’s real-time clock, the slowness of the display, and registration of input during a test session. A program for using the techniques to study asynchronous perception of printed words is described, and a demonstration program is provided as an appendix.
Machine-readable texts in the humanities are in a period of rapid growth, as are programs for analyzing them, with concomitant problems of finding and choosing the most suitable of each. No one text-base, and no one search program, proves to be wholly suitable for all projects. Three of the currently available electronic editions of Shakespeare are discussed and compared, along with three commercial programs for text analysis on microcomputers.
CorrecText, from Houghton-Mifflin, is a significant advance in grammar checkers, because it uses a full parse of sentences in its analysis. Though limited by the fact that many English sentences are syntactically ambiguous, the program can find many errors in grammar, style, and usage. Still, questions remain about where and how it can be useful. It is not terribly useful on final versions of good prose; the makers assert that it is more useful on unedited prose, where there are many inadvertent errors. An empirical study of much unedited prose would not only verify this assertion, but would help improve the program.
The research described in this paper was originally presented as a dissertation at the Università di Venezia. The objective was to establish the kind of language to which Chinese students are exposed during their primary education. The analysis was based on twelve texts of Guomin Xiaoxue Guoyu. This manuscript concentrates mainly on the techniques used to encode Chinese characters.
This paper describes two research projects, both involving Latin and Greek lexicography. They are undertaken at the University of Florence and at the Italian National Research Council respectively. The one involves the creation of a Dictionary of Justinian's constitutions based on the emperor's legislative lexicon formed in the Corpus Iuris and elsewhere. The most demanding aspect of this task has been the creation of the Dictionary of the Novellae. The other project involves the creation of a Lexicon of the Novellae in the Authenticum version.
The following is a description of a computerized version of the corpus of Latin grammarians published by Heinrich Keil in Leipzig between 1855 and 1880. The intent was to prepare an instrument which would serve both as a key to Keil's corpus and as the basis for a re-edition of the work itself. We discuss the corpus itself, the ways in which it was encoded, the pre-editing work, and how the material was organized for analysis by computer.
Research on changes in Shaw's rhetoric inMrs. Warren's Profession, Major Barbara, andHeartbreak House led me to a heuristic for gaining literary critical control over computer output. This essay describes the eleven-step process: stepping away from the data, stating first premises, developing a working hypothesis, classifying computer-sorted data, marking implicit literary sub-structures, collecting sub-structural data into tables, applying earlier statistical observations, choosing parts for detailed analysis, designing a visual method for representing the analysis, presenting segment by segment analysis of the selected data, and making larger descriptive generalizations. While describing this heuristic, the essay also reports on the Shaw research.
The possible benefits of computing in humanities research are often wasted because of the psychological barriers that computers evoke in non-specialists. This paper examines the underlying causes and suggests some ways of alleviating the problem. One approach in particular, i.e. ease through familiarity, is discussed in more detail. It is illustrated by means of a description of a database system that uses this approach: the Linguistic DataBase, which contains syntactic analysis trees of natural language data.
Based on a corpus of neologisms extracted from Catalan newspapers (published between 1979 and 1983), the author aims to show two fundamental trends in the lexical modernization of Catalan: a) the innovation of socio-political vocabulary reinforces the international character of the lexicon and the models that serve for the formation of new words, while the discourse on the norm gives a clearly purist orientation; b) in the sense of neologisms, there is also a slight tendency towards the elimination of castillianisms, that is, towards a differentiation from Spanish, while the integration of anglicisms and gallicisms is much easier. Despite this, the first trend is predominant, especially in the reduction of internal models of word formation in favor of those that are used, preferably, also in other languages. Finally, reflections are made on the methodological-theoretical value of analyzes for the establishment of a theory of language policy.
After describing the English course and the particular hypertext system that supports it at Brown University, the essay surveys the materials on Context32, that part of the system devoted to literature courses, and narrates how a student uses the system during a typical session in our electronic laboratory/classroom. Next, it presents evidence of the effects of such information technology on student performance, after which it examines the relation of hypertext to contemporary literary theory, in particular to the ideas of decentering, intertextuality, and anti-hierarchical texts. Finally, it explains the continuing developments of Intermedia.
Scholars in the humanities often have to account exhaustively for the structure of large masses of data. Tree-diagrams implemented by means of suitable computer programs can be of considerable assistance in achieving a cohesive representation of the data. This paper discusses the respective merits of the two main approaches to tree representation and introduces a new method based on the use of unrooted trees. After a detailed examination of the topological properties of such trees, two algorithms are described. The second part of the paper consists in practical applications of the method of tree representation to a corpus of contemporary English poetry. Several sets of data made up of both lexical and grammatical items (adjectives, modals, auxiliaries and personal pronouns) have been submitted to the method. The findings are assessed in terms of their heuristic value in the light of modern linguistic theory and compared with the results obtained by means of more traditional statistical procedures.
Based on the ARTFL version of theProfession and excerpts fromEmile, high frequency function and content words, as defined by Brunet, are analyzed via Pearson chi square tests. Next, four measures of narrative voice from the same populations are compared using Markovian chains and further chi square tests. In a third analysis the two orders of evidence are juxtaposed. The lexical and narratological preferences of theVicaire and theGouverneur, while not resolving the problematic of chronological composition (Burgelin, 1969), highlight the distinctiveness of each character.
The purpose of this paper is to show the utility of the application of a non-ultrametric tree-model to textual data. The first part introduces a basic topological property of the tree and the notion of neighbourhood, which reflects the structure of the tree. The second part emphasizes through illustration examples the adequacy of this model for representing different varieties of textual data.
The Société de 1789 was a political club founded in early 1790 to propagate the ideals of the Revolution and the Enlightenment. A systematic analysis of the language found in the public discourse of the Société using simple quantitative techniques suggests important distinctions in comparison to the language found in a baseline sample, a selection of the General Cahiers de doléances of 1789. It is further argued that these differences represent an Enlightened reforming tradition that carried into the French Revolution.
This paper suggests ways in which the pattern-matching capability of the computer can be used to further our understanding of stylized ballad language. The study is based upon a computer-aided analysis of the entire 595,000- word corpus of Francis James Child'sThe English and Scottish Popular Ballads (1882–1892), a collection of 305 textual traditions, most of which are represented by a variety of texts. The paper focuses on the “Mary Hamilton” tradition as a means of discussing the function of phatic language in the ballad genre and the significance of textual variation.
An interface circuit to connect a microphone to an Apple Macintosh computer is described. The Apple Macintosh mouse port is used as the input port, and the microphone activation simulates a mouse press.
This article describes some approaches to imitation analysis and the use of ready-made software for this task. Devising computer-assisted techniques for exploring the conscious literary imitation of style is an application of particular relevance to contemporary Hispanic narrative and one that can be handled with a microcomputer and readily accessible software. The article describes some approaches to imitation analysis and the use of ready-made software to assess the effectiveness of stylistic imitation of eighteenth-century historical chronicle in La renuncia del héroe Baltasar (The Renunciation of the Hero Baltasar), by the Puerto Rican novelist Rodríguez Juliá. Even when employing familiar procedures of text analysis with computer, comparing a fictional text with a multiple and diverse corpus of authentic historical documents requires somewhat unique assumptions and hypotheses, since neither authorship, influence, or authenticity are in question.
A key word with regard to a sub-corpus is a word of which the frequency in that sub-corpus is significantly higher than expected under the hypothesis that its use and the variable “part of the corpus” are mutually independent. A study in literary statistics almost invariably includes a chapter devoted to key words. However, a strong attack has been recently launched upon the way stylometry has been modelling texts since the classical works of Herdan, Guiraud or Muller. In fact statistical modelling seems as valid in stylistics as in any other field of the humanities and social sciences. What is questionable is the fact that many studies in literary statistics are more satisfied with the easy identification of monsters, i.e. literary phenomena unexplained by wrong models, than with the laborious research of models fitting the textual data well. A short examination of the mentioned controversy and the quantitative analysis of an example provided by Laclos' novelLes Liaisons dangereuses endeavour to support this argument.
The study of the history of new words in theNewOED described in this paper was undertaken in 1986-87, and is based on the material then available. Since then, theNewOED has been finished, and PAT, the inquiry system developed at the University of Waterloo for the investigation of theNewOED data base, has been much altered and improved. Nevertheless, this report should prove useful in indicating the potentiality for analyzing the computerizedNewOED and some of the problems. This project is a study of the ways in which new words are created in English at various periods of time. A chronological dictionary 's created listing words introduced into the language over 50 year increments. These words are then classified by the processes used in forming them to show, in proportional terms, if certain processes are more common at some times than at others.
Information theory offers a means for analyzing some constraints on the reading and copying process in Old English. Entropy for strings of various lengths offers a baseline measure of the uncertainty involved in transmission of Old English texts, while avoiding the pitfalls of applying models of modern reading to early medieval practice. Analysis of lengthy prose and verse texts in Old English revealed uniformly high values for entropy at all string lengths. High entropies may be the result of the language's irregular orthography, poetic koiné, and several dialects and imply that the language may have been easy to write but difficult to read. The low redundancy of the language which its high entropy values indicate suggests that the reader of Old English played an enhanced role in “decoding” a text and may provide an explanation for the high variability in the transmission of Old English verse.
The present paper is a critique of quantitative studies of literature. It is argued that such studies are involved in an act of reification, in which, moreover, fundamental ingredients of the texts, e.g. their (highly important) range of figurative meanings, are eliminated from the analysis. Instead a concentration on lower levels of linguistic organization, such as grammar and lexis, may be observed, in spite of the fact that these are often the least relevant aspects of the text. In doing so, quantitative studies of literature significantly reduce not only the cultural value of texts, but also the generalizability of its own findings. What is needed, therefore, is an awareness and readiness to relate to matters of textuality as an organizing principle underlying the cultural functioning of literary works of art.
In this paper, we discuss the potential of HyperCard for research and instruction in psychology. First we give a general overview of the HyperCard program; after that, we present two HyperCard stacks as sample solutions for two specific research applications. Surveyor, a self-contained survey tool, is a HyperCard-based vehicle for developing, administering, and processing tests and surveys. Queston demonstrates how HyperCard can be used as a data-management and data-analysis tool during the stages of questionnaire development. Both stacks illustrate how flexible HyperCard is and how easy it is to use it to manage, analyze, and process data, to transfer data to other programs, and to print reports. HyperCard, unlike traditional applications, gives the user a great degree of control over the way information is stored, mainipulated, and presented. Although both stacks are custom-made for specific purposes, the concepts underlying the design can be generally applied and adapted for other purposes.
We describe a set of color-graphics routines that permit implementing perceptual learning studies onIBM PCs. Stimulus items may be presented in 320×200 graphics mode, using 4 colors from a 16-color palette. There are 256 levels of stimulus clarification that may be either subject-paced, in which case the subject makes keypresses until the stimulus can be identified, or program-paced. The perceptual learning task may involve either mask clarification or dot clarification. A description of the present level of development of the procedures, and of several demonstration experiments, is presented.