Papers reviewed and determined not to be word norm studies. Use the flag icon to report errors or suggest re-inclusion.
18265 papers
Introduction / B. Boguraev and E. Briscoe -- Placing the dictionary on-line / Alshawi, B. Boguraev, and D. Carter -- An independent analysis of the LDOCE grammar coding system / E. Akkerman -- Utilising the LDOCE grammar codes / B Boguraev and E. Briscoe -- The derivation of a large computational lexicon o English from LDOCE / J. Carroll and G. Grover -- LDOCE and speech recognitio D. Carter -- Analysing the dictionary definitions / H. Alshawi -- Meaning an structure in dictionary definitions / P. Vossen, W. Meijs, and M. den Broede -- A tractable machine dictionary as a resource for computational semantics Y. Wilks... [et al.] -- Conclusion / B. Boguraev and E. Briscoe -- Appendic: Lexical database-user guide -- Semantic types of LDOCE verbs -- Dative alternations -- Lexicon development environment-user guide -- The Longman semantic codes -- The Longman grammar coding system.
ABSTRACT The purpose of this study was to contrast two methods of assessing commitment to equal opportunity (EO) goals. Students training at the Defense Equal Opportunity Management Institute (DEOMI) to be military EO advisors were the subjects of the study. The validity and reliability of DEOMI's measure of EO commitment, the Interpersonal Skills Development Evaluation (ISDE), were assessed. Slides were used to present cue words associated with various categories, including EO issues (e.g., discrimination and racism), to the DEOMI students. The students rated their association of these cue words with their current concerns and the emotional arousal evoked (two variables related to goal commitment; cf. Klinger, 1988). Students were then asked to recall as many of these words as they could. Although the ratings and the free-recall scores for the EO words were correlated with each other (support for the word-rating approach to measuring EO commitment), they were not significantly correlated with students' ISDE ratings. Other problems of validity and reliability for the ISDE are discussed.
Strict adherence to a low-fat diet without breast milk appears to be effective in treating infants with severe lipemia retinalis associated with exceptionally high triglycerides.
This manual addresses the linguistic issues that arise in connection with annotating texts by part of speech ("tagging"). Section 2 is an alphabetical list of the parts of speech encoded in the annotation systems of the Penn Treebank Project, along with their corresponding abbreviations ("tags") and some information concerning their definition. This section allows you to find an unfamiliar tag by looking up a familiar part of speech. Section 3 recapitulates the information in Section 2, but this time the information is alphabetically ordered by tags. This is the section to consult in order to find out what an unfamiliar tag means. Since the parts of speech are probably familiar to you from high school English, you should have little difficulty in assimilating the tags themselves. However, it is often quite difficult to decide which tag is appropriate in a particular context. The two sections 4 and 5 therefore include examples and guidelines on how to tag problematic cases. If you are uncertain about whether a given tag is correct or not, refer to these sections in order to ensure a consistently annotated text. Section 4 discusses parts of speech that are easily confused and gives guidelines on how to tag such cases, while Section 5 contains an alphabetical list of specific problematic words and collocations. Finally, Section 6 discusses some general tagging conventions. One general rule, however, is so important that we state it here. Many texts are not models of good prose, and some contain outright errors and slips of the pen. Do not be tempted to correct a tag to what it would be if the text were correct; rather, it is the incorrect word that should be tagged correctly.
The purpose of this study is to examin the relationships between the following three data sets:(1) evaluative ratings (5-point scale) on 10 most advancedscientific technologies (i. e., artificial intelligence, bio-technology, nuclear power generation, space technology, linear motor car, tube baby, 5th generation computer, super conductivity, organ transplant, and high-speed reactor), (2) familiarity ratings (5-point scale) on the same technologies, and (3) ratings on the subject's own personality (Y-G Personality Inventory). A hierarchical component analysis technique for the multiset data (Murakami, 1989) was applied. Three first order components for the image data sets were interpreted as indicating dangerous and harmful, useful and development, and personally beneficial dimensions, while the two first order components for the familiarity, as industrial technology and medical technology. Second order components showed several intersting relationships between image and familiarity which were found to be specific to particular scientific technologies. For example, industrial technologies high in the familiarity demension tended to have an image of usefulness and development.
This paper presents a specialized editor for a highly structured dictionary. The basic goal in building that editor was to provide an adequate tool to help lexicologists produce a valid and coherent dictionary on the basis of a linguistic theory. If we want valuable lexicons and grammars to achieve complex natural language processing, we must provide very powerful tools to help create and ensure the validity of such complex linguistic databases. Our most important task in building the editor was to define a set of coherence rules that could be computationally applied to ensure the validity of lexical entries. A customized interface for browsing and editing was also designed and implemented
To construct a data base (the "Penn Treebank") of written and transcribed spoken American English annotated with detailed grammatical structure. This data base will serve as a national resource, providing training material for a wide variety of approaches to automatic language acquisition, a reference standard for the rigorous evaluation of some components of natural language understanding systems, and a research tool for the investigation of the grammar and prosodic structure of naturally spoken English.
Research on the assessment of sexual orientation has been limited, and what does exist is often conflicting and confusing. This is largely due to the lack of any agreed upon definition of bisexuality. The Multidimensional Scale of Sexuality (MSS) was developed to validate and to contrast six proposed categories of bisexuality, as well as categories related to heterosexuality, homosexuality, and asexuality. This instrument includes ratings of the behavioral and cognitive/affective components of sexuality. The MSS was completed by 148 subjects, the majority of whom were from identified homosexual and bisexual populations. Although subjects' self-descriptions on the MSS were consistent with their self-descriptions on the Kinsey Heterosexual-Homosexual Scale, the MSS provided a more varied description of sexual orientation. Subject's self-described sexual orientation on the MSS was more consistent with their cognitive/affective ratings than with their behavioral ratings. With the exception of self-described heterosexuals, the frequency of cognitive/affective sexuality was greater than that of behavioral sexuality.
A program for the conversion of TIFF files of scanned images for display on IBM PCs is described. The program allows line drawings from various sources to be displayed in Turbo Pascal programs. The resultant picture files can be converted on a range of monitors, and the images displayed in different colors.
In this paper, we describe a computer-video interface capable of measuring videotaped motion without physical contact. It constructs imaginaryx-, y-, z-coordinates around a moving object by juxtaposing front and side views on one video record. A special effects generator (SEG) superimposes this two-part image onto the microcomputer’s graphic output. The user watches the SEG’s output and traces the motion of interest in real-time with a computer-generated cursor, a process that assesses the onset time, termination time, duration, linear displacement (i.e., length), and velocity of each traced movement. Reliability coefficients between two independent users were .99 for onset time, .85 for duration, .89 for displacement, and .90 for velocity. In 80 gestures, the correlation between estimated displacement and measured displacement was .73; the correlation between estimated duration (determined from frame-by-frame inspection of the videotape) and measured duration was .80. Twenty-five measurements of displacement for a motorized target traveling 75.4 in. differed from the true value by | .06 |, | .06 |, and | .30 | in. for thex, y, andz dimensions, respectively.
Ever since the pioneering work of Roberto Busa, computers have had an important work in humanistic research. The Istituto per le Scienze Religiose has been involved in such research since the 1970s. This article describes the Pope John XXIII project whose aim is to produce a computer-aided index of all the Pope's writings in both Italian and Latin; as well as future projects currently under consideration.
A cooperative team of researchers from various Italian universities are collaborating on a project for Latin lexicography. This article describes the linguistic work being done on the Latin language. The various problems — and their solutions — are discussed: allographs, homographs, source material and classification methods, word searching, etc.
This paper describes two research projects, both involving Latin and Greek lexicography. They are undertaken at the University of Florence and at the Italian National Research Council respectively. The one involves the creation of a Dictionary of Justinian's constitutions based on the emperor's legislative lexicon formed in the Corpus Iuris and elsewhere. The most demanding aspect of this task has been the creation of the Dictionary of the Novellae. The other project involves the creation of a Lexicon of the Novellae in the Authenticum version.
This article describes a second aspect of the Project for Latin Lexicography (see previous article). We here concentrate on two aspects of the project. First, we describe the morphological analyzer, which comprises a base dictionary, a table of suffixes, a table of endings and a table of postfixes. Second, we describe the lemmatization module, which operates by reference to a series of grammatical codes or information given for the base, and reference codes.
The teaching of literature through CAI raises problems of both a linguistic and instructional nature; student involvement and creativity in studying literature, and especially poetry, is difficult to build into a computer-based lesson. We have confronted these difficulties in the lessonPoetry I, which introduces undergraduates to basic concepts of poetic verse in a design using screen display, speech synthesis, and verse processing to maximize interactivity and student involvement.
The following is a description of a computerized version of the corpus of Latin grammarians published by Heinrich Keil in Leipzig between 1855 and 1880. The intent was to prepare an instrument which would serve both as a key to Keil's corpus and as the basis for a re-edition of the work itself. We discuss the corpus itself, the ways in which it was encoded, the pre-editing work, and how the material was organized for analysis by computer.
The research described in this paper was originally presented as a dissertation at the Università di Venezia. The objective was to establish the kind of language to which Chinese students are exposed during their primary education. The analysis was based on twelve texts of Guomin Xiaoxue Guoyu. This manuscript concentrates mainly on the techniques used to encode Chinese characters.
Several standard solutions have been developed for problems encountered when monitoring laboratory response sensors with microcomputers. These problems include voltage-level translation, power-supply isolation, contact debouncing, and temporary data storage. Data storage has typically been accomplished through the use of edge-sensitive devices, which introduce difficulties when a researcher wishes to monitor the duration of a response. A circuit is described that includes an optoisolator with a silicon-controlled rectifier at its output stage. The circuit acts as a latch for the detection of brief responses, and, because it is not edge-sensitive, it can also be used to record the duration of sustained responses.
A research paradigm is suggested that combines the perspectives of the humanistic scholar and the behavioral scientist: After differentiating the popularity of actual aesthetic products using archival indices and then subjecting these compositions to objective computer content analyses, further statistical treatment may divulge the intrinsic properties responsible for differences in impact. This approach is illustrated by an analysis of the 154 sonnets attributed to William Shakespeare. Each sonnet was partitioned into four consecutive units (three quatrains and a couplet), and then a computer gauged how the number of words, different words, unique words, primary process imagery, and secondary process imagery changed within each sonnet. Taking advantage of a previous objective measure of the relative aesthetic merit of the sonnets, and implementing a statistical search for interaction effects, it was demonstrated that Shakespeare's lexical choices adopt a discernible pattern in the highly popular creations that is not found in the more obscure poems. Perhaps the most fascinating aspect of this pattern shift is the distinct manner in which the poet modifies his vocabulary when composing the concluding couplet in his best sonnets.
This manuscript is a description of the research activities of the Lessico Intellettuale Europeo. The Centre works on the lexicographical analysis of philosophical and scientific texts, mainly of the seventeenth and eighteenth centuries. An important project is the Philosophical Dictionary of the Seventeenth and Eighteenth Centuries. The LIE series of publications include a number of indices, concordances and lexicons. Another project at the Centre is the Thesaurus Mediae et Recentioris Latinitatis. The Centre also publishes the proceedings of the three-yearly Colloqui internazionali.
CorrecText, from Houghton-Mifflin, is a significant advance in grammar checkers, because it uses a full parse of sentences in its analysis. Though limited by the fact that many English sentences are syntactically ambiguous, the program can find many errors in grammar, style, and usage. Still, questions remain about where and how it can be useful. It is not terribly useful on final versions of good prose; the makers assert that it is more useful on unedited prose, where there are many inadvertent errors. An empirical study of much unedited prose would not only verify this assertion, but would help improve the program.
Machine-readable texts in the humanities are in a period of rapid growth, as are programs for analyzing them, with concomitant problems of finding and choosing the most suitable of each. No one text-base, and no one search program, proves to be wholly suitable for all projects. Three of the currently available electronic editions of Shakespeare are discussed and compared, along with three commercial programs for text analysis on microcomputers.
Techniques are described for using an IBM PC equipped with a mouse to investigate the reading of character sequences of maximally one display line. Solutions are given to problems that arise and derive from the PC’s real-time clock, the slowness of the display, and registration of input during a test session. A program for using the techniques to study asynchronous perception of printed words is described, and a demonstration program is provided as an appendix.
Examined the extent to which phonological performance varied as a function of test–client congruence on 3 tests of articulation containing standard English assumptions among 10 African-American children (aged 5 yrs 11 mo to 6 yrs 11 mo) who spoke 'Black English vernacular.' Results suggest that African-American children perform differently on standardized tests as a function of the linguistic norms used to score items. A failure to take the issue of dialect variation into account substantially increased the likelihood of misdiagnosing normally speaking African-American children as having articulation disorders. (PsycINFO Database Record (c) 2016 APA, all rights reserved)
CHIP is a computer program for the automatic coding and analysis of parent-child conversational interaction. The program was developed because manual coding of large collections of computerized transcripts cannot possibly be completed within a reasonable time span. CHIP codes parent-child conversational data as stored in transcript files and computes a series of descriptive statistics based on these codes. Parental responses to child utterances and child responses to parent utterances are both coded. This allows for an analysis of the reciprocal relationship between parental and child language. Three longitudinal corpora from CHILDES (totaling 151,900 utterances) were coded and analyzed by CHIP. The results indicated a high degree of contingency between parental and child language for different word classes across a large span of development. Two main points are argued: (1) Automatic data coding and analysis programs are important new tools for transcript analysis, and (2) CHIP, as an example of such a tool, can provide detailed information concerning the exact nature of parent-child conversational interactions.
We describe a set of color-graphics routines that permit implementing perceptual learning studies onIBM PCs. Stimulus items may be presented in 320×200 graphics mode, using 4 colors from a 16-color palette. There are 256 levels of stimulus clarification that may be either subject-paced, in which case the subject makes keypresses until the stimulus can be identified, or program-paced. The perceptual learning task may involve either mask clarification or dot clarification. A description of the present level of development of the procedures, and of several demonstration experiments, is presented.
The study of the history of new words in theNewOED described in this paper was undertaken in 1986-87, and is based on the material then available. Since then, theNewOED has been finished, and PAT, the inquiry system developed at the University of Waterloo for the investigation of theNewOED data base, has been much altered and improved. Nevertheless, this report should prove useful in indicating the potentiality for analyzing the computerizedNewOED and some of the problems. This project is a study of the ways in which new words are created in English at various periods of time. A chronological dictionary 's created listing words introduced into the language over 50 year increments. These words are then classified by the processes used in forming them to show, in proportional terms, if certain processes are more common at some times than at others.
An interface circuit to connect a microphone to an Apple Macintosh computer is described. The Apple Macintosh mouse port is used as the input port, and the microphone activation simulates a mouse press.
Based on the ARTFL version of theProfession and excerpts fromEmile, high frequency function and content words, as defined by Brunet, are analyzed via Pearson chi square tests. Next, four measures of narrative voice from the same populations are compared using Markovian chains and further chi square tests. In a third analysis the two orders of evidence are juxtaposed. The lexical and narratological preferences of theVicaire and theGouverneur, while not resolving the problematic of chronological composition (Burgelin, 1969), highlight the distinctiveness of each character.
The purpose of this paper is to show the utility of the application of a non-ultrametric tree-model to textual data. The first part introduces a basic topological property of the tree and the notion of neighbourhood, which reflects the structure of the tree. The second part emphasizes through illustration examples the adequacy of this model for representing different varieties of textual data.
The Société de 1789 was a political club founded in early 1790 to propagate the ideals of the Revolution and the Enlightenment. A systematic analysis of the language found in the public discourse of the Société using simple quantitative techniques suggests important distinctions in comparison to the language found in a baseline sample, a selection of the General Cahiers de doléances of 1789. It is further argued that these differences represent an Enlightened reforming tradition that carried into the French Revolution.
The present paper is a critique of quantitative studies of literature. It is argued that such studies are involved in an act of reification, in which, moreover, fundamental ingredients of the texts, e.g. their (highly important) range of figurative meanings, are eliminated from the analysis. Instead a concentration on lower levels of linguistic organization, such as grammar and lexis, may be observed, in spite of the fact that these are often the least relevant aspects of the text. In doing so, quantitative studies of literature significantly reduce not only the cultural value of texts, but also the generalizability of its own findings. What is needed, therefore, is an awareness and readiness to relate to matters of textuality as an organizing principle underlying the cultural functioning of literary works of art.
This paper suggests ways in which the pattern-matching capability of the computer can be used to further our understanding of stylized ballad language. The study is based upon a computer-aided analysis of the entire 595,000- word corpus of Francis James Child'sThe English and Scottish Popular Ballads (1882–1892), a collection of 305 textual traditions, most of which are represented by a variety of texts. The paper focuses on the “Mary Hamilton” tradition as a means of discussing the function of phatic language in the ballad genre and the significance of textual variation.
A key word with regard to a sub-corpus is a word of which the frequency in that sub-corpus is significantly higher than expected under the hypothesis that its use and the variable “part of the corpus” are mutually independent. A study in literary statistics almost invariably includes a chapter devoted to key words. However, a strong attack has been recently launched upon the way stylometry has been modelling texts since the classical works of Herdan, Guiraud or Muller. In fact statistical modelling seems as valid in stylistics as in any other field of the humanities and social sciences. What is questionable is the fact that many studies in literary statistics are more satisfied with the easy identification of monsters, i.e. literary phenomena unexplained by wrong models, than with the laborious research of models fitting the textual data well. A short examination of the mentioned controversy and the quantitative analysis of an example provided by Laclos' novelLes Liaisons dangereuses endeavour to support this argument.
This article describes some approaches to imitation analysis and the use of ready-made software for this task. Devising computer-assisted techniques for exploring the conscious literary imitation of style is an application of particular relevance to contemporary Hispanic narrative and one that can be handled with a microcomputer and readily accessible software. The article describes some approaches to imitation analysis and the use of ready-made software to assess the effectiveness of stylistic imitation of eighteenth-century historical chronicle in La renuncia del héroe Baltasar (The Renunciation of the Hero Baltasar), by the Puerto Rican novelist Rodríguez Juliá. Even when employing familiar procedures of text analysis with computer, comparing a fictional text with a multiple and diverse corpus of authentic historical documents requires somewhat unique assumptions and hypotheses, since neither authorship, influence, or authenticity are in question.
Research on changes in Shaw's rhetoric inMrs. Warren's Profession, Major Barbara, andHeartbreak House led me to a heuristic for gaining literary critical control over computer output. This essay describes the eleven-step process: stepping away from the data, stating first premises, developing a working hypothesis, classifying computer-sorted data, marking implicit literary sub-structures, collecting sub-structural data into tables, applying earlier statistical observations, choosing parts for detailed analysis, designing a visual method for representing the analysis, presenting segment by segment analysis of the selected data, and making larger descriptive generalizations. While describing this heuristic, the essay also reports on the Shaw research.