Papers reviewed and determined not to be word norm studies. Use the flag icon to report errors or suggest re-inclusion.
16504 papers
In studies of author attribution, measurement of differential use of function words is the most common procedure, though lexical statistics are often used. Content analysis has seldom been employed. We compare the success of lexical statistics, content analysis, and function words in classifying the 12 disputedFederalist papers. Of course, Mosteller and Wallace (1964) have presented overwhelming evidence that all 12 were by James Madison rather than by Alexander Hamilton. Our purpose is not to challenge these attributions but rather to useThe Federalist as a test case. We found lexical statistics to be of no use in classifying the disputed papers. Using both classical canonical discriminant analysis and a neural-network approach, content analytic measures — the Harvard III Psychosociological Dictionary and semantic differential indices — were found to be successful at attributing most of the disputed papers to Madison. However, a function-word approach is more successful. We argue that content analysis can be useful in cases where the function-word approach does not yield compelling conclusions and, perhaps, in preliminary screening in cases where there are a large number of possible authors.
The field of natural language processing (NLP) has seen a dramatic shift in both research direction and methodology in the past several years. In the past, most work in computational linguistics tended to focus on purely symbolic methods. Recently, more and more work is shifting toward hybrid methods that combine new empirical corpus-based methods, including the use of probabilistic and information-theoretic techniques, with traditional symbolic methods. This work is made possible by the recent availability of linguistic databases that add rich linguistic annotation to corpora of natural language text. Already, these methods have led to a dramatic improvement in the performance of a variety of NLP systems with similar improvement likely in the coming years. This paper focuses on these trends, surveying in particular three areas of recent progress: part-of-speech tagging, stochastic parsing, and lexical semantics.
One experiment compared the effect of elaboration on enacted and non-enacted events. The commands were either presented in a basic form (e.g., "wave your hands") or in an enriched form. The commands were enriched by adding statements to the commands of how to perform the actions (e.g., "wave your hands as a conductor"). Free- and cued-recall data showed elaboration to have a dissociative effect on enacted and non-enacted events. Memory for the non-enacted events benefited from enrichment, whereas simple enacted events were remembered to a higher extent than complex enacted events. Lack of benefit from elaboration on memory of enacted events is suggested to be due to enactment leading to a sufficient degree of item-specific processing, and a negative effect of elaboration is suggested to occur when the way of manipulating item complexity decreases the familiarity of the actions. Familiarity ratings of the items by two independent groups of subjects supported this interpretation.
There are currently two philosophies for building grammars and parsers – Statistically induced grammars and Wide-coverage grammars. One way to combine the strengths of both approaches is to have a wide-coverage grammar with a heuristic component which is domain independent but whose contribution is tuned to particular domains. In this paper, we discuss a three-stage approach to disambiguation in the context of a lexicalized grammar, using a variety of domain independent heuristic techniques. We present a training algorithm which uses hand-bracketed treebank parses to set the weights of these heuristics. We compare the performance of our grammar against the performance of the IBM statistical grammar, using both untrained and trained weights for the heuristics. 1
Research based on a treebank is active for many natural language applications. However, the work to build a large scale treebank is laborious and tedious. This paper proposes a probabilistic chunker to help the development of a partially bracketed corpus. The chunker partitions the part-of-speech sequence into segments called chunks. Rather than using a treebank as our training corpus, a corpus which is tagged with part-of-speech information only is used. The experimental results show the probabilistic chunker has more than 92% correct rate in outside test. The well-formed partially bracketed corpus is a milestone in the development of a treebank. Besides, the simple but effective chunker can also be applied to many natural language applications.
Syntactic natural language parsers have shown themselves to be inadequate for processing highly-ambiguous large-vocabulary text, as is evidenced by their poor performance on domains like the Wall Street Journal, and by the movement away from parsing-based approaches to text-processing in general. In this paper, I describe SPATTER, a statistical parser based on decision-tree learning techniques which constructs a complete parse for every sentence and achieves accuracy rates far better than any published result. This work is based on the following premises: (1) grammars are too complex and detailed to develop manually for most interesting domains; (2) parsing models must rely heavily on lexical and contextual information to analyze sentences accurately; and (3) existing {$n$}-gram modeling techniques are inadequate for parsing models. In experiments comparing SPATTER with IBM's computer manuals parser, SPATTER significantly outperforms the grammar-based parser. Evaluating SPATTER against the Penn Treebank Wall Street Journal corpus using the PARSEVAL measures, SPATTER achieves 86\% precision, 86\% recall, and 1.3 crossing brackets per sentence for sentences of 40 words or less, and 91\% precision, 90\% recall, and 0.5 crossing brackets for sentences between 10 and 20 words in length.
The Robert Electronique is the CD-ROM version of the nine volume Grand Robert, roughly the French equivalent of the OED. This article outlines a project to produce some learning materials using this lexical database. It describes various types of exercises ranging from the semantic to the stylistic. Most of the exercises can be completed on screen within a wordprocessing application and can be done by a student working independently. The activities exploit as far as possible features specific to the CD-ROM version of the dictionary.
This paper chronicles the work of the TEI textual criticism working groups through several phases, documenting how and why the design goals were shaped by the requirements of several distinct user communities and by the nature of the textual evidence itself. Encoding schemes for the representation of physical details of textual witnesses were unified with encoding schemes for critical editing practices when it was observed that the two phenomena were inextricably layered and linked within real texts. Rationale is offered for the development teams' adherence to exceedingly general design principles: (a) the requirement that the encoding notations be neutral in text-theoretic terms; (b) the need to accommodate dramatically different text-transmission phenomena and research goals within diverse text-critical arenas; (c) the need for commensurability of the text-critical markup with encoding notations used in closely related text-analytic research. The paper also assesses the results of the effort in terms of the encoding scheme's adequacy for several scholarly purposes: suggestions are made concerning the need for programmatic testing, for refinement, and for extension of the encoding model to support a broader range of text-transmission phenomena and research objectives.
ABSTRACT: This paper analyzes the data from three questionnaires administered to Pakistani male and female journalists, teachers, and university students in Islamabad, Karachi, and Lahore during a period from 1987 to 1992. The first questionnaire deals with respondents’(320) choice of a model of English (British, American, or Pakistani). The second and third questionnaires measure the acceptability of selected Pakistani English lexical and grammatical items (150 respondents) and complementation types (165 respondents). Results show that while an exonormative model of English (British) still has considerable influence in the former colony (in both ‘ideal’ as well as in reported ‘actual’ usage), a Pakistani norm is also beginning to emerge. This trend is most evident in respondents’ acceptance of typically Pakistani features of English such as Urdu borrowings, Urdu‐English hybrids, and local morphological and syntactic innovations.
Another current major issue in lexical semantics is the definition and the construction of real-size lexical databases that will be used by parsers and generators in conjunction with a grammatical system. Word meaning, terminological knowledge representation and extraction of knowledge in machine readable dictionaries are the main topics addressed. They really represent the backbone of a lexical semantics knowledge base construction.
Adult children of alcoholics' (n = 68) perceptions of their relationships with parents were compared with those of a control sample (n = 37) to examine independent and joint influences of interpersonal status and affect on family dynamics. Visual metaphors for relationships using circle drawings and a status-affect rating scale from the Grasha-Ichiyama Psychological Size and Distance Scale were employed. Compared with the control group, adult children of alcoholics drew smaller circles to represent themselves, i.e., indicating less interpersonal status, only when assessing their relationships with their fathers. Analyses of status-affect ratings showed that the drawings of smaller circles reflected feeling less competent, i.e., having less personal knowledge and expertise, rather than perceptions of being submissive in the relationship. The distance drawn between the circles of adult children of alcoholics and their parents, i.e., psychological distance, was much larger than that of the control group. Ratings showed that perceptions of a negative emotional climate and submissiveness together accounted for 25% of the unique variance in predicting psychological distance. Perceptions of being submissive, however, were not associated with perceptions of psychological distance among adult children of nonalcoholic parents.
The purpose of this study was to examine the experience with attitudes towards, and knowledge about homosexuality of three groups of health care professionals. Subjects were 97 registeres nurses, social workers, and psychologists who responded to a six-page mailed questionnaire. Professional discipline of the subject, gender of the client, and gender of the client's lover in a fictitious scenario did not significantly affect ratings or suggested diagnoses of the client. Most subjects felt that they needed moer training in working with homosexual clients, which was consistent with their high but not perfect scores on a knowledge test. Subject's mean scores of the Attitudes Toward Lesbians (ATL) and Gay Men (ATGM) scales of Herek (1998) reflected significantly less prejudice than his college samples. More knowledgeable respondents were less prejudiced and had more positive attitudes about working with gay and lesbian clients; those with more positive attitudes toward clients also showed less prejudice on the ATL and ATGM scales. The authors argue that training health care professionals to be more knowledgeable about gay and lesbian issues would lead to more positive attitudes and better services for gay and lesbian clients.
The concept of class is studied by social psychologists as an important determinant in the development and expression of attitudes and behaviors (Schaefer, 1986). This dissertations explores existing perspectives of class including the theoretical treatments of experts and non-experts. The empirical use of the class concept is considered and past use is described as lacking some empirical basis. Study One serves to define the dimensions of the class construct as seen by non-experts. Thirty-two introductory psychology students rated eight class labels on fifty bi-polar adjectives. Multi-dimensional scaling revealed a two dimensional structure for class with ratings based on both a stereotypical hierarchical social class dimension and on a weaken ingroup bias evaluative dimension. Study Two serves to identify distinguishable class labels as well as differentiating characteristics of those labels. This allows for designation of characteristics or descriptions which are useful for social psychological research and have an empirical base. One hundred eighteen Introductory psychology students rated one class label on 52 bi-polar adjectives. Results reveal three distinct groupings, roughly corresponding to traditional upper, middle, and lower class delineation. However, working class and middle class are rated equivalently rather than as a dichotomy. The rating pattern replicates Study One by following either a traditional stereotypical delineation or self-identified class labels as more positive. In Study Three, class bias is explored using empirically designated descriptors by acquiring ratings for comparison between the two extreme class categories; upper and lower class. Other variables of interest are the valence of the description and presence or not of the class label. The question of class bias as a process distinct from race bias is explored by use of rating comparisons between class and race label-only stimuli. Responses from 246 introductory psychology student volunteers reveal that valence of the stimulus is a primary elicitor of reaction. They also indicate that class and an overt display of class membership via presence of a label affect ratings. Class bias is expressed overtly while racial bias is expressed subtlely.
There are currently two philosophies for building grammars and parsers -- Statistically induced grammars and Wide-coverage grammars. One way to combine the strengths of both approaches is to have a wide-coverage grammar with a heuristic component which is domain independent but whose contribution is tuned to particular domains. In this paper, we discuss a three-stage approach to disambiguation in the context of a lexicalized grammar, using a variety of domain independent heuristic techniques. We present a training algorithm which uses hand-bracketed treebank parses to set the weights of these heuristics. We compare the performance of our grammar against the performance of the IBM statistical grammar, using both untrained and trained weights for the heuristics.
In this paper I report on the development of an application in which HTML forms serve as a front-end to a lexical database.Lexical information and data retrieval strategies are based on the Longman Language Activator.A Visual Basic CGI application connects a front end HTML form with the back-end relational database implemented using Microsoft Access.Three aspects of the applcation are discussed in this paper: (1) the lexical database; (2) the HTML front end; and (3) the Visual Basic CGI programming necessary to connect (1) and ( 2). The Lexical DatabaseThe source for the lexical database implemented for this application is the Longman Language Activator, (LLA), subtitled The World's First Production Dictionary.The LLA was selected for its unique organization of dictionary information which is intended to make it easier for the user to find the right word or phrase for a particular context.The LLA is organized on one level according to 'Key Words' or concepts and on another level by the words and phrases which realize these concepts.The 1052 Key Words in the LLA are said to account for the basic concepts from the core of English.For each Key Word, there is what is called a 'Meaning Menu' consisting of numbered sections corresponding to major aspects of the Key Word.The numbered sections follow.Each with another `menu' of words and phrases.Further detail about each word/phrase -pronunciation, part-of-speech, definition, examples -
Lexical Collocations are frequently occurring word pairs in natural language whose presence are not always predictable by their usage. These collocations are used by native speakers of a language almost without thought; yet they must be learned by non-native speakers of that language. A native speaker of English may drink strong coffee while a non-native speaker may say either $\sp{*}$powerful coffee or $\sp{*}$sturdy coffee. Collocations tend to vary among languages and topic domains. Unfortunately, the task of correctly identifying lexical collocations, even by native speakers of that language, has been shown to be very difficult. Computer systems that translate natural languages, or Machine Translation (MT) systems, need to know about lexical collocation information in order to produce natural sounding or colloquially proper text. Natural Language Generation (NLG) is a component of an MT system which automatically produces natural sounding text in a particular target language given a language-independent meaning as input. This dissertation will demonstrate how to automatically locate and extract lexical collocations from machine-readable text for use within an MT system's NLG component. A lexical-semantic and statistical approach is adopted for the location and extraction of lexical collocations. For this approach, a computational definition is provided for lexical collocations which demonstrates that: (1) they occur as adjacent word pairs; (2) they occur more often than would be expected by chance; and (3) they comprise words for which neither word may be substituted by a synonym or hyponym. Potential collocations comprising certain adjacent part-of-speech tags are extracted from text. An on-line thesaurus and lexical database of word classes are queried for synonyms and hyponyms, respectively, for each potential collocation. These queries create potential challenger pairs, such as strong java and powerful coffee. A substitution procedure is then applied to determine if any of these challenging word pairs occur more frequently than the potential collocation. The VERIFY lexical collocation extraction system has been implemented incorporating these ideas. Results to date have been positive: using lexical-semantic knowledge, i.e., synonymy and hyponymy, within a lexical collocation extraction system outperforms a system using purely statistical knowledge. In order to compare system output to human judgments of training data, a training component was also incorporated into VERIFY. This component is able to adapt to new data. Overall system performance, measured by Recall and Precision scores, was shown to improve using this component. In order to provide a more flexible system given a user's application, a weighting mechanism was used to produce a range of Recall and Precision scores. These weights can be 'adjusted' to optimize system performance. The use of lexical-semantic knowledge has advanced the state of the art for lexical collocation extraction beyond traditional statistical approaches. Incorporation of a training component within an extraction system provides the capability of adapting to any changes within the data. Controlling overall system performance through the use of a weighting mechanism provides flexibility to the user of an extraction system. And, in an experiment to compare VERIFY'S performance to that of human performance on a particular set of data, it was shown that VERIFY outperforms humans in both Recall and Precision.
The relationship between positive and negative events and emotional well-being for depressed and nondepressed residents of a nursing home and congregate housing care facility was examined. For 30 consecutive working days, each of 79 participants was presented with the Philadelphia Geriatric Center Positive and Negative Affect rating scales. Events during the previous 24 hr were elicited by an open-ended format. Results indicated that variations in daily events (e.g., health, family, self-initiated, and social events) were related to residents' affect, and there was congruence between mood and event valence when the effects of psychopathology and residence were removed. Thus, regardless of diagnosis or residential setting, people's moods showed a relationship to the quality of daily events. Findings also indicated that ratings of residents' affect could be translated into audits for institutional quality.
Syntactic natural language parsers have shown themselves to be inadequate for processing highly-ambiguous large-vocabulary text, as is evidenced by their poor performance on domains like the Wall Street Journal, and by the movement away from parsing-based approaches to text-processing in general. In this paper, I describe SPATTER, a statistical parser based on decision-tree learning techniques which constructs a complete parse for every sentence and achieves accuracy rates far better than any published result. This work is based on the following premises: (1) grammars are too complex and detailed to develop manually for most interesting domains; (2) parsing models must rely heavily on lexical and contextual information to analyze sentences accurately; and (3) existing n-gram modeling techniques are inadequate for parsing models. In experiments comparing SPATTER with IBM's computer manuals parser, SPATTER significantly outperforms the grammar-based parser. Evaluating SPATTER against the Penn Treebank Wall Street Journal corpus using the PARSEVAL measures, SPATTER achieves 86% precision, 86% recall, and 1.3 crossing brackets per sentence for sentences of 40 words or less, and 91% precision, 90% recall, and 0.5 crossing brackets for sentences between 10 and 20 words in length.
espanolEn castellano, la silaba parece actuar como un elemento prelexico-fonologico de relacion con el nivel lexico. Su mayor o menor frecuencia determina la cantidad de palabras que se activaran en el nivel lexico. Esta cualidad de restriccion lexica nos ha sugerido la necesidad de elaborar un estudio normativo en el cual los sujetos evocaban palabras de 2 y 3 silabas a partir de una inicial dada que despues pueden ser utilizados como base para diversos estudios experimentales. Se produjeron asi 130 conjuntos de candidatos competidores lexicos (ccl), que proporcionan informacion sobre dos aspectos fundamentales: su tamano (numero de candidatos lexicos) y la accesibilidad relativa de cada una de las salidas lexicas que las componen. EnglishThe syllable in Spanish could operate as a phonological and prelexical unit with relation to lexical level. The syllable frequency determines the number of words that will be activated at lexical level. This quality of accessibility constriction suggests the need to produce some candidate set norms. In this normative study, the subjects recover two and three syllable words from one initial syllable given. In this way, we obtained 130 sets of lexical competitor candidates (CCL), which provide information about two basic aspects: the set size (i.e. number of lexical candidates), and relative accessibility of each lexical output composing the set.
Mechanisms of hypnotic analgesia were investigated by examining changes in the R-III, a nociceptive spinal reflex, during hypnotic reduction of pain sensation and unpleasantness. The R-III was measured in 15 healthy volunteers who gave VAS-sensory and VAS-affective ratings of an electrical stimulus during conditions of resting wakefulness, suggestions for hypnotic analgesia, and attempted suppression of the reflex during non-hypnotic conditions. The H-reflex was also measured to monitor and control for general changes in alpha-motoneuron excitability. Hypnotic sensory analgesia was related to reduction in the R-III after controlling for changes in the H-reflex (R2 = 0.51, P < 0.003), suggesting that hypnotic sensory analgesia is at least in part mediated by descending antinociceptive mechanisms that exert control at spinal levels in response to hypnotic suggestion. The relationship between hypnotic affective analgesia and reduction in R-III approached significance (R2 = 0.26; P = 0.053). Reduction in R-III was 67% as great and accounted for 51% of the variance in reduction of pain sensation. In turn, reduction in pain sensation was 75% as great and accounted for 77% of the variance in reduction of unpleasantness. The results suggest that 3 general mechanisms may be involved in hypnotic analgesia. The first, implicated by reductions in R-III, is related to spinal cord antinociceptive mechanisms. The second, implicated by reductions in pain sensation over and beyond reductions in R-III, may be related to brain mechanisms that serve to prevent awareness of pain once nociception has reached higher centers, as suggested by Hilgard.(ABSTRACT TRUNCATED AT 250 WORDS)
An extragrammatical sentence is what a normal parser fails to analyze. It is important to recover it using only syntactic information although results of recovery are better if semantic factors are considered. A general algorithm for least-errors recognition, which is based only on syntactic information, was proposed by G. Lyon to deal with the extragrammaticality. We extended this algorithm to recover extragrammatical sentence into grammatical one in running text. Our robust parser with recovery mechanism -- extended general algorithm for least-errors recognition -- can be easily scaled up and modified because it utilize only syntactic information. To upgrade this robust parser we proposed heuristics through the analysis on the Penn treebank corpus. The experimental result shows 68% ¸ 77% accuracy in error recovery. 1 Introduction Extragrammatical sentences include patently ungrammatical constructions as well as utterances that may be grammatically acceptable but are beyond the synta...
The present study examined factors hypothesized to influence mental health professionals' perceptions of dangerousness, predictions of violence, and decisions on patients' release. 120 mental health professionals employed in state mental hospitals were each given one of 12 patient profiles. The independent variables, manipulated within vignettes, were (a) violence history, (b) paranoid schizophrenia versus nonparanoid schizophrenia, and (c) perceived consequences in terms of liability and publicity. Type of schizophrenia did not affect ratings, but violence history of the predictee and perceived consequences to the predictor did significantly influence the ratings. Patients with actual violence histories were viewed by the subjects as having more potential for future violence, as being more globally dangerous, and as requiring a more secure placement than those with histories of threats of violence or no violence. Possible litigation following release led to a recommendation for more secure placement than did minimal legal consequences. Predictions of violence and decisions on hospital release were interpreted as dependent on both predictor and patient-related variables.
The University of Delaware and the University of Dundee are collaborating on a project that is investigating the application of spatialization and spatial metaphors to interfaces for Augmentative and Alternative Communication. This paper outlines the project's motivation, goals, and methodological considerations. It presents a number of design principles obtained from a review of the HCI literature. Finally, it describes progress on the demonstration of this approach. This application called VAL provides a computer-based word board that retains spatial equivalence to the user's paper-based system. It also allows the user to access an extended lexicon through an interface to the WordNet lexical database.
The aim of this article is to help to clarify the highly complicated and vexed issue of the use of different linguistic varieties in spoken Czech. The term 'varieties' refers both to distinctive language types and to hybrid forms of expression, which are determined by the social characteristics of the speaker and by the context of a particular utterance or series of remarks. The article is intended primarily for non-native speakers, to whom deviations from the standard, somewhat literary-sounding form of Czech, known as spisovna ceStina, can pose quite considerable difficulties. It is accepted implicitly that all scholars who wish to appreciate something of the richness and diversity of spoken Czech must at the very least learn to recognize the characteristic morphological, phonological and lexical differences between spisovna ceStina and the more informal styles of speech widely used for the purpose of everyday communication. The question of the co-existence of different linguistic varieties is not of paramount importance to the majority of native speakers. Most adult Czechs have little need to reflect consciously on the way they express themselves since they have learnt through extensive practical experience to adopt language styles which they (rightly or wrongly) consider to be appropriate in given circumstances. The same, however, cannot be said of Czech children or of foreigners who can sometimes feel ill-at-ease in unfamilar settings and ill-equipped linguistically to deal with the demands made of them in new social surroundings. All teachers of Czech (whether teaching Czech to native speakers or as a foreign language) are aware of the difficulties posed by the co-occurrence and inter-relationship of different styles of speech. Yet there is a considerable divergence of opinion amongst scholars with respect to the status, functions and definitions of the various non-standard forms of the spoken language. Whilst all scholars recognize that spisovna cestina is inherently more stable than any other form of Czech, they remain divided over the precise role and characteristics of the different types of language used in everyday conversation. In all languages non-literary forms of expression are prone to fluctuation and inevitably change much more quickly than the codified norms. Not only is there a great deal of variability within the verbal repertoire of a given community, but even individual speakers can show considerable inconsistency in their use of language, in particular with respect to lexical items. Czechs frequently combine
We describe a series of three experiments in which supervised learning techniques were used to acquire three different types of grammars for English news stories. The acquired grammar types were: 1) context-free, 2) context-dependent, and 3) probabilistic context-free. Training data were derived from University of Pennsylvania Treebank parses of 50 Wall Street Journal articles. In each case, the system started with essentially no grammatical knowledge, and learned a set of grammar rules exclusively from the training data. Performance for each grammar type was then evaluated on an independent set of test sentences using Parseval, a standard measure of parsing accuracy. These experimental results yield a direct quantitative comparison between each of the three methods.
Computer technology is applied extensively to biblical studies and many programs are available for all different categories of people interested in the Bible. This paper argues that text databases which offer additional features, such as links to morphological analyses or lexica, should influence the teaching of biblical languages. However, linguistic information in these databases is at present very limited and should be expanded to include the full spectrum of linguistic knowledge. Linguistic databases, developed specifically for researchers, should be integrated with concordance software, online grammars, etc., as well as cultural-historical material, to create a comprehensive biblical information system. Existing information should be included in such a system, but in converting paper documents to electronic publications, value should be added to the products by means of creating sophisticated methods of access and by integrating the material with other sources. Problems of compatibility and standardization are also briefly addressed.
The Documentation Project is a cooperative project between Faculties of Arts in the Norwegian universities. It aims to produce “the Norwegian universities' databases for language and culture” from the paper-based archives at the participating institutions. The project has been active on a national basis since 1992. This paper describes the methodologies involved and ongoing subprojects.
How educators and researchers define and study school effectiveness continues to be shaped by two divided camps. The policy mechanics attempt to identify particular school inputs, including discrete teaching practices, that raise student achievement. They seek universal remedies that can be manipulated by central agencies and assume that the same instructional materials and pedagogical practices hold constant meaning in the eyes of teachers and children across diverse cultural settings. In contrast, the classroom culturalists focus on the implicitly modeled norms exercised in the classroom and how children are socialized to accept particular rules of participation and authority, linguistic norms, orientations toward achievement, and conceptions of merit and status. It is the culturally constructed meanings attached to instructional tools and pedagogy that sustain this socialization process, not the material character of school inputs per se. This article reviews how these two paths of school-effects research are informed by work conducted within developing countries. First, we discuss the school’s aggregate effect, relative to family background, within impoverished settings. Second, we review recent empirical findings from the Third World on achievement effects from discrete school inputs. An emerging extension of this work also is reviewed: How input effects are conditioned by the social rules of classrooms. Third, we illustrate how future work in the policy-mechanic tradition will be fruitless until cultural conditions are taken into account. And the classroom culturalists may reach a theoretical dead end until they can empirically link classroom processes to alleged effects. We put forward a culturally situated model of school effectiveness—the implications of which are discussed for studying ethnically diverse schools within the West. By bringing together the strengths of these two intellectual camps, researchers can more carefully condition their search for school effects.
This paper presents a method for constructing deterministic Prolog parsers from corpora of parsed sentences. Our approach uses recent machine learning methods for inducing Prolog rules from examples (inductive logic programming). We discuss several advantages of this method compared to recent statistical methods and present results on learning complete parsers from portions of the ATIS corpus. Introduction Recent approaches to constructing robust parsers from corpora primarily use statistical and probabilistic methods such as stochastic context-free grammars (Black et al., 1992; Pereira and Schabes, 1992). Although several current methods learn some symbolic structures such as decision trees (Black et al., 1993) and transformations (Brill, 1993), statistical methods still dominate. In this paper, we present a method that uses recent techniques in machine learning to construct symbolic, deterministic parsers from parsed corpora (treebanks). Specifically, our approach is implemented in...
Examining the complex relationship between law and language enhances our understanding of the marginalization and subordination of linguistic Outsiders. This nexus between law and language has many manifestations. In this essay I discuss the biases about language that constrain traditional legal discourse while I explore strategies for its reframing by using the languages of Outsiders. Succinctly stated, this essay posits that traditional language norms create images or maintain stereotypes that stultify public discourse as well as impose cultural integration and linguistic assimilation with destructive consequences. The essay proposes that linguistic norms in law schools can be refashioned through pedagogical innovations to minimize their subordinating effects.
I review evidence for the claim that syntactic ambiguities are resolved on the basis of the meaning of the competing analyses, not their structure. I identify a collection of ambiguities that do not yet have a meaning-based account and propose one which is based on the interaction of discourse and grammatical function. I provide evidence for my proposal by examining statistical properties of the Penn Treebank of syntactically annotated text.
The functional role of counterfactual thoughts ("might have been" reconstructions of the past) was explored in three laboratory experiments. Specifically, counterfactual thoughts were posited to serve two possible functions: an affective function (feeling better) and a preparative function (preparing for the future via avoiding the recurrence of negative events). It is argued that counterfactuals as a generic class of cognitions generally serve these two functions, but that specific types of counterfactuals may in particular do so. Two dimensions are described alone which counterfactuals may be classified: direction (upward vs downward) and structure (additive vs subtractive). Upward counterfactuals focus on an alternative that is better than reality, whereas downward counterfactuals focus on an alternative that is worse than reality. Additive counterfactuals focus on the addition of antecedent elements that were not present in the past, whereas subtractive counterfactuals focus on the deletion of antecedent elements that were present in the past.;In all three experiments, these two variables are manipulated, forming, along with self-esteem, 2 x 2 x 2 factorial designs. In Experiment 1, subjects recalled negative life events, generated counterfactuals, and rated their current affect. Direction but not structure influenced affect ratings, such that downward counterfactuals resulted in more positive affect than upward counterfactuals. In Experiment 2, subjects recalled a disappointing examination performance, generated counterfactuals, rated their current affect, and rated their intentions to perform success-facilitating behaviours. Again, direction but not structure influenced affect rating in their same manner as in Experiment 1. Direction but not structure also influenced intention ratings, such that upward counterfactual generation resulted in stronger intentions to perform success-facilitating behaviours. In Experimental 3, subjects engaged in a computer-administered anagram task. Although the affective effects were not significant, both direction and structure influenced performance: upward as well as additive counterfactual generation resulted in greater improvement on the anagram task.;These findings provide initial support for a functional theory of counterfactual thinking: people may strategically use downward counterfactuals to make themselves feel better (an affective function), and they may strategically use upward and additive counterfactuals to improve performance in the future (a preparative function). The present studies suggest that the mechanism underlying the preparative function represents a causal link from counterfactuals to intentions to overt behaviours. Implications for current theory and future research are considered.
We propose a lexical organisation for multilingual lexical databases (MLDB). This organisation is based on acceptions (word-senses). We detail this lexical organisation and show a mock-up built to experiment with it. We also present our current work in defining and prototyping a specialised system for the management of acception-based MLDB.
The validity effect is the increase in perceived validity of repeated statements. In the first experiment, subjects rated repeated and nonrepeated statements for validity, familiarity, and source recognition. Validity and familiarity were enhanced by repetition, but source dissociation was not. A path analysis suggested that familiarity mediates perceived validity. In Experiment 2, statements presented in a natural setting were later rated for perceived validity, familiarity, and source recognition. Repetition had parallel effects on validity and familiarity ratings, but source dissociation was unaffected. Controlling for familiarity statistically eliminated the validity effect. In Experiment 3, the effect of prior knowledge on validity judgments was studied. Subjects rated the validity of statements that were related or unrelated to their field of expertise. Those most knowledgeable about the topic were most likely to exhibit the validity effect. Overall, the results suggest that familiarity is the basis of judged validity.
We describe a generative probabilistic model of natural language, which we call HBG, that takes advantage of detailed linguistic information to resolve ambiguity. HBG incorporates lexical, syntactic, semantic, and structural information from the parse tree into the disambiguation process in a novel way. We use a corpus of bracketed sentences, called a Treebank, in combination with decision tree building to tease out the relevant aspects of a parse tree that will determine the correct parse of a sentence. This stands in contrast to the usual approach of further grammar tailoring via the usual linguistic introspection in the hope of generating the correct parse. In head-to-head tests against one of the best existing robust probabilistic parsing models, which we call P-CFG, the HBG model significantly outperforms P-CFG, increasing the parsing accuracy rate from 60% to 75%, a 37% reduction in error.
Assessing the impact of disease and treatment of the emotional state and temperament of preschool children has been limited by the lack of sensitive and objective measurement techniques. To construct such a measure for longitudinal use, a sample of 179 children with febrile seizures and 85 normal children were used to develop the Minnesota Preschool Affect Rating Scales (MN-PARS). Video-taped play sessions were a source of behaviorally anchored ratings on 12 scales. Factor analysis yielded three factors of Negative Affect, Positive Affect, and Self-regulation with additional individual scales of Dependency and Activity Level. These scales and factors yield reliable ratings, as measured by interrater agreement and split-half techniques, as well as initial evidence of concurrent validity. Although they were developed for measuring the behavioral effects of phenobarbital on children with febrile seizures, these scales provide an objective means of measuring emotional expression and self-regulation useful for other studies.
We propose a lexical organisation for multilingual lexical databases (MLDB). This organisation is based on acceptions (word-senses). We detail this lexical organisation and show a mock-up built to experiment with it. We also present our current work in defining and prototyping a specialised system for the management of acception-based MLDB. Keywords: multilingual lexical database, acception, linguistic structure. Introduction Needs for large scale lexical resources for Natural Language Processing (NLP) in general and for Machine Translation (MT) in particular, increase every day. These resources are considered to represent the most expensive part of almost any NLP system. Hence, an increasing interest in the development of reusable dictionaries can be observed. To develop a Multilingual Lexical Database (MLDB), we think of two main approaches. First, the transfer approach where the links between the languages are realized via unidirectional bilingual dictionaries. This approach is...
An automatic treebank conversion method is proposed in this paper to convert a treebank into another treebank. A new treebank associated with a different grammar can be generated automatically from the old one such that the information in the original treebank can be transformed to the new one and be shared among different research communities. The simple algorithm achieves conversion accuracy of 96.4% when tested on 8,867 sentences between two major grammar revisions of a large MT system.
Two experiments were undertaken to examine whether facial responses to odors correlate with the hedonic odor evaluation. Experiment 1 examined whether subjects (n = 20) spontaneously generated facial movements associated with odor evaluation when they are tested in private. To measure facial responses, EMG was recorded over six muscle regions (M. corrugator supercilii, M. procerus, M. nasalis, M. levator, M. orbicularis oculi and M. zygomaticus major) using surface electrodes. In experiment 2 the experimental group (n = 10) smelled the odors while they were visually inspected by the experimenter sitting in front of the test subjects. The control group (n = 10) performed the same experimental condition as those subjects participating in experiment 1. Facial EMG over four mimetic muscle regions (M. nasalis, M. levator, M. zygomaticus major, M. orbicularis oculi) was measured while subjects smelled different odors. The main findings of this study may be summarized as follows: (i) there was no correlation between valence rating and facial EMG responses; (ii) pleasant odors did not evoke smiles when subjects smelled the odors in private; (iii) in solitude, highly concentrated malodors evoked facial EMG reactions of those mimetic muscles which are mainly involved in generating a facial display of disgust; (iv) those subjects confronted with an audience showed stronger facial reactions over the periocular and cheek region (indicative of a smile) during the smelling of pleasant odors than those who smelled these odors in private; (v) those subjects confronted with an audience showed stronger facial reactions over the M. nasalis region (indicative of a display of disgust) during the smelling of malodors than those who smelled the malodors in private. These results were taken as evidence for a more social communicative function of facial displays and strongly mitigates the reflexive-hedonic interpretation of facial displays to odors as supposed by Steiner.
The input and organization of terms play a central role in a dictionary-based translation system. The quality of the output text depends, for the most part, on the development of the dictionary. However, a lexical database cannot solve all problems of ambiguity. Moreover, it cannot take into account the effective usage of terms in specialized texts. This paper describes how terminology management is carried out in a machine-translation environment. We list a series of problems posteditors have to cope with and offer partial solutions to these problems. The material is based on work done in a translation firm where all translators are required to postedit machine translations.