Papers reviewed and determined not to be word norm studies. Use the flag icon to report errors or suggest re-inclusion.
16504 papers
The authors documented a linguistic norm account of direction of comparison asymmetry effects in relational judgments (e.g., seeing hyenas as more similar to dogs than dogs are similar to hyenas). The asymmetry effect is magnified by discrepancies in prominence between subject and reference, and has previously been explained using A. Tversky's (1977) feature-matching model. Given a linguistic norm to place more prominent objects in the referent position, violation of this norm might reduce sentence clarity, which then weakens the magnitude of subsequent relational judgments. 53 undergraduates completed a study that tested the degree to which both clarity and feature-matching predicted the magnitude of relational judgments by construction in 3 regression models, on for each of the 3 types of relation statements: similarity, difference, and spatial relation. Clarity perceptions predict the magnitude of relational judgments independently of the cognitive manipulation of the features of the compared objects. The pattern of findings suggests that a linguistic norm interpretation may account for variance in relational judgments independently of Tversky's feature-matching model. (PsycINFO Database Record (c) 2016 APA, all rights reserved)
Previous research on the effects of age of acquisition on lexical processing has relied on adult estimates of the age at which children learn words. The authors report 2 experiments in which effects of age of acquisition on lexical retrieval are demonstrated using real age-of-acquisition norms. In Experiment 1, real age of acquisition emerged as a powerful predictor of adult object-naming speed. There were also significant effects of visual complexity, word frequency, and name agreement. Similar results were obtained in reanalyses of data from 2 other studies of object naming. In Experiment 2, real age of acquisition affected immediate but not delayed object-naming speed. The authors conclude that age-of-acquisition effects are real and suggest that age of acquisition influences the speed with which spoken word forms can be retrieved from the phonological lexicon.
We tested the validity of the paddle method for measuring both the kinesthetic and visual-kinesthetic perception of inclination. In three conditions, subjects performed three different tasks: (1) rotating a manual paddle to a set of verbally given inclinations (blindfolded subjects), (2) rotating a manual paddle to the same set of verbally given inclinations after specific kinesthetic training (blindfolded subjects), and (3) rotating the paddle to a set of fixed visual inclinations after the kinesthetic training. The results showed a high degree of accuracy and precision in the second and third task but not in the first one. When subjects were asked to rotate a manual paddle to a set of verbally given inclinations, they used three main anchors (0°, 45°, 90°). Furthermore, the paddle method is biased by a kinesthetic deficiency, namely a rotational problem of the wrist that can be corrected by means of specific training.
We investigated the feasibility of a computer-graphics-based method of assessing stereomotion thresholds (Silicon Graphics Stereoview stereoscopic system). Stereomotion thresholds for a rectangle oscillating in depth were determined with the use of a dual randomly interleaved staircase design. In a group of 31 naive observers, the average thresholds of 5.97′ of arc forcrossed stereomotion and 6.00′ of arc foruncrossed stereomotion were comparable to those assessed in earlier work done with optics-based techniques. By assessing the thresholds for a rectangle that was defined either by lateral motion or by changing size, in a group of experienced observers, we were able to show that any potential residual translational motion present in the display would not have influenced the stereomotion thresholds. Our findings suggest that this computer-graphics-based technique may be a reasonable alternative to optics-based methods of assessing stereomotion thresholds.
Psychologists have used artificial neural networks for a few decades to simulate perception, language acquisition, and other cognitive processes. This paper discusses the use of artificial neural networks in research on semantics—in particular, in the investigation of abstract noun meanings. It is widely acknowledged that a word’s meaning varies with its contexts of use, but it is a complex task to identify which context elements are relevant to a word’s meaning. The present study illustrates how connectionist networks can be used to examine this problem. A simple feedforward network learned to distinguish among six abstract nouns, on the basis of characteristics of their contexts, in a corpus of randomly selected naturalistic sentences.
In two studies, the alternate-form reliability of the Snodgrass picture fragment completion test of implicit memory (Snodgrass, Smith, Feenan, & Corwin, 1987) was examined. In this test, identification thresholds are established for fragmented pictures. The same fragmented pictures are then shown again, intermixed with new fragmented pictures. Implicit memory is indicated by a decrease in identification threshold from the first to second presentation. Alternate-form reliability was low to moderate, depending on the measure used, regardless of the length of the test. A third study showed that explicit memory instructions did not increase the reliability. Recommendations for use of the test in correlational and experimental research are presented.
412 LANGUAGE, VOLUME 74, NUMBER 2 (1998) machine translation. Its substantial bibliography and its author and subject indexes are useful research tools in themselves. However, it cannot be considered a pedagogically-oriented introduction to the theory and practice of translation for the earlier stages of translator training. The book results from a lecture series given by the author in Finland in 1993 and is pitched at quite a demanding level although the claim that the 'individual chapters are relatively self-contained [and] can be read largely independently' (xiii) is justified. For this reason, it will be a very useful source of supplementary readings in advanced translator training. Also, teachers of translation, translation critics, and translators looking for some time out to reflect on the nature of their craft will benefit from it. They may end up agreeing, however, that at times W is not completely innocent of the 'pretentious, glutinous, heavily metaphorical or extremely abstract prose' (3) that he chides others who write about translation for using. Also, readers who know German will benefit more than those who don't from the fairly large sections of German text that are sometimes included to exemplify a point. While we should be surprised at a book on translation that doesn't include at least some other-language material, the growing internationalization of TS means that authors can no longer assume familiarity with particular languages by their readers. In such cases, interlinear glosses and/or literal back translations should be provided. While rejecting the impossibility of translation, W' s consideration of the text-related and translatorrelated problems inherent in translation in no way diminishes his appreciation of the complexity of the task. However, his balanced perspective always encourages a positive and hopeful outlook that interlingual communication is indeed possible, e.g. translation 'contains both culture-specific and cultureuniversal components' (90). Compensatory linguistic behaviors and adaptive skills and strategies can be acquired to enable the interlingual/intercultural gap to be bridged where necessary. W has little use for (sentence-based) generative theory which allows for a linguistic creativity that is both inapplicable and uninteresting to TS. On the other hand, modern cognitive linguistics is seen as having a most useful input into (text-based) translation theory and practice which should seek to operate in 'an interdisciplinary, cognitively embedded framework' (xiii). Such an approach will allow for the creativity displayed in linguistic performance to be given the central significance it deserves. W reminds us that the modern phase of 'TS... is still a fairly young and methodologically unstable field of research' (2), but this volume is indeed a most useful advance. [Robert Early, University of the South Pacific, Vanuatu.] Dictionary of Caribbean English usage. Ed. by Richard Allsopp. (French and Spanish supplement edited by J. E. Allsopp.) Oxford: Oxford University Press, 1996. Pp. lxxviii, 697. This dictionary (hereafter referred to as DCE) is a groundbreaking publication derived from fieldwork and approximately one thousand bibliographic sources. It presents extensive data on English spoken in the Anglophone West Indies (including the Bahamas, Belize, and Guyana) and is certain to become a valuable resource in the fields of both Creole and English studies. Allsopp should be congratulated for finishing what must have seemed a daunting project when begun more than 25 years ago. The aims of DCE are different from an earlier landmark in lexicography, the Dictionary ofJamaican English (ed. by F. G. Cassidy and R. B. Le Page, Cambridge University Press, 1967). That dictionary established the goal ofhistorically describing the lexicon of Jamaican English, including both creóle varieties as well as more standard forms. DCE, however, is less devoted to historical principles than to language planning, seeking to establish a norm for Caribbean English while identifying some regional variation. What is identified as 'Caribbean English' is in fact a narrowly-defined representation of lexical entries and idioms associated primarily with 'acrolectal ' and 'mesolectal' varieties. DCE purposely excludes entries associated with 'deeper' creóle forms, and consequently, it assumes an uncomfortably (and admittedly) prescriptive tone (see xxvi). A should keep in mind that the bundle of features which creolists often consider as constituting the so-called basilect, mesolect, and ACROLECT are ambiguous at best and artificial at...
Using a corpus to investigate empirically grammatical phenomena prior to writing grammatical rules or constraints for a disambiguating tagger is important. The paper shows how even case distinctions on pronouns are used more diversely than is usually assumed. Both in English and Norwegian nominative pronouns are used in more positions than the expected Subject one. Although the other uses are statistically less frequent, they may be important to the users of the resulting tagged corpus – who are often theoretical linguists. A tagger should therefore tag correctly also the more infrequent constructions. The paper shows how this can be done in a Constraint Grammar type tagger.
In many Oceanic languages in northwest Melanesia the default attribute construction ('a big house') is one whose morphosyntax looks like that of a possession construction: the attribute occupies the (possessed) head slot, the noun the (possessor) modifier slot ('a big one of a house'), that is, the opposite of the cross-linguistic norm and a rare phenomenon worldwide. I briefly describe these constructions, which are morphosyntactically varied, then examine their history, proposing that a major factor in their genesis was the presence in Proto-Oceanic of a small class of adjectival nouns whose reflexes in languages scattered across Oceania either may still behave as noun phrase heads or retain features reflecting this earlier status. The adjectival noun class had a small membership but high token frequency, and provided the template for a pattern extension that in a number of northwest Melanesian languages drew in the much larger adjectival verb class. I also address the question of why this change occurred in northwest Melanesia but not elsewhere in Oceania
This experimental study investigated three components of the lexical-acquisition process: fast mapping, word learning, and word extension. Thirty preschool-age children with specific language impairment (SLI) and 30 age- and gender-matched normal language (NL) controls participated. Two types of low-frequency words were used to name objects: phonologically simple and phonologically complex. Two types of objects were used: semantically familiar and semantically unfamiliar. Comprehension and production were probed across components, with trials to criterion calculated for word learning. During the word-extension task, recognition was also assessed. In addition, norm-referenced receptive and expressive vocabulary tests were administered. On tasks where group differences were found, many children with SLI performed as well as controls. The fast-mapping task revealed no group differences (p On this and the word-learning task, comprehension exceeded production for both groups. On the word-learning task, the SLI group comprehended and produced fewer words than did controls. In contrast to the NL group, the SLI group produced more words for semantically unfamiliar than familiar objects. Overall, the SLI learned more phonologically simple than complex words, but the NL group showed no difference. Although the SLI group required more trials than controls to comprehend words, no group differences were found in trials to production. On the word-extension task, the SLI group scored lower than controls only for production. Vocabulary-test scores did not accurately identify children with SLI or predict number of words learned; but these scores did predict a small amount of variance for fast-mapping and word-extension performance. Fast mapping performance accounted for 25% of word-learning variance; fast-mapping performance and word-learning performance combined accounted for 54% of word-extension variance. Different predictor variables were found for each language group. Findings suggest that some, but not all, children with SLI demonstrate poor word-learning and word-extension performance. Overall, the SLI group had greater difficulty than controls on production measures and required more trials to achieve learning criterion for comprehension of words than the NL group. The semantic familiarity of the target objects affected productive word learning by the NL group but did not appear to have a similar effect on the SLI group.
u radu se obrađuju neke morfoloske osobitosti imenica, zamjenica i pridjeva u prozi Janka Polica Kamova
In the present article, we investigated the reading ability of CP, a pure alexic patient, using an experimental paradigm that is known to elicit the viewing position effect in norm al readers. The viewing position effect consists of a systematic variation of word recognition performance as a function of fixation location w ithin a word: Word recognition is best when the eyes fixate slightly left from the word centre and decreases when the eyes deviate from this optimal viewing position. A mathematical model (Nazir, O'Regan, & Jacobs, 1991), which provides a good description and quantification of the prototypical shape of the viewing position effect, served to interpret CP's reading performance. The results show ed that, like normal readers, CP was able to process all letters of a w ord in one fixation. However, in contrast to normal readers, reading performance was optimal when CP w as fixating the right half of the word. This somewhat abnormal pattern of performance was due to (1) poor perceptual processing in the right visual field, and (2) poor processing of letters situated towards the end of the word, independent of visual field presentation. A similar pattern of perform ance w as obtained with normal readers under experimental conditions in which lexical know ledge was of restricted use. We suggest that CP's reading impairment stems from a dysfunction in the coupling between incoming visual information and stored lexical information. This dysfunction is thought to uncover a prelexical level of word processing, where letter information is weighted differently as a function of letter position in a word-centred space.
The introduction of vocabulary checklists for infant acquisition has allowed us to gather detailed information about early lexical growth from a broader number of bilingual children than before and to begin exploring the relation of growth in one language to growth in the other in a range of bilingual learning circumstances. Still, no standardized instruments to date give an adequate picture of normal bilingual development. Norms and milestones based on monolingual experience underestimate bilingual abilities in that they tap only a portion of the bilinguals' knowledge and credit them with less complex conceptualizations than what they actually possess. Double-language measures, like those proposed by Pearson, Fernández, and Oller (1993) and Muñoz-Sandoval, Cummins, Alvarado, and Ruef (1998), are an improvement over single-language measures as they encompass a greater portion of the bilinguals' knowledge, but they do not address the greater complexity of bilingual mental representations. What is needed are norms derived from observations of typically-developing bilingual children, followed up with measures of concurrent and predictive validity. However, bilinguals as a group are so diverse, it will be difficult to decide which subgroup(s) would be the appropriate reference to use for a standardization. In this review of recent studies, the difficulties involved in assessing bilingual vocabularies and recommendations for clinical practice are discussed.
Problems in retrieval of names form large data bases and in nominal record linkage are discussed with respect to computational solutions. The quest for robust methods that can handle the typical variability of historical nominal information is discussed, with some emphasis on probabilistic methods. It is argued that comparison and assessment of different systems used on the same data could enhance our understanding of methodological issues.
Dans son ouvrage intitule Etudes sur les Tchinghianes ou Bohemiens de l'Empire Ottoman (1870), Paspati a examine les particularites des dialectes tsiganes au niveau de la structure semantique de plusieurs unites lexicales et au niveau de la construction de groupes de mots et de phrases. Il a ainsi contribue de maniere appreciable a la description des interferences dans la langue tsigane des Balkans, qui - grâce au bilinguisme - se manifestent dans la prise en charge de plusieurs types d'emprunts formes d'apres des modeles grecs et dans l'appropriation de regles et de normes grammaticales et syntaxiques grecques. Dans cet article, l'A. examine les similitudes et les concordances grammatico-typologiques et lexico-semantiques entre les dialectes tsiganes examines par Paspati et plusieurs langues des Balkans, et montre que, malgre des doutes repetes, ce domaine de recherche est une mine d'informations sur les contacts de langues dans l'aire linguistique des Balkans
Finding simple, non-recursive, base noun phrases is an important subtask for many natural language processing applications. While previous empirical methods for base NP identification have been rather complex, this paper instead proposes a very simple algorithm that is tailored to the relative simplicity of the task. In particular, we present a corpus-based approach for finding base NPs by matching part-of-speech tag sequences. The training phase of the algorithm is based on two successful techniques: first the base NP grammar is read from a "treebank" corpus; then the grammar is improved by selecting rules with high "benefit" scores. Using this simple algorithm with a naive heuristic for matching rules, we achieve surprising accuracy in an evaluation on the Penn Treebank Wall Street Journal.
We present a method for the extraction of stochastic lexicalized tree grammars (SLTG) of different complexities from existing treebanks, which allows us to analyze the relationship of a grammar automatically induced from a treebank wrt. its size, its complexity, and its predictive power on unseen data. Processing of different S-LTG is performed by a stochastic version of the two-step Early-based parsing strategy introduced in (Schabes and Joshi, 1991). 1 Introduction In this paper we present a method for the extraction of stochastic lexicalized tree grammars (S-LTG) of different complexities from existing treebanks, which allows us to analyze the relationship of a grammar automatically induced from a treebank wrt. its size, its complexity, and its predictive power on unseen data. The use of S-LTGs is motivated for two reasons. First, it is assumed that S-LTG better capture distributional and hierarchical information than stochastic CFG (cf. (Schabes, 1992; Schabes and Waters, 1996)),...
This research documented a linguistic norm account of direction of comparison asymmetry effects in relational judgments (e.g., seeing hyenas as more similar to dogs than dogs are similar to hyenas). The asymmetry effect is magnified by discrepancies in prominence between subject and referent, and has previously been explained using Tversky's (1977) feature-matching model. Given a linguistic norm to place more prominent objects in the referent position, violation of this norm might reduce sentence clarity, which then weakens the magnitude of subsequent relational judgments. This research showed that clarity perceptions predict the magnitude of relational judgments independently of the cognitive manipulation of the features of the compared objects. The pattern of findings suggests that a linguistic norm interpretation may account for variance in relational judgments independently of Tversky's (1977) feature-matching model.
Although the influence of emotional arousal on declarative memory has been documented behaviorally, the mechanisms underlying arousal-memory interactions and their representation in the human brain remain uncertain. One route through which arousal achieves its effects on memory performance is by regulating consolidation processes. Animal research has revealed that the amygdala strengthens hippocampal-dependent memory consolidation in a limited time window following participation in an arousing task. To examine whether this integrative function of amygdalo-hippocampal structures extends to the human brain, we tested unilateral-temporallobectomy patients on an adaptation of a classic paradigm in which levels of physiological arousal at encoding modulate retention over time. Subjects rated emotionally arousing (taboo) and neutral words on an arousal scale while their skin conductance responses (SCRs) were monitored. Recall for the words was assessed immediately and after a 1-hr delay. Both temporal-lobectomy patients and control subjects generated enhanced SCRs and arousal ratings for the arousing words at the time of encoding. However, only control subjects exhibited an increase in memory for the arousing words over time. This group difference in the effect of arousal on the rate of forgetting suggests that the role of medial temporal lobe structures in memory consolidation for arousing events is conserved across species.
Most bipolar models of affective processing in social psychology assume that positive and negative valent processes are represented along a single continuum that rangesfrom very positive to very negative. Recent research has raised the possibility, however, that the motivational systems for positive/approach and negative/defensive valent processing (positivity and negativity, respectively) are separable. In this article, the authors use unipolar positivity, negativity, and ambivalence ratings and bipolar valence, dominance, and arousal ratings of 472 slides from the International Affective Picture System to examine several aspects of the bivariate model of evaluative space. Analysis confirmed a positivity offset and negativity bias in the activation functions of the valent systems as wel as multiple modes of evaluative activation (e.g., reciprocal, uncoupled positivity, uncoupled negativity). Together, these data suggest that the bipolar structure of affective processes should be tested rather than assumed.
Participants judged the believability of simulated accounts varying in perceptual and emotional detail. In Experiment 1, both younger and older adults' tendency to believe that an account described an actually experienced event increased as either type of detail was added. In Experiment 2, while younger adults in a low suspicion condition judged the accounts with added details as more likely to be of actually-experienced events than the impoverished accounts, those given instructions designed to induce suspicion about the speakers' honesty found the more detailed accounts less believable. In Experiment 3, for both younger and older adults, both types of detail again increased believability ratings under low-suspicion conditions but did not affect ratings under high-suspicion conditions. In addition, there were systematic differences in the types of details the high- and low-suspicion participants reported using to make their judgments. Results are discussed in relation to the source monitoring framework (Johnson, Hashtroudi, & Lindsay, 1993) and reality monitoring and credibility judgments.
Using the cold pressor test, three experiments were conducted to investigate the effects of water temperature and labeling on three dependent measures in college women: behavioral pain tolerance (BPT), a sensory rating of the pain experience (SR) and a parallel affective rating of the experience (AR). Temperature of the cold pressor was varied as the physical factor; labels (discomfort, pain, vasoconstriction pain) were varied as the psychological factor. Experiment I varied only water temperature; colder temperatures led to significantly lower BPT scores and significantly higher SR and AR scores. Experiment 2 varied only labeling and demonstrated that BPT decreased and AR increased as labels became more painful-sounding; in contrast, SR was unaffected by labeling. In Experiment 3 both the psychological and physical factors were varied simultaneously. Results indicated significantly higher BPT scores as the water temperature increased and the pain label became more benign. In addition, both SR and AR were sensitive to changes in temperature, whereas only AR was affected by changes in labeling.
Children's judgements about pain at age 8–10 years were examined comparing two groups of children who had experienced different exposure to nociceptive procedures in the neonatal period: extremely low birthweight (ELBW) ≤ 1000 g ( N = 47) and full birthweight (FBW) ≤ 2500 g ( N = 37). The 24 pictures that comprise the Pediatric Pain Inventory, depicting events in four settings: medical, recreational, daily living, and psychosocial, were used as the pain stimuli. The subjects rated pain intensity using the Color Analog Scale and pain affect using the Facial Affective Scale. Child IQ and maternal education were statistically adjusted in group comparisons. Pain intensity and pain affect related to activities of daily living and recreation were significantly higher than psychosocial and medically related pain on both scales in both groups of children. Although the two groups of children did not differ overall in their perceptions of pain intensity or affect, the ELBW children rated medical pain intensity significantly higher than psychosocial pain, unlike the FBW group. Also, duration of neonatal intensive care unit stay for the ELBW children was related to increased pain affect ratings in recreational and daily living settings. Despite altered response to pain in the early years reported by parents, on the whole at 8–10 years of age ELBW children judged pain in pictures similarly to their term peers. However, differences were evident, which suggests that studies are needed of biobehavioural reactivity to pain beyond infancy, as well as research into beliefs, attitudes, and perceptions about pain during the course of childhood in formerly ELBW children.
Part-of-speech tagging methodology has succeeded, but on problems that may lack real-world application. Redirection of the field is indicated, toward potentially more useful, but harder and more sophisticated tagging tasks: (1) using much more detailed tagsets (semantically and syntactically); (2) testing performance on treebanks reflecting the huge gamut of domains, etc., characterizing real-world applications; (3) understanding the magnitude of the unknown-word and unknown-tag problems, then overcoming them. Tagging results are presented on two versions of a new, highly variegated treebank, featuring tagsets of 2720 and 443 tags, respectively, and utilizing a dictionaryless, decision-tree tagger.
This paper compares different methods of generating intonation for an American English Text-to-Speech synthesis system. We look at a primarily rule-based approach and two data-driven approaches. For data-driven modeling we used two separate data sets, each representing a somewhat different prosodic style. One database was recordings of a portion of 1989 Wall Street Journal text from the Penn Treebank Project. The second database was recordings of interactive prompts used in telephone network services. Both were read by the same female speaker. Approximately two and one-half hours of speech was phonetically and prosodically segmented and labeled (first automatically, and subsequently verified manually). The prosodic labeling used ToBI [7] tones and breaks. Three different intonation models were compared: (1) a predominantly rule-based model based on ToBI labels [3]; (2) a parametric model using the Tilt approach [8]; and (3) a Vector Quantized model based on an underlying parametric re...
A series of four experiments were conducted to examine viewer perceptions of three sets of five nonrepresentational paintings. Increased complexity was embedded in the hierarchical structure of each set by carefully selecting colors and ordering them in each successive painting according to certain rules of transformation which created hierarchies. Experiment 1 supported the hypothesis that subjects would discern the hierarchical complexity underlying the sets of paintings. In Experiment 2 viewers rated the paintings on collative (complexity, disorder) and affective (pleasing, interesting, tension, and power) scales, and a factor analysis revealed that affective ratings were tied to complexity (Factor 1) but not to disorder (Factor 2). In Experiment 3, a measure of exploratory activity (free looking time) was correlated with complexity (Factor 1) but not with disorder (Factor 2). Multidimensional scaling was used in Experiment 4 to examine perceptions of the paintings seen in pairs. Dimension 1 contrasted Soft with Hard-Edged paintings, while Dimension 2 reflected the relative separation of figure from ground in these paintings. Together these results show that untrained viewers can discern hierarchical complexity in paintings and that this quality stimulates affective responses and exploratory activity.
We argue that the current dominant paradigm in parser evaluation work, which combines use of the Penn Treebank reference corpus and of the Parseval scoring metrics, is not well-suited to the task of general comparative evaluation of diverse parsing systems. We propose an alternative approach which has two key components. Firstly, we propose parsed corpora for testing that are much flatter than those currently used, whose "gold standard" parses encode only those grammatical constituents upon which there is broad agreement across a range of grammatical theories. Secondly, we propose modified evaluation metrics that require parser outputs to be `faithful to', rather than mimic, the broadly agreed structure encoded in the flatter gold standard analyses. 1. Introduction Interest in the evaluation of language technology has grown immensely in the past few years. This interest varies depending on the perspective one has on the technology: users and suppliers want to know how accurate, usabl...
`Linguistic annotation' is a term covering any transcription, translation or annotation of textual data or recorded linguistic signals. While there are several ongoing efforts to provide formats and tools for such annotations and to publish annotated linguistic databases, the lack of widely accepted standards is becoming a critical problem. Proposed standards, to the extent they exist, have focussed on file formats. This paper focuses instead on the logical structure of linguistic annotations. We survey a wide variety of annotation formats and demonstrate a common conceptual core. This provides the foundation for an algebraic framework which encompasses the representation, archiving and query of linguistic annotations, while remaining consistent with many alternative file formats. 1. INTRODUCTION `Linguistic annotation' is a cover term for any orthographic, phonetic or prosodic transcription; any speech, part-of-speech, disfluency or gestural annotation; and any free or word-level tr...
Machine learning techniques can be used to make lexicons adaptive. The main problems in adaptation are the addition of lexical material to an existing lexical database, and the recomputation of sublanguage-dependent lexical information when porting the lexicon to a new domain or application. Inductive lexicons combine available lexical information and corpus data to alleviate these tasks. In this paper, we introduce the general methodology for the construction of inductive lexicons, and discuss empirical results on a case study using the approach: prediction of the gender of nouns in Dutch. 1. Introduction In computational lexicography, lexicons of language engineering applications should come with acceptable lexical coverage, and with the information necessary for the intended applications. They should also come equipped with methods for the automatic extension and adaptation of the lexicon with new or modified lexical entries. Computational lexicology should therefore try to solve t...
Word recognition and generation is a fundamental part of the processing of natural language and it requires computationally effective morphological processors, especially for languages with rich morphology such as Modern Greek. Various models have been proposed for developing computerized systems to accomplish the task of recognition of morphosyntactic features of words In the work presented here, the lazy tagging approach was examined, in which taggers are expected to work in the simplest possible way. The model of functional decomposition was extended and adapted for Modern Greek as a target language, following the lazy word-parsing approach, in order to cover a number of morphological phenomena that are encountered in Modern Greek, namely inflection, affixation, and longdistance dependencies. To achieve a more efficient word recognition, several automata of different levels of computing power, based on the original model, were introduced and evaluated according to the criteria of complexity, recognition speed, and accuracy of the results. The proposed system was used for processing a large-scale corpus, and the results are presented and discussed. To accomplish their task, taggers can rely upon large lexical databases, which are expected to be organized in such a way as to provide rapid access to the stored data and efficient memory management. Directed graphs can be used to describe and organize a lexical database of large magnitude in a compact manner. These data structures are named here matrix lexica, where the letters are described as nodes of directed graphs and the lemmata as paths (set of edges). It is expected that matrix lexica will support a tagger efficiently by providing a high speed of resolution, sound mathematical foundation, low memory requirements, and ability to handle distorted input in future developments.
Inscriptions, a type of source which has hitherto hardly been taken into account in historical linguistics, can provide data for a whole series of questions thrown up by research into the linguistic geography of Early New High German and the Late Middle Low German of the same period. In respect of the pressure towards standardization exerted by M. Luther's bible, various degrees can be differentiated. The greatest willingness to adopt this linguistic norm is found in the case of New High German monophthongization and diphthongization, as well as in the lowering of u to o before nasals and the use of the verbal prefix ver-. In these cases, the symbols on the appropriate maps show agreement with Luther 's usage. In some other cases, namely, the use of for MHG /ei/, the occurrence of vowel forms without unrounding, the use of the uncontracted form of the word nicht and the use of the subjunctive in the predicates of concessive clauses, we find that the form favoured by Luther is the one which appears most frequently in the inscriptions. Nevetheless, especially in sixteenth-century inscriptions, we find several examples of the replacement of standard forms by deviant variants valid at the same time. In contrast, the number of deviations from Luther's language visibly decreases in the seventeenth century. In some cases, the treatment of the norms of the Luther's bible in the inscriptions indicates that the linguistic usage of the reformer had no or only indirect significance for the course of development here. Thus, in the case of the feminine inflectional system, while it is true that the epigraphic sources largely follows Luther's practice, but it is also the case that from about 1600 on a tendency towards deviation from his usage emerges. Here, restructuring is a process which was clearly not influenced by Luther. Again, in the case of abstract suffix -nis, the course of development seems to have been determined by factors inherent in the language rath
Recent approaches to statistical parsing include those that estimate an approximation of a stochastic, lexicalized grammar directly from a treebank and others that rebuild trees with a number of tree-constructing operators, which are applied in order according to a stochastic model when parsing a sentence. In this paper we take an entirely different approach to statistical parsing, as we propose a method for parsing using a Hidden Markov Model. We describe the stochastic model and the tree construction procedure, and we report results on the Wall Street Journal Corpus.
This paper describes a method in determining syntactic structure for coordinate constructions. It is based on the information taken from semantic similarities, selectional restrictions, and some other linguistic cues. We discuss the role the information plays in resolving ambiguities that appear in coordinate constructions, describe the means of acquiring the necessary information automatically from two on-line corpora and a lexical database, and devise two algorithms for disambiguating coordinate constructions. An experiment that follows shows effectiveness of our method and its applicability to resolving ambiguities in some other syntactic structures.
This paper describes the combination compound unit (CU) recognizer with syntactic verifier using partial parsing mechanism. The recognizer finds all the CUs, combined concept including collocations, idioms, and compound nouns, in input sentence. CU information reduces the search space of syntactic analysis and a portion of Part-Of-Speech (POS) ambiguities. Syntactic verification is to obtain precise CU recognition results by means of pruning wrongly recognized units that are caused by improper variable hypotheses. The experimental results show the precision of CU recognition is increased to 99.69% with 31 CFG rules on cyclic trie structure for 1,268 WSJ articles in the Penn Treebank. They also show CU recognition increases the understandability of translation for Web documents.
We describe a system for extracting concepts from unstructured text. We do this by clustering document words and then assembling a structure which relates these words semantically. The clustering process identifies words which co--occur across a set of documents and creates groups of words which suggest a semantic context common across the document set. This context is formalized by identifying semantic relationships between the cluster words using a lexical database to build a Semantic Relationship Graph (SRG). This SRG is a directed graph which conveys a robust representation of the sub-- and super--class relationships between the correct word senses. We show how this process can be applied to a user--selected set of HTML documents; the SRGs can subsequently aid in searching the World Wide Web for documents which are similar. 1 Introduction Mining textual information presents challenges over data mining of relational or transaction databases because there are no predefined fields...
Children's judgements about pain at age 8-10 years were examined comparing two groups of children who had experienced different exposure to nociceptive procedures in the neonatal period: extremely low birthweight (ELBW) <or = 1000 g (N = 47) and full birthweight (FBW) > or = 2500 g (N = 37). The 24 pictures that comprise the Pediatric Pain Inventory, depicting events in four settings: medical, recreational, daily living, and psychosocial, were used as the pain stimuli. The subjects rated pain intensity using the Color Analog Scale and pain affect using the Facial Affective Scale. Child IQ and maternal education were statistically adjusted in group comparisons. Pain intensity and pain affect related to activities of daily living and recreation were significantly higher than psychosocial and medically related pain on both scales in both groups of children. Although the two groups of children did not differ overall in their perceptions of pain intensity or affect, the ELBW children rated medical pain intensity significantly higher than psychosocial pain, unlike the FBW group. Also, duration of neonatal intensive care unit stay for the ELBW children was related to increased pain affect ratings in recreational and daily living settings. Despite altered response to pain in the early years reported by parents, on the whole at 8-10 years of age ELBW children judged pain in pictures similarly to their term peers. However, differences were evident, which suggests that studies are needed of biobehavioural reactivity to pain beyond infancy, as well as research into beliefs, attitudes, and perceptions about pain during the course of childhood in formerly ELBW children.
198 LANGUAGE, VOLUME 74, NUMBER 1 (1998) cussion of specific examples reveals some characteristic aspects of agrammatic output. Ch. 4, 'The grammar of connected agrammatic speech', completes the summary of grammatical characteristics of nonfluent aphasia. It argues against the label 'telegraphic speech' and against the theory of 'economy of articulatory effort' in explaining agrammatic behavior. Ch. 5, 'Speech, writing, and oral reading', discusses differences and similarities of impairment among these different output modes. Examples of spontaneous writing illustrate the possibility of dramatic dissociation between written and spoken language. Examples from oral reading show that a patient may tend to make the same types of errors in reading as in speech. Ch. 6, 'Bilingual and polyglot aphasia', by Loraeme K. Obler, José Centeno, and Nancy Eng, discusses the similarities and differences between languages in aphasies who are multilingual. They point out that deficits and recovery are usually parallel but that variance can occur (resulting, for example, from structural differences between languages). The authors then discuss special behaviors and brain organization of the bilingual aphasie. They end the chapter with advice to speech-language pathologists on diagnosis of patients who speak a language or dialect unfamiliar to the clinician. Ch. 7, 'Inventing therapy for aphasia', by Audrey L. Holland and Claire Penn, is a fascinating discussion of how to devise therapy when the patient and therapist do not share the patient's primary language. They address such issues as how to choose which language to work with and the importance of cross-cultural influences on therapy design. In sum, this book is primarily designed to inform clinicians in their efforts to provide therapy. However, the liberal exemplification of agrammatic output should be of interest to even nonclinical researchers. [Sherri K. Shaw, University of Texas, Austin. ] Xenismen: Die Nachahmung fremder Sprachen. By Wolfgang Moser (Europ äische Hochschulschriften, Reihe 21, 159.) Frankfurt am Main Peter Lang, 1996. Pp. 284. Xenisms are phenomena that characterize somebody or something as foreign. They can be nonlinguistic (a kimono, national anthems, schnitzel) or linguistic (Chinese characters, a foreign accent, loan words). Most importantly, xenisms must be considered as typical of a foreign people or language, even if this doesn't coincide with reality: most Germans don't wear leather breeches, and Chinese people do have an ItI phoneme distinct from IV. In his PhD thesis (University of Graz, Austria), Moser gives a detailed linguistic and semiotic analysis ofxenisms that imitate foreign languages, limiting the scope of his study to intentionally used interlingual xenisms in written texts. He then discusses the role of foreignness and points out that on a scale 'total ignorance—intimate knowledge', foreignness is more or less near to ignorance but not identical to it. You have to know at least something (even something wrong) about other people to see them as foreign and different from yourself. You must know that Fritz is a German name to evoke 'germanness', but you need not know more (e.g. that hardly anybody in Germany would call their son Fritz nowadays, but rather Michael or Kevin). This explains why linguistic xenisms do not appear at random but are associated with the languages and cultures in contact with one's own. The introductory chapter (1 3-29) is completed by a series ofprivative dichotomies: Xenisms can be cotextual or implanted, reduced or fully decodable, interlingual or intralingual, spontaneous or conventional, foreign or pseudo-foreign, foreign elements (mostly words) or foreign usages (e.g. word order rules). Finally, the relation of xenisms to a certain language can be more or less vague or precise. Apart from the introductory chapter, the book is divided into two main parts: The first part (31-129) analyzes the linguistic structure of xenisms; the second part (131-254) deals with xenisms as semiotic phenomena. In the structural analysis, M explains xenisms as deviations from linguistic norms (following Eugenio Coseriu's definition of norm). Their forms range from the imitation of hieroglyphs to 'typical' name endings and are broadly illustrated with examples from Asterix and its translations. The second part of this chapter analyzes the various languages that are used to characterize different people in Jaroslav Hasek's The good soldier Svejk and how the xenistic effects...
Two barriers to the use of the Thematic Apperception Test (TAT) in motivation research were addressed: its low internal consistency and its time-consuming coding system. Sixty males and 60 females wrote five stories to TAT pictures either on the computer or by hand. Half of each group were timed and half untimed. The writing of stories was guided by four sets of questions, and stories were coded for need for power (n Pow) by the corresponding four paragraphs. Cronbach’s alpha for the five stories was .46; for the 20 paragraphs, Cronbach’s alpha was .65. We conclude that, to the extent that measuring internal consistency is appropriate for a thought-sampling instrument like the TAT, internal consistency should be calculated by paragraphs. Significantly more words were produced in the untimed condition, but n Pow did not differ by gender, hand-written versus computer-written, or timed versus untimed conditions. The five pictures elicited significantly different amounts of n Pow. It is recommended that researchers who give the TAT on the computer use the untimed condition. Suggestions are made for increasing the scoring validity and for using the computer to decrease the time required for human coders.
A well-known problem in the domain of quantitative linguistics and stylistics concerns the evaluation of the lexical richness of texts. Since the most obvious measure of lexical richness, the vocabulary size (the number of different word types), depends heavily on the text length (measured in word tokens), a variety of alternative measures has been proposed which are claimed to be independent of the text length. This paper has a threefold aim. Firstly, we have investigated to what extent these alternative measures are truly textual constants. We have observed that in practice all measures vary substantially and systematically with the text length. We also show that in theory, only three of these measures are truly constant or nearly constant. Secondly, we have studied the extent to which these measures tap into different aspects of lexical structure. We have found that there are two main families of constants, one measuring lexical richness and one measuring lexical repetition. Thirdly, we have considered to what extent these measures can be used to investigate questions of textual similarity between and within authors. We propose to carry out such comparisons by means of the empirical trajectories of texts in the plane spanned by the dimensions of lexical richness and lexical repetition, and we provide a statistical technique for constructing confidence intervals around the empirical trajectories of texts. Our results suggest that the trajectories tap into a considerable amount of authorial structure without, however, guaranteeing that spatial separation implies a difference in authorship.
The lateralized readiness potential (LRP) is an electrophysiological indicator of the central activation of motor responses. Procedures for deriving the LRP on the basis of event-related brain potential (ERP) waveforms obtained over the left and right motor cortices are described, and some findings are summarized that show that the LRP is likely to reflect activation processes within the motor cortex. Two experiments investigating spatial S-R compatibility effects are reported that demonstrate that, because of systematic overlaps of motor and nonmotor asymmetries, LRP waveforms derived by the double subtraction method cannot always be interpreted unequivocally in terms of response activation. Such confounds can be detected when LRP waveforms are compared with difference waveforms obtained by the double subtraction method from ERPs elicited at other lateral scalp sites.
Forty-nine young adults (M age = 22 years) and 30 elderly adults (M age = 69 years) rated the 60 pictorial stimuli from the Boston Naming Test (BNT) on familiarity, providing the first such normative data for these stimuli along this dimension. Participants also made speeded lexical decisions about the word item representations of each BNT picture. B NT word frequency values were also examined in relation to BNT familiarity and speeded lexical decision performance. For both young and elderly adults, lexical decision reaction times to the word representations of BNT stimuli were negatively related to word frequency and familiarity of the BNT pictures. These patterns suggest that increases in word frequency and picture familiarity facilitate (i. e., speed up) the processing of BNT word representations. Furthermore, speed of processing appears to be a relevant dimension of BNT performance, at least when young and elderly adults free from clinical aphasia are involved.
In “Response to Elliott and Valenza, 'And Then There Were None'”, (1996) Donald Foster has taken strenuous issue with our Shakespeare Clinic's final report, which concluded that none of the testable Shakespeare claimants, and none of the Shakespeare Apocrypha poems and plays – including Funeral Elegy by W.S. – match Shakespeare. Though he seems to accept most of our exclusions – notably excepting those of the Elegy and A Lover's Complaint – he believes that our methodology is nonetheless fatally flawed by “worthless figures ... wrong more often than right”, “rigorous cherry–picking”, “playing with a stacked deck”, and “conveniently exil[ing] ... inconvenient data.” He describes our tests as “foul vapor” and “methodological madness.”
This article describes how an independent commercial academic publisher initiated its electronic publishing programme. It outlines the range of electronic activities under development and some of the issues addressed during the creation of electronic resources. Case studies of two early projects are included: a multimedia teaching too, A Right to Die? The Dax Cowart Case; and an SGML textbase, the Arden Shakespeare CD-ROM. In addition, the Routledge Encyclopedia of Philosophy is discussed as an example of the second generation of electronic projects at Routledge, highlighting lessons learned from previous projects and some of the issues relating to the production of a simultaneous print and electronic resource.
BOOK NOTICES 207 sure's concept ofmotivation should not be associated with words but rather with cotext and context. Everything is relative in language, and the systematic character of language does not lie in separate phonemic, lexical, grammatical, and textual systems. In fact the systematic and universal features of language are made up by the human capabilities of thinking and experiencing. Speakers are able to use a limited number of signs to express highly complicated ideas, and they can decode expressions in very complex situations. To do this requires applying the basic principle of language and language description—the idea of economy. This elementary principle of human behavior can be seen when a speaker tries to avoid unnecessary redundancy and when simpler ways of pronunciation are preferred to more difficult ones. In addition language economy can be seen on the deeper level of reinterpreting linguistic units new to a certain speaker. This gives sense to an utterance since under the principle ofeconomy, a speaker must assume that no text is uttered without meaning. As a result of new expressions and new interpretations produced by the principle ofeconomic use, language might change over time. Considering language change again, D emphasizes how the individual reflects about language. For the individual the main goal of language and speaking is to impart and to decode sense or meaning. D rejects the idea of the 'invisible hand phenomenon' as well as the concept of teleology in language change. Again he links the systematic character of language to the speaker's purposeful acts rather than to the whole speech community or to language as an abstract system. In short, D questions the traditional structuralist conception of system in language. He emphasizes the systematic cognitive behavior of each individual that uses the relative means of language. In order to support his opinion D illustrates his ideas with many detailed examples mainly taken from German, English, and French. [Dieter Aichele, Fachhochschule Neubrandenburg.] The Oxford English-Hebrew Dictionary. Ed. by N. S. Doniach and A. Kahane. Oxford: Oxford University Press, 1996. Pp. xxiii, 1091. The late N. S. Doniach (d. 16 April 1994), the chief editor ofthis dictionary, is well known to Semitologists for editing the excellent Oxford EnglishArabic dictionary ofcurrent usage (1972). The introduction by Professor A. Kahane explains D's goal of using various styles of modern Hebrew in this volume, including colloquial language and slang. From abacus to Zulu, this dictionary, happily, has it all! It even has the f-word with many of its most common idiomatic usages, such as '__ up' and '__ off!' (351). The tome's particularly noteworthy features include up-to-date terminology ofall sorts, such as that dealing with computers. However, some inconsistencies can be found. 'Software' is written as toxna with a vav (883), but it is spelled without a vav (using a kamats katan) under 'hardware' (400). And curiously, the name of the vowel kamats is written kamatz, yet the vowel chataf kamats is spelled differently on the very next line (ix). AU Hebrew words are given in their fully vocalized or pointed forms, including dagesh, which is said to have the phonetic value 'stress mark', a puzzling statement (ix). Phonological matters on the whole, however, have been handled well. It was a wise decision for the editors to give preference to a word's modern pronunciation if it differs from its traditional pointing (xxiii). Along these lines, I wanted to check the vocalization and pronunciation of the irregular Classical Hebrew plural for bayit 'house'—battiim or (bottiim); to my surprise, however, it was not given (426). British English has been chosen as the norm of this dictionary, a reasonable and expected decision by Oxford University Press. An American has little difficulty getting used to British spellings such as 'programme'; however, 'program' is also listed, but with the stipulation, 'US and Comput.' (724). However, it may be difficult at first for an American to appreciate 'farther' and 'father' transcribed exactly the same. There are a few discrepancies to report. American English 'buggy' is given as eglat-tinok (114), but 'pram' lists only eglat-tinokot, its plural (708), whereas the more formal 'perambulator' (called 'formal ' by the editors) lists...
Chadwyck-Healey has a long tradition of electronic publishing. Beginning with production of CD-based literary corpora, it has recently moved many of its products to a web-accessible online environment. The article reflects on experiences with both CD and web-based publications.
The article reports on one of the more sophisticated critical editions ever to be published in electronic format. The Wife of Bath is richly encoded, provides access to literally thousands of manuscript images, and enables users to assess the relationships between the numerous extant manuscript editions. The authors assess the methods used in the edition's development and the lessons learned through its production.
“On-line communities” (and especially MUDs—“multiuser domains”) are a popular, growing Internet phenomenon. This paper provides an overview of a project designed to provide a careful characterization of what “life” is like in LambdaMOO—a classic social MUD—for most, or at least many, members. A “convergent-methodologies” approach embracing qualitative and quantitative, subjective and objective methods was used to generate a large and rich database on this on-line community in terms of four general categories: (1) users and use, (2) sociality, (3) identity, and (4) spatiality. The evidence thus far appears to debunk some of the more provocative claims of widespread MUD addiction and rampant identity fragmentation on line. While supporting the primary importance of sociality in the MUD, the results also demonstrate the strong prevalence of personal, one-on-one social interactions over larger social gatherings. Finally, some close correspondences between patterns of spatial behavior and spatial cognition “in real life” and in LambdaMOO were found.