Papers reviewed and determined not to be word norm studies. Use the flag icon to report errors or suggest re-inclusion.
18265 papers
Hebrew Studies 40 (1999) 269 Reviews grams that serve as excellent visual aids. The work is written in a highly technical language that assumes familiarity with the terminology. shorthand, and conceptual framework of Generative Grammar. and this limits its accessibility and appeal to a wide readership. The inclusion of a glossary andlor a brief overview of this method of linguistic study would help address this problem. The book's title is somewhat inaccurate since it does not offer a truly comparative study of Hebrew and Arabic syntax. It is principally an analysis of Hebrew which makes reference to other languages to explicate and illustrate Shlonsky's ideas on Hebrew. While Arabic is cited more frequently than any other language besides Hebrew. there are lengthy sections of the book in which it never figures in the discussion. For example. there is not a single reference to Arabic in the thirty-page chapter treating subject-verb inversion. A further problem concerns the inconsistent manner in which Shlonsky appeals to the Arabic evidence. Throughout the book, he argues his case by referring to several different forms of the language. including Standard Arabic and the dialects from Palestine. Southern Palestine. Morocco, and Egypt. Each of these is a unique linguistic system that is distinct from the others but Shlonsky does not pay sufficient attention to the differences among them. This method hinders the purported comparative focus of his work since the reader is not sure which Arabic is meant to be the primary point of comparison. Consistent reliance upon one form of the language would have been a more beneficial approach to adopt. Such relatively minor problems do not significantly detract from the many strengths of this volume and the important contribution it makes. Shlonsky's book is required reading for anyone interested in serious study of Hebrew and will be a major work in the field. John Kaltner Rhodes College Memphis, TN 38112 kaltner@rhodes.edu THE SEMANTICS OF ASPECT AND MODALITY: EVIDENCE FROM ENGLISH AND BIBLICAL HEBREW. By Galia Hatav. Studies in Language Companion Series 34. Pp. x + 224. Philadelphia. PA: John Benjamins, 1997. Cloth, $85.00. Originally a dissertation at Tel-Aviv University. this work "aims to provide a general (semantic) theory for temporality...but it also systemati- Hebrew Studies 40 (l999) 270 Reviews cally examines the verb system in Biblical-Hebrew...which lacks tenses, as will be demonstrated, and thus enables us to see the nature of aspect and modality more clearly" (p. I). The author's theoretical assumption is "that TAM, i.e., the Tense-Aspect-Modal system in language, should be defmed within truth conditional semantics, in terms of temporality, rather than within a pragmatic approach which deals with it in terms of perspective, attitude, and the like" (although pragmatics is not ignored altogether, p. 195). Moreover, she seeks to analyze the data within the framework of a threefold distinction. Here she builds on the work of Hans Reichenbach, who argued that the contrast between the time of speech (S-time) and the time in which the event actually took place (E-time) is not sufficient to account for verbal uses. A third category is needed, namely, the time of reference (R-time), a somewhat fuzzy concept that Hatav defmes as a time unit that contains (or is contained in, or is ordered with) the E-time (pp. 3-5). In the introductory chapter the author, after explaining these and other assumptions, surveys previous attempts to account for the verbal system of biblical Hebrew; unfortunately, she seems unaware of Bruce K. Waltke and M. O'Connor, Biblical Hebrew Syntax, which gives considerable attention to verbal aspect She further informs us that she examined sixty-two chapters taken from the Pentateuch and the Former Prophets (excluding poetic material, since "poetry often violates otherwise valid linguistic norms," p. 24) and gives us some details regarding her method. Following the lead of discourse-analysis scholars, such as R. Longacre, she argues that "the biblical Hebrew system organizes the text into sequential and non-sequential material," but that in addition to sequentiality, three other temporal parameters are needed. These four parameters are individually considered in the following chapters. Chapter 2, accordingly, deals...
1019 Aminophylline-based creams are marketed as fat reducing agents for the thighs. The purposes of this study were to determine if: 1) a thigh reducing cream would influence body image, 2) product price would influence perception of effectiveness and, 3) the product would decrease thigh size. Serving as their own control, 11 women with thigh cellulite were randomly assigned to a double-blinded, counterbalanced cream treatment with 2% aminophylline (A) on one leg and a placebo (P) on the other (age: 26±7 yrs.; BMI: 23±2 kg/m2; Body Fat (BF): 24±4%; VO2: 39±4 ml/kg/min). They were also randomly assigned into a fictitious expensive (E) or inexpensive (I) cream treatment group. In the lab, subjects massaged 4.5 g of either A or P cream for 1 min into each thigh 5 d/wk for 6 wks. Pre/post testing was done 1 wk prior and immediately after the 6 wk intervention. Dependent measures included thigh girth and skinfolds (distal, mid, proximal to patella) and psychological indices (mood, body image, rating of effectiveness). Results: No significant differences (NSD) were found at baseline between the E vs. I groups for VO2, %BF, or BMI (Independent T's, p>.20). NSD were found for the right vs. left thigh measures at baseline, nor for pre-post measures of VO2, %BF, or BMI (Paired T's, p>.22). NSD were found for skinfold and girth measures between A vs. P-treated thighs (Paired T's, p>.10). Subjects in both the E and I groups reported improved thigh image (Friedman 2-way ANOVA, p=.046), but neither the E nor I group were influenced by product cost (Mann-Whitney U-Wilcoxon Rank Sum Test, p>.17). Conclusions: The 2% A cream was not effective in reducing thigh size, however, the subjects felt more positive about their thighs and overall body build. (Partially supported by FAU Foundation)
767 Electroencephalographic (EEG) frontal asymmetry has been proposed as both a physiologic index and a diathesis or disposition for emotional responding (Davidson, 1992; 1994). This study used a counter-balanced repeated measures design to examine the effect of 30 minutes of cycling exercise at ∼50%V̇O2peak, compared to 30 minutes of rest, on changes in emotional response ∼30 minutes after condition to standardized negative and positive images presented by slide projection (International Affective Picture System, Lang et al., 1988a). Emotional response was measured by EEG and self-ratings of valence and arousal (Self-Assessment Manakin, Lang et al., 1988b) in 13 females and 19 males (23±3 and 25±4y) having moderate levels of cardiorespiratory fitness (V̇O2peak = 41±9 and 53±11 ml·kg−1·min−1, respectively). Consistent with Davidson's diathesis model, response frontal α (lognμV2/Hz) right-minus-left asymmetry was predicted by resting α asymmetry, independent of neutral response (negative: t(29)=1.72, p=.09; positive:t(29)=2.35, p<.03). Valence ratings were weakly related to resting α asymmetry (p<.03). The cycling condition did not alter emotional response as indicated by [1] α asymmetry (negative: F(1,28)=0.04, p=.84, η2=.002; positive: F(1,28)=0.48, p=.49, η2=.017), [2] valence (negative: F(1,28)=0.23, p=.64, η2=.008; positive: F(1,28)=0.88, p=.36, η2=.030), or [3] arousal (negative: F (1,28)=0.15, p=.71, η2=.005; positive: F(1,28)=0.18, p=.67, η2=.007). The β1 and β2 EEG spectra were similarly unaffected by condition. Also, cycling exercise failed to alter positive or negative affect (PANAS, Watson et al., 1988), but it decreased state anxiety (STAI-Y1, 10-item)(F(1,28)=6.56, p=.02, η2=.190). Results did not differ according to gender. In the cycling condition, state anxiety change was not influenced by trait anxiety (STAI-Y2) or pre-condition resting α power asymmetry (F(2,29)=1.21,p=.31, R2adj=.014). These results indicate that moderate intensity cycling exercise lasting 30 minutes does not alter emotional response as measured, despite reducing state anxiety.
Reversing a One-Way Bilingual Dictionary* Leonard Newmark "One day we will go back to Kosov[a]. That's our land." — Ramada Shaqiri, 30 March 1999 Afi fter completing a ten-year project to write an Albanian dicLtionary, published in March 1998 by Oxford University Press as the Oxford Albanian-English Dictionary and designed specifically for users who want to read Albanian and whose access language is English, I decided to prepare a companion English-Albanian dictionary for users who want to write Albanian, but I did not want to devote another ten years to that compilation. I wondered whether I could produce a useful bilingual dictionary with the reverse orientation by automatic conversion of the entries in the data files from which the first dictionary was generated. This paper is a report on the degree to which the attempt succeeded and the degree to which human intervention was required. Examples are provided to illustrate some rather surprising results, and a general conclusion is drawn for bilingual lexicography. For languages of limited worldwide commercial importance, like Albanian, it seems particularly important to use computational techniques to derive new dictionaries from lexical data files compiled for some other purpose, especially if those files are extensive and have information otherwise difficult to come by in machine-readable form. The richness of the lexical data files from which my Albanian-English dictionary was generated is evidenced by that dictionary's 75,000 entries and subentries, more than are found in any other dictionary of Albanian. Those files already provide a number of features that dis- *This paper is a reworked and expanded version of the paper I presented at the 8th EURALEX International Congress (4-8 August 1998) and published in the proceedings of that congress. 38Leonard Newmark tinguished this dictionary from many other bilingual dictionaries: 1) inclusion of large numbers of nonstandard items (marked by asterisks ) as well as all attested standard stems; 2) marking of morpheme boundaries in Albanian words; 3) inclusion of some 16,700 phrasal expressions, in particular, phrasal names, collocations, idioms, and proverbs; 4) use of large numbers of bipartite definitions with a discursive description of the sense followed, after a colon, by English synonyms exemplifying that sense; 5) inclusion of large numbers of terms for grasses, flowers, birds, and fish with their scientific definitions; 6) inclusion of a modest amount of encyclopedic information to explain words whose strictly lexical meaning would not make their use in Albanian contexts intelligible; 7) listing of the various stem forms of lexemes as separate entries in their own alphabetical position to enable readers to decipher otherwise mystifying forms encountered in actual texts; 8) indication of the specific limits of variation that leave idiomatic senses intact (e.g., in phrasal expressions, marking a verb that can appear in any of its inflected forms by giving it in citation form with a symbol (·) at the end of the stem); 9) elaborate labeling of Albanian distinctions in domain and register; 10) rendition of phrasal expressions by stylistically similar English expressions, frequently supplemented by literal translations (enclosed in quotation marks) to enable more nuanced understanding. Each entry in the plain text, flat data files from which the Albanian-English dictionary was generated is a line consisting of an Albanian word or phrase followed by a definition in English (or by a cross-reference to another line). Each line is embedded with simple visible two-letter formatting codes (e.g., HW [headword], DF [definition], TK [technical name], CO [collocation] ) immediately preceded by a period (.) and immediately followed by what will get the formatting assigned by that code. The easily redefinable codes are later translated by a set of UNIX scripts into formatting instructions in TeX, which can go directly to a printer or indirectly by translation into Post-Script files. The simplicity of such transparent and flexible coding for entering the data, in contrast with elaborate schemes requiring complex coding by experts into predefined structures,1 was initially dictated by limits typical of languages that attract little commercial interest and 'For example, those used in the architecture described by Willy Martin and Anne Tamm in "OMBI: An editor for constructing reversible lexical databases," EURALEX '96 Proceedings...
In order to preserve the basic unity of a language like Spanish, which is spoken in 20 countries, all the speakers must adopt a respectful and careful attitude when they use it orally. In the area of phonetics, educated Spanish speakers, even those who study the language, express themselves in a somewhat careless way and what is most surprising is that such negligence is accepted in Spain by the educated linguistic norm.
Multiway trees (MT, henceforth) are a common and well-understood data structure for describing hierarchical linguistic information. With the availability of large treebanks, retrieval techniques for highly structured data now become essential. In this contribution, we investigate the efficient retrieval of MT structures at the cost of a complex index---the Treegram Index.We illustrate our approach with the VENONA retrieval system, which handles the BHt (Biblia Hebraica transcripta) treebank comprising 508,650 phrase structure trees with maximum degree eight and maximum height 17, containing altogether 3.3 million Old-Hebrew words.
In order to preserve the basic unity of a language like Spanish, which is spoken in 20 countries, all the speakers must adopt a respectful and careful attitude when they use it orally. In the area of phonetics, educated Spanish speakers, even those who study the language, express themselves in a somewhat careless way and what is most surprising is that such negligence is accepted in Spain by the educated linguistic norm.
中文句結構樹資料庫(Sinica Treebank)建構的主要目的是提供中文自然語言處理研究一個具有標記語料庫的研究素材,我們可以從這個中文句結構樹資料庫中抽取語法知識,也藉由語法知識的抽取與瞭解使我們的剖析系統功能更趨完善。本文介紹中文句結構樹資料庫構建方法和步驟,從五百萬詞的中央研究院平衡語料庫(Sinica Corpus),抽取句子,以訊息為本格位語法(Information-based Case Grammar, ICG)的表達模式為基本架構,經由電腦自動剖析成結構樹,可以盡量維持結構標記的一致性,最後並加以人工修正、檢驗,以維持標記的正確性。對於歧義的句法結構形式及詞類標記,我們也提出處理的原則。
This paper explores the automatic construction of a multilingual Lexical Knowledge Base from pre-existing lexical resources. We present a new and robust approach for linking already existing lexical/semantic hierarchies. We used a constraint satisfaction algorithm (relaxation labeling) to select --among all the candidate translations proposed by a bilingual dictionary-- the right English WordNet synset for each sense in a taxonomy automatically derived from a Spanish monolingual dictionary. Although on average, there are 15 possible WordNet connections for each sense in the taxonomy, the method achieves an accuracy over 80%. Finally, we also propose several ways in which this technique could be applied to enrich and improve existing lexical databases.
We introduce a notion of training methodology space (TM space) for specifying training methodologies in the different disciplines and teaching traditions associated with computational linguistics and the human language technologies, and pin our approach to the concept of operational model; we also discuss different general levels of interactivity. A number of operational models are introduced, with web interfaces for lexical databases, DFSA matrices, finite---state phonotactics development, and DATR lexica.
This paper uses an information-based approach to conduct feature types selection for language modeling in a systematic manner.We describe a quantitative analysis of the information gain and the information redundancy for various combinations of feature types inspired by both dependency structure and bigram structure through analyzing an English treebank corpus and taking word prediction as the object.The experiments yield several conclusions on the predictive value of several feature types and feature types combinations for word prediction, which are expected to provide reliable reference for feature type selection in language modeling.
The Internet presents a potentially revolutionary tool in the dissemination of scientific information, offering many advantages to authors and audiences. However, this resource has been underutilized in psychological research because of several factors: unfamiliarity with required technology, lack of peer review, absence of an efficient centralized accessibility resource, concerns about copyright issues, and financial considerations. The present article describes the advantages of on-line presentation of research, as well as discusses various concerns about on-line publishing and the developing solutions to deal with those concerns.
To evaluate the predictive utility of Russell's two-dimensional model of affect to the experience of depression and anxiety, self-report ratings of pleasure and arousal were obtained from 200 undergraduates using the Affect Grid. Ratings of Pleasure and Arousal each accounted for significant variance in predicting depression scores on the Beck Depression Inventory and Profile of Mood States. Only ratings of Pleasure, however, were predictive of Anxiety scores on the Profile of Mood States, whereas the relationship between Arousal ratings and Anxiety scores was more complex, demonstrating possible moderation by variables consistent with a third dimension of Dominance–Submissiveness, as postulated by other investigators.
A new ambiguity representation scheme SPR(structure preference relation) is proposed in this paper, which consists of useful quantitative distribution information for ambiguous structures. Two automatic acquisition algorithms, (1) acquired from treebank, (2) acquired from raw texts, are introduced, and some experimental results which prove the availability of the algorithms are also given. At last, some SPR applications in linguistics and natural language processing are introduced and some future research directions are proposed in this paper.
ResumenCuando durante el proceso de adquisición los niños vascos comienzan a construir enunciados de dos o más palabras, atraviesan un período en el que no utilizan conocimientos gramaticales o sintácticos al construir sus producciones lingüísticas. Tras este período presintáctico en el que la producción lingüística es construida basándose en principios semántico-pragmáticos, los niños vascos, hacia la edad de 2;00, inician un desarrollo sintáctico gradual y de algún modo calificable como uniforme. El presente trabajo analiza la producción lingüística de tres niños vascos, dos bilingües vasco-castellanos y un monolingüe, que fueron videograbados quincenalmente desde 1;06 hasta 3;00 de edad, durando las sesiones unos 30 minutos. Dadas las características morfológicas del euskera, resulta relativamente sencillo identificar la utilización o no de las mismas por parte de los niños. Se observarán las unidades lingüísticas fundamentales: determinación y estructura del sintagma nominal, casos declinativos, morfología verbal, órdenes de las preguntas Qu, etc. También conviene señalar que el posterior desarrollo sintáctico que tienen lugar no surge de la utilización de los principios semántico-pragmáticos, sino que más bien es independiente de ellos.AbstractWhen Basque children begin to produce multiword utterances during their language acquisition process, they go through a first period characterized by a lack of grammatical or syntactic knowledge when constructing their language productions. After this presyntactic period, when language production is mainly based on semantic-pragmatic principles, Basque children show, towards the age of 2;00, a gradual syntactic development which could be considered as uniform. The present study analyses the language production of three Basque children, two of them Basque-Spanish bilinguals, and the third a Basque monolingual, videotaped fortnightly from the age of 1;06 until 3;00, in 30 minute sessions. Due to the pronounced nature of Basque morphology, it is not difficult to determine whether it is used or not by these children. We will take note of the most important morphological characteristics: determination and structure of noun phrases, case markings, verbal morphology, word order in Wh-questions, etc. We would also like to point out that the subsequent syntactic development does not emerge from semantic-pragmatic principles. On the contrary, it is quite independent.Extended SummaryIt can be affirmed that in the process of acquisition of Basque, children go through a pre-syntactic phase during which they do not use grammatical properties, nor do they base their language production on syntactic rules. The subsequent syntactic development begins towards the age of two and seems to develop gradually, according to basic sentence structure. it is the mental, neurological maturity, activating the biologically inherited principles of Universal Grammar, which makes possible the development of the grammar process, since grammar is independent of the semantic-pragmatic principles which children seem to be using during the previous pre-syntactic period.I have elaborated this working hypothesis on the basis of other studies and investigations analysing different L1 acquisition processes. I would like to point out the importance of studies carried out in this area by: Bickerton (1990), theorising on general acquisition processes; Radford (1989), who, analysing early English, observed the lack of INFL and COMP functional categories; Platzack (1990), who arrived at similar conclusions observing the process of acquisition of Swedish; Meisel (1992), who, investigating the acquisition of French and German by bilingual children, established an early stage of language lacking the above-mentioned functional categories, etc.In order to give support to our hypothesis I have observed the language production of three Basque children, two of them bilinguals and the third one monolingual. They have been video-taped fortnightly from the age of 1;06 to the age of 3;00, in 30 minute sessions.Basque is a morphologically clearly marked language as to case markings as well as verbal conjugation. This fact simplifies the identification of the use of differentiated morphological elements. I have analysed determiners and internal structure in the process of NP acquisition; case marking: so-called grammatical markings (with verbal agreement) as well as non-grammatical ones; the verbal aspect; triple verbal agreement (subject, direct object, and indirect object); verbal tense and mode; subordinating conjunctions (temporal and non-temporal); word order in WH-questions, and word order in the first declarative utterances.After describing these acquisition processes, i can state that there evidently exists an early, pre-syntactic stage, during which the children do not use case markings or syntactic structures. After this period, the three children analysed show a gradual syntactic development, which can somehow be qualified as uniform, towards the age of 2;00. The three children first start using the VP structure and NP determiners, difference case markings and auxiliary verbs. The use of the functional category INFL, which permits them to use subject agreement, comes later. Subsequently, the acquisition of the functional category COMP, permits them to use subordinating temporal conjunctions or the construction of WH-questions according to adult norms. i should like to point out that due to the complexity of INFL in Basque, it seems evident that the use of this category by the children during the language acquisition process will be gradual.As to the pre-syntactic period, during which not only lexical acquisition but also the acquisition of different lexical categories is evident, I can say that a great part of the construction of two-or-more word utterances follows semantic-pragmatical principles, which are quite normal in adult spoken Basque. Since these principles are used during the whole of the following syntactic development, always adjoining the theme to the left of the structure used at each moment, i.e., at the beginning of the utterance, it seems evident that the acquisition of syntax is independent of the semantic-pragmatic principles used.Palabras clave: Adquisición de lenguajecategorías funcionales (INFL, COMP)desarrollo sintáctico gradualeuskeraperíodo presintácticoprincipios semántico-pragmáticosKeywords: Language acquisitionfunctional categories (INFL, COMP)gradual syntactic developmentBasque languagepresyntactic periodsemantic-pragmatic principles
Statistical approaches to processing Lexical Functional Grammars (LFG-DOP, [1]) require large corpora of text annotated with c-structure and f-structure representations. To date, such corpora that exist are constructed manually or semi-automatically. Manual construction is both time-consuming and error-prone. Semi-automatic construction usually proceeds as follows: an existing LFG grammar is used to parse input text. Typically, for each sentence in the input text, parsing will produce a large number of c- and f-structure analyses. A linguistic expert then inspects the analyses and for each sentence in the input text selects the single best analysis for the case at hand. For large grammars this can involve inspection of hundreds or thousands of proposed analyses for a single input sentence. In the present paper we develop an alternative, semi-automatic methodology that as much as possible avoids manual inspection of analyses for best fit. As input, the method requires a treeb...
This article examines the reliance of U.S. campuses on international teaching assistants (ITA) for staffing undergraduate course and the strategies that may affect ratings of their speaking competence. This increasing reliance has led to student complaints about incomprehensibility of ITA. This problem has been examined by looking through the eyes of the students, administrators and taxpayers. Therefore, the responsibility had been placed on the ITA, whose burden it was to learn the language and culture more fully. The goal of having ITA learn the language and culture better was eclipsing another important issue that needs consideration, the issue of teaching assistant's feelings of loss of control over the students' perception about themselves.
This paper provides normative data for Australian school children on a modified version of the Castles Word/Non-Word Tests (Castles, 1993). The tests were designed to isolate the lexical and nonlexical reading procedures. Data were collected from 298 school children in Perth and combined with data provided by Coltheart and Leahy (1996) for 420 school children in Sydney. Norms for the Sydney sample have been published previously (Coltheart & Leahy, 1996). Norms for the combined sample are reported in 12-month age bands from 7 to 12 years, in the form of normalised standard scores. Issues surrounding subtyping research in dyslexia are reviewed, and a way to subclassify research samples using the provided norms is outlined and evaluated.
We present a novel approach for measuring body size estimation in normal and eating-disordered women and men. Clinical categories of body types were used as prototypes. By comparing the subjective appearance of a person’s body with prototypes, we can understand how different attributes of his or her body shape contribute to perception of body size. After lifelike random distortions have been applied to parts of their body image, individuals adjust their body shapes until they converge on their perceived veridical appearance. Exaggeration and minimization of particular body areas measured with respect to their true shape and with different prototypes can be expressed as numerical deviations. In this way, perceived body size and body attractiveness can be appraised during the course of diagnosis and treatment of eating disorders.
The paper presents the results of a series of Principal Components Analyses of the frequencies of very common words in the dialogue of characters in plays by Ben Jonson. The first Principal Component in the data, the most important axis of differentiation, proves in each case to be a spectrum from elaborate, authoritative pronouncements to a dialogue style of reaction and interchange. Reference to other quantitative studies, literary and otherwise, suggests that a version of this axis may often be among the most important in stylistic difference generally. In Jonson it has a chronological aspect -- there is a shift over his career from one end to the other -- and there is often significant change within the idiolects of his characters as well. Successive segments of Volpone and Mosca's parts (they are protagonist and antagonist of Volpone, perhaps Jonson's best-known comedy) change markedly along this axis, beginning far apart but coming by the end of the play to resemble each other very closely on this measure.
There is a general concern within the field of word sense dusamb~guatmn about the rater-annotator agreement between human annota tors. In thus paper, we examine th~s msue by comparing the agreement rate on a large corpus of more than 30,000 sense-tagged instances Thin corpus us the mtersectmn of the WORDNET Semcor corpus and the DSO corpus, which has been independently tagged by two separate groups of human annotators The contribution of this paper us two-fold First, ~t presents a greedy search algori thm tha t can automatical ly derive coarser sense classes based on the sense tags assigned by two human annotators The resulting derived coarse sense classes achmve a h~gher agreement rate but we s t f l!mamtam as many of the original sense classes as posmble Second, the coarse sense grouping derived by the algorithm, upon verification by human, can potent ial ly serve as a better sense inventory for evaluating automated word sense d~samb~guatmn algori thms Moreover, we examined the derived coarse sense classes and found some interesting groupings of word senses that correspond to human mtmtlve judgment of sense granularity 1 I n t r o d u c t i o n. It us widely acknowledged that word sense d~samblguatmn (WSD) us a central problem m natural language processing In order for computers to be able to understand and process natural language beyond simple keyword matching, the problem of d~samblguatmg word sense, or dlscermng the meamng of a word m context, must be effectively dealt with Advances in WSD v, ill have slgmficant Impact on apphcatlons hke information retrieval and machine translation For natural language subtasks hke part-of-speech tagging or s)ntactm parsing, there are relatlvely well defined and agreed-upon cnterm of what it means to have the correct part of speech or syntactic structure assigned to a word or sentence For instance, the Penn Treebank corpus (Marcus et a l, 1993) pro~ide~,t large repo.~tory of texts annotated w~th partof-speech and s}ntactm structure mformatlon Tv.o independent human annotators can achieve a high rate of agreement on assigning part-of-speech tags to words m a g~ven sentence Unfortunately, th~s us not the case for word sense assignment F~rstly, it is rarely the case that any two dictionaries will have the same set of sense defimtmns for a g~ven word Different d~ctlonanes tend to carve up the semantic space m a different way, so to speak Secondly, the hst of senses for a word m a typical dmtmnar~ tend to be rather refined and comprehensive This is especmlly so for the commonly used words which have a large number of senses The sense dustmctmn between the different senses for a commonly used word m a d~ctmnary hke WoRDNET (Miller, 1990) tend to be rather fine Hence, two human annotators may genuinely dusagree m their sense assignment to a word m context The agreement rate between human annotators on word sense assignment us an Important concern for the evaluatmn of WSD algorithms One would prefer to define a dusamblguatlon task for which there us reasonably hlgh agreement between human annotators The agreement rate between human annotators will then form the upper ceiling against whmh to compare the performance of WSD algorithms For instance, the SENSEVAL exerclse has performed a detaded s tudy to find out the raterannotator agreement among ~ts lexicographers taggrog the word senses (Kllgamff, 1998c, Kllgarnff, 1998a, Kflgarrlff, 1998b) 2 A C a s e S t u d y In th i s -paper, we examine the ~ssue of raterannotator agreement by comparing the agreement rate of human annotators on a large sense-tagged corpus of more than 30,000 instances of the most frequently occurring nouns and verbs of Enghsh This corpus is the intersection of the WORDNET Semcor corpus (Miller et a l, 1993) and the DSO corpus (Ng and Lee, 1996, Ng, 1997), which has been independently tagged wlth the refined senses of WORDNET by two separate groups of human annotators The Semcor corpus us a subset of the Brown corpus tagged with ~VoRDNET senses, and consists of more than 670,000 words from 352 text files Sense taggmg was done on the content words (nouns, ~erbs, adjectives and adverbs) m this subset The DSO corpus consists of sentences drawn from the Brown corpus and the Wall Street Journal For each word w from a hst of 191 frequently occurring words of Enghsh (121 nouns and 70 verbs), sentences containing w (m singular or plural form, and m its various reflectional verb form) are selected and each word occurrence w ~s tagged w~th a sense from WoRDNET There ~s a total of about 192,800 sentences in the DSO corpus m which one word occurrence has been sense-tagged m each sentence The intersection of the Semcor corpus and the DSO corpus thus consists of Brown corpus sentences m which a word occurrence w is sense-tagged m each sentence, where w Is one of.the 191 frequently oc,currmg English nouns or verbs Since this common pomon has been sense-tagged by two independent groups of human annotators, ~t serves as our data set for investigating inter-annotator agreement in this paper
For online yellow pages and product catalogs, structured content representations coupled with linguistic ontologies can increase both recall and precision of content based retrieval. Our OntoSeek system adopts a language of limited expressiveness for content representation and uses a large ontology based on WordNet for content matching. WordNet is a linguistic database formed by synsets-terms grouped into semantic equivalence sets, each one assigned to a lexical category (noun, verb, adverb, adjective). Each synset represents a particular sense of an English word and is usually expressed as a unique combination of synonymous words.
This research assessed an interactive satellite-based training program integrating interactive audiovisual experiences with face-to-face interactions. Key elements were content created by experts, high-quality video segments, satellite-based interaction, off-line interactions among teams of parents and caregivers, workshops, and team building exercises. For pragmatic reasons, it was necessary to develop brief assessment instruments concurrently with training. A large set of survey items were created from draft materials and reduced empirically through piloting to those with the best psychometric properties. To avoid the appearance of traditional testing, knowledge was assessed with Likert items. Surveys measured participant satisfaction, knowledge, attitudes, and the application and articulation of concepts. Participant satisfaction was high. Participants increased positive attitudes and learned appropriate vocabulary. Training was more effective than no training or watching videotapes. The program appears to represent a viable model of training that could successfully be applied to Internet technologies.
Object Relocation is a computer program for Windows 95, with which experiments on spatial memory for object locations can be designed, run, and analyzed. Because of its clear graphical user interface, no long and complex command syntax is needed. Basically, a stimulus consists of a frame that contains a chosen number of locations (i.e., the actual spatial layout) to which objects can be assigned. When the experiment is run, these stimuli are presented to the subject for a variable period of time. Subsequently (either with or without a delay), the objects are presented in a row above the frame and have to be relocated to the correct positions. Finally, the raw data can be analyzed efficiently, using various error scores, and an SPSS-ready output file can be produced. Object Relocation is a very flexible program: New objects and positions can easily be added, and various options for presentation and relocation are present.
Rapid expansion in the digitization of image and image collections has vastly increased the numbers of images available to scholars and researchers through electronic means. This research review will familiarize the reader with current research applicable to the development of image retrieval systems and provides additional material for exploring the topic further, both in print and online. The discussion will cover several broad areas, among them classification and indexing systems used for describing image collections and research initiatives into image access focusing on image attributes, users, queries, tasks, and cognitive aspects of searching. Prospects for the future of image access, including an outline of future research initiatives, are discussed. Further research in each of these areas will provide basic data which will inform and enrich image access system design and will hopefully provide a richer, more flexible, and satisfactory environment for searching for and discovering images. Harnessing the true power of the digital image environment will only be possible when image retrieval systems are coherently designed from principles derived from the fullest range of applicable disciplines, rather than from isolated or fragmented perspectives.
We are experimenting with the representation of a DTD and associated documents (i.e., documents conformant to the DTD) in a knowledge representation (KR) system, in order to provide more sophisticated query and retrieval from TEI documents than current systems provide. We are using CLASSIC, a frame-based representation system developed at AT&T Bell Laboratories. Like many KR systems, CLASSIC enables the definition of structured concepts/frames, their organization into taxonomies, the creation and manipulation of individual instances of such concepts, and inference such as inheritance, relation transitivity, inverses, etc. In addition, CLASSIC provides for the key inferences of subsumption and classification. By representing a document as an individual instance of a hierarchy of concepts derived from the DTD, and by allowing the creation of additional user-defined concepts and relations, sophisticated query and retrieval operations can be performed. This paper describes CLASSIC and the formalism of description logic that underlies it, and demonstrates how it can be used for enhanced retrieval from richly encoded documents.
This paper report on some of the concrete outcomes of a larger research project on the study of syntactic change. In this part of the project, we are collecting and encoding historical texts and tagging them for syntactic analysis. We have so far produced a TEI-conformant version of an Old French text, La Vie de Saint Louis written by Jehan de Joinville around 1305, and we are in the process of adding syntactic tags to this text. Those syntactic tags are derived from the Penn-Helsinki coding scheme, which had been devised for the syntactic encoding of Middle English texts, and have been translated into TEI.
Mylonas and Renear introduce a volume of selected papers from The Text Encoding Initiative 10th Anniversary Conference, held at Brown University in November 1997. The Text Encoding Initiative (TEI), was launched in 1987 and sponsored by the Association for Computers and the Humanities, the Association for Literary and Linguistic Computing, and the Association for Computational Linguistics. It had as its original objective the development of an interchange language for textual data. This effort was completely successful and the TEI Guidelines are now widely accepted as the standard interchange format for textual data. Mylonas and Renear also note that the TEI has accomplished two other major achievements: it has produced a powerful new data description language (which is influencing the development of new WWW standards); and, most importantly, it has motivated the development of an entirely new research community, focused on understanding the role of text structure and markup in the use of emerging information technologies in culture, scholarship, and communication.
Latent semantic analysis (LSA) serves as both a theory and a method for representing the meaning of words based on a statistical analysis of their contextual usage (Foltz, 1996; Landauer & Dumais, 1997). In experiments in the domains of psychology and history, we compared the representation of readers’ knowledge structures of information learned from texts with the representation generated by LSA. Results indicated that LSA’s representation is similar to readers’ representations. In addition, the degree to which the reader’s representation is similar to LSA’s representation is indicative of the amount of knowledge the reader has acquired and of the reader’s reading ability. This approach has implications both as a model of learning from text and as a practical tool for performing knowledge assessment.
This article explores the importance of register variation for analyses of grammar and discourse. The general theme is illustrated through consideration of variability in the form and use of English complement clauses. First, the patterns of use for four related grammatical constructions are considered: that-clauses and to-clauses, headed by verbs and by nouns. The differing discourse functions of each construction type are explored by considering their lexico-grammatical associations (i.e. the verbs or nouns most commonly occurring as the head of each type). However, it is shown that the characteristic uses of each type are conditioned by register. That is, each construction type has a different distribution across spoken and written registers, with a different set of associated lexical heads. A second study provides an even more striking illustration of this interaction between grammar, discourse, and register: the contextual factors conditioning the retention vs omission of the complementizer that. In this case, it is shown that each register has an overall norm, and that contextual factors are influential only when they work in opposition to that register norm. These case studies are presented to make the general point that analyses of grammar and discourse are often inadequate and misleading when they disregard register differences. Instead, a register perspective is required to capture the range of variability associated with grammatical patterns of use.
We present a memory-based learning (MBL) approach to shallow parsing in which POS tagging, chunking, and identification of syntactic relations are formulated as memory-based modules. The experiments reported in this paper show competitive results, the F-value for the Wall Street Journal (WSJ) treebank is: 93.8% for NP chunking, 94.7% for VP chunking, 77.1% for subject detection and 79.0% for object detection.
This paper develops a solution to the problem of importing existing TEI data into an existing object-oriented database schema without changing the TEI data or the database schema. The solution is based on architectural processing. Two meta-DTDs are used, one to define the architectural forms for the object model and another to map the existing SGML data onto those forms. A full example using a critical text in TEI markup is developed.
Elementary dependency relationships between words within parse trees produced by robust analyzers on a corpus help automate the discovery of semantic classes relevant for the underlying domain. We introduce two methods for extracting elementary syntactic dependencies from normalized parse trees. The groupings which are obtained help identify coarse-grain semantic categories and isolate lexical idiosyncrasies belonging to a specific sublanguage. A comparison shows a satisfactory overlapping with an existing nomenclature for medical language processing. This symbolic approach is efficient on medium size corpora which resist to statistical clustering methods but seems more appropriate for specialized texts.
A study commissioned by the Canadian Institute for Historical Microreproductions produced some interesting secondary findings about the attitudes of the Canadian research community towards digitized facsimile collections. In written responses to a questionnaire designed primarily to elicit advice about the subject content and focus of future projects, and in structured follow-up interviews, many respondents demonstrated a marked ambivalence towards the concept of digitized collections. Furthermore, if faced with a choice between fully searchable text and digitized facsimile images with traditional points of access (subject, author, title, etc.), there appears to be a preference for the latter means of access.
Electronic texts are claimed to exhibit features distinct from their more tangible cousins. The Snapshot project aims to observe and capture language usage in an electronic medium by creating an open corpus of World Wide Web documents. These documents are re-encoded using the TEI guidelines to create a flexible, persistent and portable data repository. This report gives an overview of the decisions made with respect to the re-encoding of HTML documents, and with the structuring the overall corpus.
The journals of the Psychonomic Society have served as outlets for numerous stimulus norms and ratings. Such norms are useful to researchers in a variety of areas for manipulating and controlling stimulus attributes. This article presents an index of 142 norms published in the Society’s journals, categorized according to the types of materials and ratings that are included in each.
This paper describes morphing techniques to manipulate two-dimensional human face images and three-dimensional models of the human head. Applications of these techniques show how to generate composite faces based on any number of component faces, how to change only local aspects of a face, and how to generate caricatures and anticaricatures of faces. These techniques are potentially useful for many psychological studies because they permit realistic images to be generated with precise control.
Boundary extension refers to a tendency to remember seeing a greater expanse of a scene than was shown in a photograph. It is hypothesized that the view shown in the stimulus activates expectations about the scene’s layout just outside the picture’s borders. Following presentation, the viewer remembers having seen this expected information, and this yields boundary extension. We provide photographs and instructions for conducting two brief demonstrations of the phenomenon and provide materials for a related class experiment on the journal’s World-Wide Web site. These demonstrations of boundary extension provide graphic illustrations of the role of schematic expectancies in the representation of scenes and help to illustrate the role of real-world knowledge in cognition.
A number of studies in perception, attention, and memory employ signal detection theory (SDT) to assess the accuracy of an observer’s detection or discrimination performance. Some of the problems that students have with understanding and using SDT are associated with the calculations needed to obtain SDT parameters and predictions. All of these calculations, plus the simulation of SDT processes, can be performed using a spreadsheet application program, such as Excel or Quattro Pro. This paper offers a short tutorial on how to use a spreadsheet program to increase your students’ knowledge and understanding of SDT.
The use of computerized psychological assessment is a growing practice among contemporary mental health professionals. Many popular and frequently used paper-and-pencil instruments have been adapted into computerized versions. Although equivalence for many instruments has been evaluated and supported, this issue is far from resolved. This literature review deals with recent research findings that suggest that computer aversion negatively impacts computerized assessment, particularly as it relates to measures of negative affect. There is a dearth of equivalence studies that take into account computer aversion’s potential impact on the measurement of negative affect. Recommendations are offered for future research in this area.