Papers reviewed and determined not to be word norm studies. Use the flag icon to report errors or suggest re-inclusion.
16504 papers
Reviewed by: Syntax of the Modern Greek Verbal System: The Use of the Forms, Particularly in Combination with θα and να Panayiotis A. Pappas Rolf Hesse, Syntax of the Modern Greek Verbal System: The Use of the Forms, Particularly in Combination with θα and να. Second revised edition. Copenhagen: Museum Tusculanum Press. 2003. Pp. 141. $37.00. This monograph is an in-depth description of the usage of the verbal forms of Modern Greek. It is an update of the 1980 edition in several important aspects, such as the corpus of texts, which now includes examples from more modern texts, and in its terminology which takes into account some of the more recent developments in the approach to Greek grammar. The book is organized into ten chapters and also includes a foreword, which gives the reasons for publishing a revised edition, a selected bibliography and an eclectic index of Greek words and English grammatical terms. The brief introduction deals mostly with issues of terminology, while chapters 2 and 3 are terse accounts of the inflectional categories of Greek, and the usage of person and number respectively. The book hits its stride in chapter 4, which discusses the distinction between the active and the mediopassive voice and demonstrates how the two forms are used. Chapter 5 begins with a detailed explanation of aspect in Greek, which not only reviews the conceptual differences between perfective and imperfective but also presents the syntactic constraints that are associated with their use. The author proceeds by listing the different verb forms according to aspect, tense and mood (e.g., imperfective non-past, perfective non-past, imperative, etc.). For each form, he provides a brief description to how each form is used and follows it with several examples. The chapter also includes a relatively lengthy discussion of the meaning of the perfective non-past and an excursus on special constructions. Chapter 6 contains some general remarks on the markers θα, να and αvν, which are the topics of chapters 7 and 8, forming the main body of the book (65 of 141 pages). Chapter 7 begins with a description of the usage of θα constructions to express futurity versus their role in inferential expressions, and then continues by cataloguing the function of θα constructions that employ specific verb forms (θα + imperfective past, θα + perfect etc.). Chapter 8 is organized according to the grammatical elements that "govern" the να clause, while each subsection is structured around the use of να with specific verb forms or other lexical items. The author is very thorough in enumerating the possible combinatory possibilities with να and takes great care in combing through all the possible functions and meanings that να constructions can have in a sentence. The book concludes with two more brief chapters, one on indirect speech (chapter 9) and one on the use of negative words (chapter 10). The book is intended for advanced students of the language and scholars. One needs to be fairly well-acquainted with the complexities of the Greek verb system in order to benefit from its detailed and well-exemplified discussion of usage norms. The author's laconic style of writing, which lets the examples carry the weight of the description, also makes this a book best suited for experienced learners of Greek. Those who delve into it, I am certain will find several points of interest, especially in the description of the various constructions that employ θα and να. In these chapters, the organization of the content according to the function of each construction, but also by the forms that are combined (e.g., θα [End Page 213] + imperfective past) makes for a handy reference tool. The use of a single verb () to showcase the distinctions between the active and mediopassive voice (22), the observation that only the perfective non-past form and not the imperfective has the meaning "it's gone, it's over" (35), and the discussion of lesser-known imperative constructions (54-55), are just some of the interesting observations that one finds in this book. The intended audience will certainly appreciate the detailed exploration in terms of function and meaning. However, the book suffers from several flaws. The production is far from perfect. First, the back cover has να misspelled as υα, which is rather unfortunate because the...
It is important that a book on adolescence should contain a chapter on slang and swearing because these are two features of adolescent linguistic behavior that attract more attention than they perhaps warrant, symbolizing as they do, the freedom that young people have at this stage of their lives to challenge linguistic norms, and at the same time to test interpersonal bonds and institutional constraints with their parents as they seek to establish new identities and relationships within their changing worlds. This chapter will start by providing a brief background description of the linguistic characteristics of slang and swearing before moving to a more focused discussion of why (although such language is also commonly used by many adult groups who find themselves living in close proximity in institutionalized contexts such as prisons or military camps), it may be particularly typical of adolescent speech. Subsequent sections of the chapter explore how the use of slang and expletives by young people can be viewed as a means of building cultural capital within their networks, while at the same time establishing boundaries between new adolescent in-groups ('us') and parents and nonmembers ('them'). A section on the use of pejorative terms further develops this theme of 'us' and 'them.' In addition, I also consider gender-based patterns of the ways young people use slang and expletives, and the subtle coercive effect such words have in regulating typical patterns of in-group behavior. Finally some suggestions are made regarding research questions that still require attention. (PsycINFO Database Record (c) 2016 APA, all rights reserved)
Using role play and verbal‐report data, this study investigates the sequential organization of politeness strategies of 24 learners of Spanish and whether the learners’ ability to negotiate and mitigate a refusal was influenced by length of residence in the target community. Refusal sequences were examined throughout the interaction (head acts, pre‐ and postrefusals) and across conversational turns. Results showed more frequent attempts at negotiation and greater use of lexical and syntactic mitigation among learners who had spent more time in the target community and also revealed a preference for solidarity and indirectness, which approximated native Spanish speaker norms. It is suggested that the variables of proficiency and length of residence should be considered independently. Finally, learners’ perceptions of social status are discussed.
М. И. ШАПИР... ЭСТЕТИКА НЕБРЕЖНОСТИ В ПОЭЗИИ ПАСТЕРНАКА (Идеология одного идиолекта)1... © 2004 г.... В первой части работы анализируются случаи непроизвольных двусмысленностей в поэзии Пастернака (лексико-фразеологических, грамматических, стилистических); во второй части они рассматриваются в ряду разного рода коллоквиализмов, нарушающих привычные нормы книжного языка и классической ритмики; наконец, в третьей части статьи делается попытка понять психологическую, эстетическую и социальную подоплеку общей установки поэта на опрощение стихотворного языка и его сближение с разговорной речью.... The first section of this article analyzes cases of involuntary ambiguity (lexical, phraseological, grammatical, and stylistic) in the poetry of Pasternak. The second section examines these cases in the context of various colloquialisms which violate the conventional norms of standard literary Russian and Russian classical prosody. The third and the last section attempts to understand the psychological, aesthetic and social background of Pasternak's general tendency towards the simplification of poetic language and its convergence with free colloquial speech.......А ты прекрасна без извилин...... Много лет назад Уильям Эмпсон, незаурядный английский поэт и филолог, расценил семантическую неопределенность (ambiguity) как неотъемлемое свойство поэзии [I]2. В последнее время интерес к поэтической неоднозначности растет и у российских лингвистов: недавно специальное исследование ей посвятил Н. В. Перцов [2] (ср. [3]). В своей книге и в предшествующих статьях я тоже обращался к этой теме (см. [4, с. 12 - 19] и др.). Но до сих пор филологи сосредоточивались главным образом на преднамеренном двоении смыслов; что же касается двусмысленностей непроизвольных (либо кажущихся таковыми), то им должного внимания не уделялось. Это упущение мне бы хотелось восполнить: сначала предметом моего анализа станет такое ветвление у Пастернака, которое, насколько можно судить, не входило в расчеты автора; затем найденные факты я попробую поставить в более широкий лингвистический и наконец - в идеологический контекст. Таким образом, против обыкновения я буду изучать не информацию, а шум, который, однако, на свой лад оказывается весьма информативным.... Размышляя над примерами, постараемся не терять из виду суть проблемы: дело не в том, что какой-то фрагмент текста не допускает верной интерпретации, - дело в том, что он объективно допускает интерпретацию неверную. Именно ощущение неадекватности вторых и третьих смыслов позволяет нам выделять оговорки среди других случаев неоднозначности. Разумеется, это ощущение может сбивать с толку: насчет авторского замысла нам дано лишь строить догадки. Наивно было бы верить, что в поэтическом тексте намеренное всегда надежно отличается от ненамеренного: в душе писателя мы читать не умеем, но попытаться его понять - обязаны3.... Явление, о котором пойдет речь, еще не имеет адекватного терминологического выражения.... 1 Исследование выполнено при поддержке Российского гуманитарного научного фонда (проект 04 - 04 - 00055а). Исправляя и дополняя исходный вариант статьи, автор имел счастливую возможность пользоваться советами и замечаниями М. В. Акимовой, С. Г. Болотова, М. Л. Гаспарова, Ф. Н. Двинятина, В. З. Демьянкова, А. А. Добрицына, И. Г. Добродомова, А. К. Жолковского, Вяч. Вс. Иванова, А. А. Илюшина, Т. М. Левиной, Т. М. Николаевой, А. Б. Пеньковского, И. А. Пильщикова, Н. В. Перцова, О. Ронена, Т. В. Цивьян. Особая признательность Е. Б. Пастернаку и Е. В. Пастернак, помогавшим автору и его поддерживавшим на протяжении всей работы.... 2 Латинское слово ambiguitas 'двусмысленность' соответствует древнегреческому (амфиболия), усвоенному русской научной терминологией.... 3 На пушкинском пленуме Союза писателей (1937) Пастернак заявил: не только намеренных двусмысленностей, но и таких провалов последнего сорта, которые бы давали повод для двусмысленного понимания и в неумышленном плане, - я за собой не помню. Вообще двусмысленности при настоящей любви к искусству немыслимы [5, т. 4, с. 644; 6, с. 401, 404 примеч. 36].... стр. 31... В арсенале испытанных средств филологического метаязыка наиболее подходящим к случаю могло бы стать понятие авторской глухоты, закрепленное в Поэтическом словаре А. П. Квятковского. Это условный термин, предложенный М. Горьким; понимаются под ним явные стилистические и смысловые ошибки..> не замеченные автором. Их можно трактовать по-разному: иногда авторская глухота - результат небрежности или неряшливости, иногда она возникает непроизвольно, когда увлечение главной задачей заслоняет отдельные детали. Явления г, - продолжает Квятковский, - свойственны не только рядовым писателям, но и большим мастерам [7, с. 10]. Он приводит примеры из Пушкина, Лермонтова, Плещеева, Фета, Маяковского, Багрицкого и Уткина. Завершается статья указанием на то, что к А г можно отнести явления сдвига, и ссылками на тематически близкие статьи: Амфиболия, Анаколуф, Солецизм [7, с.
No AccessPerspectives on School-Based IssuesArticle1 Dec 2004Evidence-Based Practices: Empirical Evidence John R. Muma and Steven J. Cloud John R. Muma University of Southern MississippiHattiesburg, MS Google Scholar and Steven J. Cloud University of Southern MississippiHattiesburg, MS Google Scholar https://doi.org/10.1044/sbi5.4.16 SectionsAboutFull TextPDF ToolsAdd to favoritesDownload CitationTrack Citations ShareFacebookTwitterLinked In References American Psychological Association (1974). Standards for educational and psychological tests.: Washington, DC: American Psychological Association. Google Scholar Ainsworth, M. (1969). Object relations, dependency and attachment: A theoretical view of the mother-child relationship.Child Development, 40, 969–1025. Google Scholar Ainsworth, M. (1972). Attachment and dependency: A comparison.In J Gewirtz (Ed.), Attachment and dependency (pp. 97–137). Washington, DC: Winston. Google Scholar Ainsworth, M. (1973). The development of infant-mother attachment.In B. Caldwell & H Ricciuti(Eds.), Review of child development research (vol. 3, pp. 1–94). Chicago: University of Chicago Press. Google Scholar Ainsworth, M. (1974). Infant-mother attachment and social development: Socialization as a product of reciprocal responsiveness to signals.In M. Edwards (Ed.), The integration of the child into the social world.: Cambridge, UK: Cambridge University Press. Google Scholar Bloom, L. (1973). One word at a time.: The Hague: Mounton. Google Scholar Bowlby, J. (1969). Attachment and loss: Vol. I. Attachment.: New York: Basic Books. Google Scholar Bowlby, J. (1973). Attachment and loss: Vol. II. Separation.: New York: Basic Books. Google Scholar Brown, R. (1973). A first language: The early stages.: Cambridge, MA: Harvard University Press. Google Scholar Brown, R. (1977). Introduction.In C Snow & C Ferguson(Eds.), Talking to children (pp. 1–27). New York: Cambridge University Press. Google Scholar Brown, R. (1978). (Disclaimer in Figure 6.In J Muma,Language handbook (p. 189). Englewood Cliffs, NJ: Prentice-Hall. Google Scholar Brown, R. (1988). Appendix: Roger Brown, An autobiography in the third person.In F Kessel(Ed.), The development of language and language researchers (pp. 395–400). Hillsdale, NJ: Erlbaum. Google Scholar Bruner, J. (1981). The social context of language acquisition.Language & Communication, 1,155–178. CrossrefGoogle Scholar Bruner, J. (1986). Actual minds, possible worlds.: Cambridge, MA: Harvard University. Google Scholar Cazden, C. (1988). Environmental assistance revisited: Variation and functional equivalence.In F Kessel(Ed.), The development of language and language researchers (pp. 281–298). Hillsdale, NJ: Erlbaum. Google Scholar Conant, S. (1987). The relationship between age and MLU in young children: A second look at Klee and Fitzgerald's data.Journal of Child Language, 14,169–173. MedlineGoogle Scholar Eisenberg, S., McGovern Fersko, T., & Lundgren, C. (2001). The use of MLU for identifying language impairment in preschool children: A review.American Journal of Speech-Language Pathology, 10, 323–342. LinkGoogle Scholar Ferguson, C., & Garnica, O. (1975). Theories of phonological development.In E Lenneberg & E Lenneberg (Eds.), Foundations of language development (pp. 149–180). New York: Academic. Google Scholar Fey, M. (1986). Language intervention with young children.: San Diego, CA: College-Hill Press. Google Scholar Gleitman, L. (1994). The structural sources of verb meanings.In P Bloom(Ed.), Language acquisition: Core readings (pp. 174–221). Cambridge, MA: MIT Press. Google Scholar Greenfield, P., & Smith, J. (1976). Communication and the beginnings of language.: New York: Academic Press. Google Scholar Guion, R. (1977). Content validity: Three years of talk—What's the action? Public Personnel Management, 6, 407–414. Google Scholar Halliday, M. (1975). Learning how to mean.In E Lenneberg & E Lenneberg(Eds.), Foundations of language development: A multidisciplinary approach (pp. 239–266). New York: Academic Press. Google Scholar Ingram, D. (1976). Phonological disability in children.: New York: Elsevier. Google Scholar Lahey, M. (1988). Language disorders and language development.: New York: MacMillan. Google Scholar Lahey, M. (1994). Grammatical morpheme acquisition: Do norms exist?.Journal of Speech and Hearing Research, 37, 1192–1194. Google Scholar Leonard, L. (1981). Facilitating linguistic skills in children with specific language impairment.Applied Psycholinguistics,, 2, 89–118. CrossrefGoogle Scholar Leonard, L. (1998). Children with specific language impairment.: Cambridge, MA: MIT Press. Google Scholar Messick, S. (1975). The standard problem: Meaning and values in measurement and evaluation.American Psychologist, 30,955–966. Google Scholar Messick, S. (1980). Test validity and the ethics of assessment.American Psychologist, 35,1012–1027. Google Scholar Miller, J., & Chapman, R. (1981). The relation between age and mean length of utterance in morphemes.Journal of Speech and Hearing Research, 24, 154–161. LinkGoogle Scholar Muma, J. (1978). Language handbook.: Englewood Cliffs, NJ: Prentice-Hall. Google Scholar Muma, J. (1981). Language primer.: Lubbock, TX: Natural Child. Google Scholar Muma, J. (1983). Speech-language pathology: Emerging clinical expertise in language.In T Gallagher & C Prutting(Eds.), Pragmatic assessment and intervention issues in language (pp. 195–214). San Diego, CA: College-Hill Press. Google Scholar Muma, J. (1986). Language acquisition: A functionalistic perspective.: Austin, TX: PRO-ED. Google Scholar Muma, J. (1991). Experiential realism: Clinical implications.In T Gallagher (Ed.), Pragmatics of language (pp. 229–247). San Diego, CA: Singular. Google Scholar Muma, J. (1993). The need for replication.Journal of Speech and Hearing Research, 36, 927–930. LinkGoogle Scholar Muma, J, (1998). Effective speech-language pathology: A cognitive socialization approach.: Mahwah, NJ: Erlbaum. Google Scholar Muma, J. (2003). Construct validity: The essence of assessment.American Speech-Language-Hearing Association Convention, short course. Google Scholar MumaJ. (2004). Clarity needed on philosophical views.The ASHA Leader, March 2. Google Scholar Muma, J.,Morales, A.,Day, K.,Tackett, A.,Smith, S.Daniel, B.,Logue, B., & Morriss, D. (1998). Language sampling: Grammatical assessment.In J. Muma, Effective speech-language pathology: A cognitive socialization approach,: Appendix D. Mahwah, NJ: Erlbaum. Google Scholar Ninio, A.,Snow, C.,Pan, B., & Rollins, P. (1994). Classifying communicative acts in children's interactions.Journal of Communicative Disorders,, 27, 157–187. Google Scholar Priestly, T. (1980). Homonymy in child phonology.Journal of Child Language, 7, 413–427. CrossrefMedlineGoogle Scholar Rees, N., & Snope, J. (1983). National conference on undergraduate, graduate, and continuing education.Asha, 25, 49–59. Google Scholar Rogoff, B. (1990). Apprenticeship in thinking: Cognitive development in social context.: New York: Oxford University Press. Google Scholar Schwartz, R., & Leonard, L. (1982). Do children pick and choose? An examination of phonological selection and avoidance in early lexical acquisition.Journal of Child Language, 9, 319–336. CrossrefGoogle Scholar Searle, J. (1992). The rediscovery of the mind.: Cambridge, MA: MIT Press. Google Scholar Slobin, D. (1970). Suggested universals in the ontogenesis of grammar (Working Paper No. 32). Stanford University: Language Behavior Research Laboratory. Google Scholar Sperber, D., & Wilson, D. (1986). Relevance: Communication and cognition.: Cambridge, MA: Harvard University Press. Google Scholar Vihman, M., & Greenlee, M. (1987). Individual differences in phonological development: Ages one and three years.Journal of Speech and Hearing Research, 30,503–521. LinkGoogle Scholar Winer, B. (1962). Statistical principles in research design.: New York: McGraw-Hill. Google Scholar Additional Resources FiguresReferencesRelatedDetails Volume 5Issue 4December 2004Pages: 16-20 Get Permissions Add to your Mendeley library History Published in issue: Dec 1, 2004 Metrics Downloaded 37 times Topicsasha-topicsasha-article-typesasha-sigsCopyright & PermissionsCopyright © 2004 American Speech-Language-Hearing AssociationPDF DownloadLoading...
This study reports two experiments assessing the spelling performance of French first graders after 3 months and after 9 months of literacy instruction. The participants were asked to spell high and low frequency irregular words (Experiment 1) and pseudowords, some of which had lexical neighbours (Experiment 2). The lexical database which children had been exposed to was strictly controlled. Both a frequency effect in word spelling accuracy and an analogy effect in pseudoword spelling were obtained after only 3 months of reading instruction. The results suggest that children establish specific orthographic knowledge from the very beginning of literacy acquisition. (PsycINFO Database Record (c) 2016 APA, all rights reserved)
There is growing interest in the use of semantic collections in order to identify and analyse domain knowledge. This paper describes some technical issues to consider when contemplating research which incorporates small-to-medium domain-specific word sets. The purpose of the corpus construction described was to provide an external word collection which could be transformed to a numeric frequency scale which could take the place of an “expert ” in order to evaluate the lexical content of aircraft Visual Landing Approach concept maps. Although this paper is based on research in the field of aviation education, the underlying principles are more widely applicable. General Corpora Definitions and Uses The study of naturally occurring word frequencies has been a focus of computational linguistics, and in particular of the field of corpus linguistics. A word collection known as a corpus is constructed from some set of texts in order to determine what is characteristic of that text set through the identification of vocabulary patterns that either differ or conform to a norm (Ide & Walker, 1993). In contrast to Chomskyan generative linguistics, which focuses on internal knowledge of language structures (Chomsky, 1957), empirical corpus linguistics seeks to describe
Recent evaluation techniques applied to corpus-based systems have been introduced that can predict quantitatively how well surface realizers will generate unseen sentences in isolation. We introduce a similar method for determining the coverage on the Fuf/Surge symbolic surface realizer, report that its coverage and accuracy on the Penn TreeBank is higher than that of a similar statistics-based generator, describe several benefits that can be used in other areas of computational linguistics, and present an updated version of Surge for use in the NLG community.
Introduction The CorpusEye project (http://corp.hum.sdu.dk ) at the University of Denmark aims at designing and programming an internet based corpus search interface that (1) offers standardised search tools and a unified descriptive formalism across different corpus types and different languages, and (2) allows users to exploit grammatical information in annotated corpora in a user-friendly and menubased way. All corpora in CorpusEye have been annotated with VISL's Constraint Grammar based parsers, in the case of treebanks using an additional PSG module or equivalent (Bick 2003). At the time of writing, the material covers 8 languages and ca. 600 million words.
There are three types of sentences that form all existing natural languages: verbal sentences (e.g."I read the book."),copulative sentences (e.g."The book is on the table."),and existential sentences (e.g."There is a book on the table.").Syntactic and semantic recognition of these sentence types are crucially important in computational linguistics although there has not been any significant work towards this end.This thesis, in an attempt to fill this evident gap, is on identifying and assigning semantic categories of Turkish existential sentences in print.Existential sentences in Turkish are minimally characterized by the two existential particles var, meaning there is/are, and yok, meaning there is/are no.In addition to these most basic meanings, other senses of existential particles are possible, which can be categorized into groups such as case existentials and possession existentials.Our system does shallow semantic parsing in defining the predicate-argument relationships in an existential sentence on a word-byword basis, via utilizing Support Vector Machines, after which it proceeds with the semantic categorization of the whole sentence.For both of these tasks, our system produces promising results, in terms of accuracy and precision/recall, respectively.Part of this research contributes to the annotation of the METU-Sabanc Turkish Treebank with semantic information.
This paper takes a look at how information competency, information processing, and receiver apprehension, affect rating behavior when rating speeches. The paper is broken down into two studies. The first study looks at information competency, which is measured by the Information Competency Assessment Instrument, and how it affects rating behavior. The second study looks at information processing, which is broken down into three scales based on the Inventory of Learning Processes then each is analyzed based on rating behavior, and receiver apprehension which is measured using the Receiver Apprehension Test and how these two concepts affect rating behavior. Students were given surveys to measure information competency, information processing, and receiver apprehension and also asked to evaluate speeches. A negative correlation between receiver apprehension and information processing was found along with support for the idea that information competency and information processing have an affect on rating behavior.
The preceding articles by Piek Vossen (PV), Willy Martin (WM), and Marc van Campenoudt (MC) give a much more detailed account on their respective multilingual lexical database designs than the article by myself in this same journal (MJ). At the same time, they indicate some points of concern regarding the set-up of the SIM<it>u</it>LLDA system. Rather than responding directly to the points raised, this response elaborates on the two aspects of the SIM<it>u</it>LLDA system that seem to form the main sources of these issues: the status of the definitional attributes, and the practical usability of the system. The issues raised in the preceding articles will be explicitly addresses in the course of this elaboration. For even more details on these topics, see Janssen (2002).
The annotation of the Prague Dependency Treebank (PDT) is conceived of as a multilayered scenario that comprises also dependency representations (tectogrammatical tree structures, TGTS’s) of the underlying structure of the sentences. TGTS’s capture three basic aspects of the underlying structure of sentences: (a) the dependency tree structure, (b) the kinds of dependency syntactic relations, and (c) the basic characteristics of the topic-focus articulation (TFA). Since the PDT is a large collection and the annotations on the deepest layer are to a large extent performed by several human annotators (based on an automatic preprocessing module), it is more than necessary to observe the consistence of annotators and the agreement among them. In the present paper, we summarize the results of the evaluation of parallel annotations of several samples taken from PDT and the measures accepted to improve the consistency of annotations.
To assign semantic roles in building Treebanks, there is a need for annotators having a guideline in determining semantic relations between phrasal head and its modifiers or arguments. Semantic roles are hard to have clear-cut definitions. It is not always easy to determine thematic relations between two concepts. This paper aims to introduce an integrated nominal modifier system. Basically we adopt other scholars &apos; incisive idea in analyzing semantic roles that modify general nouns. We use the approach of building a fine-grain taxonomy of role system. The taxonomy of fine-grain thematic roles makes the role determination easier for human annotators, since the meaning of a fine-grain semantic role is self explanatory and a higher-level semantic role is described by its hyponyms. The proposed taxonomy has been attested during construction of Sinica TreeBank and HowNet definitions of nominal concepts and proven to be more applicable than conventional flat structures. 1
The subject matter of this article concerns objectivity of lexicographic description. Here objectivity is being associated with the names of “functional” places, in D. Le Pesant’s understanding of the term. The author presents principles of the object-oriented description of lexicosemantic data, according to the concept by W. Banyś. Later on, she distinguishes a couple of basic classes of appropries predicates, which are typical for the locative nouns suggested by D. Le Pesant. The elements of the object-oriented approach are confronted with their counterparts in such lexicographic theories as: I. Melczuk and A.K. Zholkovsky’s frames, J. Pustejovsky’s qualia structure system, G. Gross’s object classes and unique beginners from the WordNet lexical database. In the suggested descriptive scheme of entries the author places results of her own analyses concerning descriptions of names of functional places, in this case-buildings.
This report presents an approach to enriching flat and robust predicate argument structures with more fine-grained semantic information, extracted from underspecified semantic representations and encoded in Minimal Recursion Semantics (MRS). Such representations are provided by a hand-built HPSG grammar with a wide linguistic coverage. A specific semantic representation, called linked predicate argument structure (LPAS), has been worked out, which describes the explicit embedding relationships among predicate argument structures. LPAS can be used as a generic interface language for integrating semantic representations with different granularities. Some initial experiments have been conducted to convert MRS expressions into LPASs. A simple constraint solver is developed to resolve the underspecified dominance relations between the predicates and their arguments in MRS expressions. LPASs are useful for high-precision information extraction and question answering tasks because of their fine-grained semantic structures. In addition, I have attempted to extend the lexicon of the HPSG English Resource Grammar (ERG) exploiting WordNet and to disambiguate the readings of HPSG parsing with the help of a probabilistic parser, in order to process texts from application domains. Following the presented approach, the HPSG ERG grammar can be used for annotating some standard treebank, e.g., the Penn Treebank, with its fine-grained semantics. In this vein, I point out opportunities for a fruitful cooperation of the HPSG annotated Redwood Treebank and the Penn PropBank. In my current work, I exploit HPSG as an additional knowledge resource for the automatic learning of LPASs from dependency structures.
This paper presents an editor for compiling a multilingual machine readable lexicon, like the Simple-lexicon. This editor has proven to be a useful tool in linking several languages in one lexical database, and to edit the entries in a consistent and convenient way. The editor has been designed for linking Danish, Swedish and Norwegian in the Simple Scan-project, but might easily be extended to include all the languages in the Simple project. The editor may also be modified for similar machine readable lexical projects. 1.
In this study, a technique called semantic self-organization is used to scale up the subsymbolic approach by allowing a network to optimally allocate frame representations from a semantic dependency graph. The resulting architecture, INSOMNet, was trained on semantic representations of the newly-released LinGO Redwoods HPSG Treebank of anno-tated sentences from the VerbMobil project. The results show that INSOMNet is able to accurately represent the semantic dependencies while demonstrating expectations and defaults, coactivation of multiple interpretations, and robust process-ing of noisy input. The cognitive plausibility of the model is underscored by the collective modelling of four experiments from the visual worlds paradigm to show the model’s ability to adapt to context.
We present a method to approximate a LTAG grammar by a CFG. A key process in the approximation method is finite enumeration of partial parse results that can be generated during parsing. We applied our method to the XTAG English grammar and LTAG grammars which are extracted from the Penn Treebank, and investigated characteristics of the obtained CFGs. We perform CFG filtering for LTAG by the obtained CFG. In the experiments, we describe that the obtained CFG is useful for CFG filtering for LTAG parser. 1
The OmniPaper project has implemented several information retrieval prototypes in the area of electronic news publishing. One prototype uses SOAP as communication protocol between the central system and a number of distributed news archives. The second prototype uses an RDF metadata database, enabling direct metadata queries to the central system. Finally the Topic Map prototype uses query expansion and semantic linking for smart metadata search. The Topic Map prototype enhances the search experience by implementing a knowledge layer that combines the semantic content of a lexical database, consisting of concepts and keywords, with a metadata-set of newspaper articles. After developing and testing three smaller prototypes, the OmniPaper consortium has combined these prototypes in one. In this final prototype a kind of “enhanced full-text search” engine is implemented. This means that the prototype is an interface on top of existing search engines. When a user submits a query, this query is forwarded to several distributed news archives to retrieve relevant news articles. Next to this, the system: 1) translates queries to enable multilingual search, 2) provides a query refinement mechanism, both in graphic and text-based form, allowing users to adapt their query and 3) provides uniform result ranking algorithm across the different news archives. In this prototype querying and navigation are considered as alternative methods to find relevant information. Both interact with each other and together they produce a combined user experience that can be expressed as find what you were looking for and then browse away from it. In fact, the prototype considers both querying and navigation as a kind of search action and tries to integrate both. In concrete, keywords in a query are looked up in a dictionary and shown to the user. In the background, the keywords are translated and expanded to related terms. These expanded queries are sent to the underlying full-text search engine(s) in all requested languages. In the graphical tool (“web of concepts”) users can redefine the meaning of their query words, resulting in an updated query and result set. Both disambiguation (choosing one meaning of a word out of many) and refinement (browsing to related words) are possible. Figure 1 shows the web of concepts for the query “poll Indonesia”. The word “Indonesia” is recognized in only one concept, “Dutch East Indies”, whereas the word “poll” has many different meanings. If you select the meaning “canvass” for example, this word is replacing the word “poll” in the original query. After selection the concept “canvass” can again be expanded to related concepts, be it more general or more specific in meaning. In the textual tool only refinement is possible. The user gets a list of words that are related to the words appearing in the query, grouped into more similar, more specific and more general terms. Then the user can change his/her query using these proposed words.
The claim made in this paper is that in a formal description of language, it is possible and useful to work with dependency-based underlying representations of sentences (tectogrammatical representations) meeting the condition of projectivity. The reasons for the inclusion of this condition into the definition of the tectogrammatical representations are both formally and empirically sound (Section 1). An analysis of the material offered by the Prague Dependency Treebank with annotations of the underlying syntactic structure of sentences (described in Section 2) has led to an interesting classification of non-projective constructions in Czech (Section 3). It documents that most (types of) constructions that appear to be non-projective in the surface shape of sentences can be described by means of projective trees. The realization of the surface word order (with the use of movement rules) is then relegated to the morphemic level, where the representation of the sentence has the shape of a string rather than a tree.
We discuss existing approaches to train LR parsers, which have been used for statistical resolution of structural ambiguity. These approaches are nonoptimal, in the sense that a collection of probability distributions cannot be obtained. In particular, some probability distributions expressible in terms of a context-free grammar cannot be expressed in terms of the LR parser constructed from that grammar, under the restrictions of the existing approaches to training of LR parsers. We present an alternative way of training that is provably optimal, and that allows all probability distributions expressible in the context-free grammar to be carried over to the LR parser. We also demonstrate empirically that this kind of training can be effectively applied on a large treebank.
We present an algorithm for generating referring expressions in open domains. Existing algorithms work at the semantic level and assume the availability of a classification for attributes, which is only feasible for restricted domains. Our alternative works at the realisation level, relies on Word-Net synonym and antonym sets, and gives equivalent results on the examples cited in the literature and improved results for examples that prior approaches cannot handle. We believe that ours is also the first algorithm that allows for the incremental incorporation of relations. We present a novel corpus-evaluation using referring expressions from the Penn Wall Street Journal Treebank.
The epidemic of mesothelioma in Cappadocia, Turkey, is unprecedented in medical history. In three Cappadocian villages, Karain, Tuzkoy and "old" Sarihidir, about 50% of all deaths (including neonatal deaths and traffic fatalities) have been caused by mesothelioma. No other epidemic in medical history has caused such a high incidence of death. This is even more unusual when considering that (I) epidemics are caused by infectious agents, not cancer, and (II) mesothelioma is a rare cancer. World-wide mesothelioma incidence varies between 1/10<sup>6</sup> in areas with no asbestos industry to about 10-30/10<sup>6</sup> in areas with asbestos industry. This article reviews how the mesothelioma epidemic was discovered in Cappadocia by Dr. Baris (my mentor), how we initially linked the epidemic to erionite exposure, and later (with Dr. Carbone) to the interaction between genetic predisposition and environmental exposure. Our team's work had an important positive impact on the lives of those living in Cappadocia and also in many genetically predisposed families living around the world. I will discuss how the work that started in three remote Cappadocian villages led to the award of a NCI P01 grant to support our studies. Our studies proved that genetics modulates mineral fiber carcinogenesis and led to the discovery that carriers of germline <i>BAP1</i> mutations have a very high risk of developing mesothelioma and other malignancies. A new, very active field of research developed following our discoveries to elucidate the mechanism by which <i>BAP1</i> modulates mineral fiber carcinogenesis as well as to identify additional genes that when mutated increase the risk of mesothelioma and other environmentally related cancers. I am the only surviving member of this research team who saw all the phases of this research and I believe it is important to provide an accurate report, which hopefully will inspire others.
This paper shows how finite approximations of long distance dependency (LDD) resolution can be obtained automatically for wide-coverage, robust, probabilistic Lexical-Functional Grammar (LFG) resources acquired from treebanks. We extract LFG subcategorisation frames and paths linking LDD reentrancies from f-structures generated automatically for the Penn-II treebank trees and use them in an LDD resolution algorithm to parse new text. Unlike (Collins, 1999; Johnson, 2000), in our approach resolution of LDDs is done at f-structure (attribute-value structure representations of basic predicate-argument or dependency structure) without empty productions, traces and coindexation in CFG parse trees. Currently our best automatically induced grammars achieve 80.97% f-score for f-structures parsing section 23 of the WSJ part of the Penn-II treebank and evaluating against the DCU 1051 and 80.24% against the PARC 700 Dependency Bank (King et al., 2003), performing at the same or a slightly better level than state-of-the-art hand-crafted grammars (Kaplan et al., 2004).
We present a new methodology for the semiautomated maintenance of a treebank built from analyses of a computational grammar and gauge the effort required for each update cycle. Based on a decade of large-scale grammar engineering experience, we propose a tight integration of treebank maintenance with the continuous evolution of a ‘deep’ computational grammar.
Abstract. The use of lexicons and corpora advances both linguistic re-search and performance of current natural language processing (NLP) systems. We present a tool that exploits such resources, specifically En-glish and German lexical databases and the World Wide Web to recognise English inclusions in German newspaper articles. The output of the tool can assist lexical resource developers in monitoring changing patterns of English inclusion usage. The corpus used for the classification covers three different domains. We report the classification results and illustrate their value to linguistic and NLP research. 1
The requirements of the depth and precision of annotation vary for different intended uses of the corpus but it has been commonly accepted nowadays that the standard annotations of surface structure are only the first steps in a more ambitious research program, aiming at a creation of advanced resources for most different systems of natural language processing and for testing and further enrichment of linguistic and computational theories. Among the several possible directions in which we believe the standard annotation systems should go (and in some cases already attempt to go) beyond the POS tagging or shallow syntactic annotations, the following four are characterized in the present contribution: (i) predicateargument representation of the underlying syntactic relations as basically corresponding to a rooted tree that can be univocally linearized, (ii) the inclusion of the information structure using very simple means (the left-to-right order of the nodes and three attribute values), (iii) relating this underlying structure (rendering the ”linguistic meaning,” i.e. the semantically relevant counterparts of the grammatical means of expression) to certain central aspects of referential semantics (reference assignment and coreferential relations), and (iv) handling of word sense disambiguation. The first three issues are documented in the present paper on the basis of our experience with the development of the structure and scenario of the Prague Dependency Treebank which provides for syntactico-semantic annotation of large text segments from the Czech National Corpus and which is based on a solid theoretical framework.
Parsing, the task of identifying syntactic components, e.g., noun and verb phrases, in a sentence, is one of the fundamental tasks in natural language processing. Many natural language applications such as spoken-language understanding, machine translation, and information extraction, would benefit from, or even require, high accuracy parsing as a preprocessing step. Even though most state-of-the-art statistical parsers were initially constructed for parsing in English, most of them are not language-specific, in that they do not rely on properties of the language that are specific to English. Therefore, construction of a parser in a given language becomes a matter of retraining the statistical parameters with a Treebank in the corresponding language. The development of the Chinese treebank [Xia et al. 2000] spurred the construction of parsers for Chinese. However, Chinese as a language poses some unique problems for the development of a statistical parser, the most apparent being word segmentation. Since words in written Chinese are not delimited in the same way as in Western languages, the first problem that needs to be solved before an existing statistical method can be applied to Chinese is to identify the word boundaries. This is a step that is neglected by most pre-existing Chinese parsers, which assume that the input data has already been pre-segmented. This article describes a character-based statistical parser, which gives the best performance to-date on the Chinese treebank data. We augment an existing maximum entropy parser with transformation-based learning, creating a parser that can operate at the character level. We present experiments that show that our parser achieves results that are close to those achievable under perfect word segmentation conditions.
The sense of smell has been traditionally assumed to be different from other sensory modalities in that odors are encoded perceptually, without a semantic component. Recent findings of improved odor memory upon encoding with verbal cues question this view. Furthermore, familiar odors are easier to remember and discriminate than are unfamiliar ones, and odor familiarity is reported to predict odor naming. To investigate whether familiar odors are processed by different cerebral structures than those that process unfamiliar odors, (15)O H(2)O-positron emission tomography (PET) measurements of cerebral blood flow were carried out in 14 healthy men. The task was passive, birhinal, smelling of familiar odors (FAM), unfamiliar odors (uFAM), and odorless air (AIR). Significant activations (P < 0.05) were calculated using the contrasts FAM-AIR, uFAM-AIR, and FAM-uFAM, and deactivations running these contrasts in the opposite direction. In relation to AIR, both FAM and uFAM activated amygdala, piriform cortex, and parts of anterior cingulate cortex. FAM activated, in addition, left frontal cortex (Brodmann's areas 44,45,47), left parietal cortex incorporating precuneus, and right parahippocampus. Clusters covering parahippocampus and precuneus were observed also in FAM-uFAM. The activation of left frontal cortex and right parahippocampus was positively correlated with familiarity ratings. Smelling of familiar but not unfamiliar odorants seems to engage cerebral circuits mediating memory and language functions, in addition to the engagement of olfactory cortex. Already the most elemental form of odor processing, passive perception thus seems to engage semantic circuits. This is achieved by the ability of odorants to immediately elicit associations and judgments of odor characteristics.
We present an automatic semantic roles labeling system for structured trees of Chinese sentences. It adopts dependency decision making and example-based approaches. The training data and extracted examples are from the Sinica Treebank, which is a Chinese Treebank with semantic role semantic roles including thematic roles, such as ‘agent’; ‘theme’, ‘instrument’, and secondary roles of ‘location’, ‘time’, ‘manner ’ and roles for nominal modifiers. The design of role assignment algorithm is based on the different decision features, such as head-argument/modifier, case makers, sentence structures etc. It labels semantic roles of parsed sentences. Therefore the practical performance of the system depends on a good parser which labels the right structures of sentences. The system achieves 92.71 % accuracy in labeling the semantic roles for pre-structure- bracketed texts which is considerably higher than the simple method using probabilistic model of head-modifier relations. 1.
We introduce a simple method to build Lexicalized Hidden Markov Models (L-HMMs) for improving the precision of part-of-speech tagging. This technique enriches the contextual Language Model taking into account a set of selected words empirically obtained. The evaluation was conducted with different lexicalization criteria on the Penn Treebank corpus using the TnT tagger. This lexicalization obtained about a 6% reduction of the tagging error, on an unseen data test, without reducing the efficiency of the system. We have also studied how the use of linguistic resources, such as dictionaries and morphological analyzers, improves the tagging performance. Furthermore, we have conducted an exhaustive experimental comparison that shows that Lexicalized HMMs yield results which are better than or similar to other state-of-the-art part-of-speech tagging approaches. Finally, we have applied Lexicalized HMMs to the Spanish corpus LexEsp.
Abstract. This paper describes experiments on using inductive machine learning to guide a deterinistic dependency parser for unrestricted natural language text. Using data from a small treebank of Swedish, an eager probabilistic learning algorithm is used to induce context-sensitive parse tables. Evaluation shows a significant improvement over the baseline, which uses a table without contextual information. 1
Positive and negative affects may bias behavior toward approach to rewards and withdrawal from threat, particularly when the contingencies are ambiguous. The hypothesis was that positive and negative affects would associate predictably with identification of happy, disgusted, or angry expressions that may signal potentially rewarding or aversive social interactions. Healthy volunteers (n=86) completed affect ratings and a facial emotion task that employed morphed continua in which emotional expressions gradually decreased in ambiguity. Relations between mood and intensity thresholds for emotion identification were computed. Anhedonia (low positive affect) predicted thresholds for happy expressions (r=0.24; P=.026) whereas negative affect predicted thresholds for disgust (r=-0.25; P=.022). Even within a normal range of mood, mood predicted emotion identification, supporting constructs of positive and negative affect derived originally from self-report measures.
This paper presents an overview of a project to acquire wide-coverage, probabilistic Lexical-Functional Grammar (LFG) resources from treebanks. Our approach is based on an automatic annotation algorithm that annotates “raw” treebank trees with LFG f-structure information approximating to basic predicate-argument/dependency structure. From the f-structure-annotated treebank we extract probabilistic unification grammar resources. We present the annotation algorithm, the extraction of lexical information and the acquisition of wide-coverage and robust PCFGbased LFG approximations including long-distance dependency resolution. We show how the methodology can be applied to multilingual, treebank-based unification grammar acquisition. Finally we show how simple (quasi-)logical forms can be derived automatically from the f-structures generated for the treebank trees.
Abstract We present a computational language learning algorithm that can induce massively probabilistic grammars from treebank data. The paper is based on chapter 6 in my forthcoming thesis “Discontinuous Grammar: A dependency-based model of human parsing and language learning.” We start in section 1 by introducing the prerequisites from probability theory and statistics that are needed in the rest of the paper. In section 2, we define weighted grammars and massively probabilistic grammars, and discuss their relationship to standardly used grammars. In section 3, we address the problem of estimating probability distributions for hierarchically structured categorical data, and present an estimation algorithm that selects a hierarchical partition model by means of local search. We also describe a simulation study that shows that the algorithm performs well unless the distribution that generated the data is highly symmetric. Finally, in section 4, we outline how the algorithm can be used to learn massively probabilistic dependency grammars, exemplified by the task of learning the probabilities of complement structures.
Automatic extraction and reasoning over temporal properties in natural language discourse has not had wide use in practical systems due to its demand for a rich and compositional, yet inference-friendly, representation of time. Motivated by our study of temporal expressions from the Penn Treebank corpora, we address the problem by proposing a two-level constraint-based framework for processing and reasoning over temporal information in natural language. Within this framework, temporal expressions are viewed as partial assignments to the variables of an underlying calendar constraint system, and multiple expressions together describe a temporal constraint-satisfaction problem (TCSP). To support this framework, we designed a typed formal language for encoding natural language expressions. The language can cope with phenomena such as under-specification and granularity change. The constraint problems can be solved using various constraint propagation and search methods, and the solutions can then be used to answer a wide range of time-related queries.
We describe the automatic conversion of English Penn Treebank (PTB) annotations into Language Neutral Syntax (LNS) (Campbell and Suzuki, 2002a,b). In this paper, we describe LNS and why it is useful, describe the conversion algorithm, present an evaluation of the conversion, and discuss some uses of the converted annotations and the potential for extending the coverage to other languages. The work described here is in the spirit of other automatic re-annotations of PTB trees (e.g. Frank, 2000 and Meyers, 2001), but differs in the nature of the output.
Hate crime laws are a highly controversial legal approach in society's response to intergroup violence. Argument acceptance, knowledge, and individual differences were examined in relationship to attitudes about these laws. These variables were also considered in terms of efforts to influence a peer's beliefs about hate crime laws. One‐hundred and sixty‐seven participants completed a measure of knowledge of human rights laws, Gough's Pr scale, the Selznick and Steinberg anti‐Semitism scale, and Cuellar's Machismo scale. Hate crime attitudes were measured on an affect rating scale and six statements reflecting arguments favoring and opposing hate crime laws. Peer influence was examined on Interpersonal Power Inventory (IPI). Results showed that while most participants endorsed positive attitudes about hate crime laws, men—and both women and men who endorsed machismo attitudes—were more likely to agree with media distortion and identity politics arguments opposing hate crime laws. The Pr and machismo scales predicted greater effort on the IPI to influence peer attitudes about hate crime laws, after controlling for demographic differences of the participants. These findings indicate that more explicitly biased individuals were more effortful in trying to change the attitudes of peers concerning the legitimacy of hate crime laws.
We present the first application of the head-driven statistical parsing model of Collins (1999) as a simultaneous language model and parser for large-vocabulary speech recognition. The model is adapted to an online left to right chart-parser for word lattices, integrating acoustic, n-gram, and parser probabilities. The parser uses structural and lexical dependencies not considered by n-gram models, conditioning recognition on more linguistically-grounded relationships. Experiments on the Wall Street Journal treebank and lattice corpora show word error rates competitive with the standard n-gram language model while extracting additional structural information useful for speech understanding.