Papers reviewed and determined not to be word norm studies. Use the flag icon to report errors or suggest re-inclusion.
18265 papers
The aim of the paper is to show that a subset of Text Encoding Initiative Guidelines is a reasonable choice as a standard for stand-off XML encoding of syntactically annotated corpora. The proposed TEI schema — actually employed in the National Corpus of Polish — is compared to other such candidate standards, including TIGER-XML, SynAF and PAULA. 1
Background: The variety of ways in which faces are categorized makes face recognition challenging for both synthetic and biological vision systems. Here we focus on two face processing tasks, detection and individuation, and explore whether differences in task demands lead to differences both in the features most effective for automatic recognition and in the featural codes recruited by neural processing. Methodology/Principal Findings: Our study appeals to a computational framework characterizing the features representing object categories as sets of overlapping image fragments. Within this framework, we assess the extent to which task-relevant information differs across image fragments. Based on objective differences we find among task-specific representations, we test the sensitivity of the human visual system to these different face descriptions independently of one another. Both behavior and functional magnetic resonance imaging reveal effects elicited by objective task-specific )
The Portal da Lingua Portuguesa is a website containing information about the Portuguese language oriented towards the general public. The largest part of the information on the Portal is lexical information concerning formal characteristics of words, such as orthography, derivations, loanwords and gentiles. The lexical information comes from a lexical database called MorDebe-or more precisely, a network of lexical databases called the Open Source Lexical Information Network (OSLIN). This abstract shows the general set-up of and major functions of MorDebe Admin, which is the lexicon management system for OSLIN. MorDebe Admin provides an easy and secure way of updating and editing the content of the different databases of OSLIN. Furthermore, much of the data on the Portal are organised as mini-dictionaries and MorDebe Admin provides an integrated collection of tools dedicated to the maintenance of these mini-dictionaries, as well as a built-in neologism tracking system. The software demonstration will illustrate these functions from a user perspective, and how easy ir is to maintain the data behind the Portal.
This paper deals with a multilingual relational lexical database of proper name, Prolexbase, a free resource available on the CNRTL website. The Prolex model is based on two main concepts: firstly, a language independent pivot and, secondly, the prolexeme (the projection of the pivot onto particular language), that is a set of lemmas (names and derivatives). These two concepts model the variations of proper name: firstly, independent of language and, secondly, language dependent by morphology or knowledge. Variation processing is very important for NLP: the same proper name can be written in different instances, maybe in different parts of speech, and it can also be replaced by another one, a lexical anaphora (that reveals semantic link). The pivot represents different referent's points of view, i.e. language independent variations of name. Pivots are linked by three semantic relations (quasi-synonymy, partitive relation and associative relation). The prolexeme is a set of variants (aliases), quasi-synonyms and morphosemantic derivatives. Prolexemes are linked to classifying contexts and reliability code.
Spoken language resources (SLRs) are essential for both research and application development. In this article we clarify the concept of SLR validation. We define validation and how it differs from evaluation. Further, relevant principles of SLR validation are outlined. We argue that the best way to validate SLRs is to implement validation throughout SLR production and have it carried out by an external and experienced institute. We address which tasks should be carried out by the validation institute, and which not. Further, we list the basic issues that validation criteria for SLR should address. A standard validation protocol is shown, illustrating how validation can prove its value throughout the production phase in terms of pre-validation, full validation and pre-release validation.
We survey the evaluation methodology adopted in information extraction (IE), as defined in a few different efforts applying machine learning (ML) to IE. We identify a number of critical issues that hamper comparison of the results obtained by different researchers. Some of these issues are common to other NLP-related tasks: e.g., the difficulty of exactly identifying the effects on performance of the data (sample selection and sample size), of the domain theory (features selected), and of algorithm parameter settings. Some issues are specific to IE: how leniently to assess inexact identification of filler boundaries, the possibility of multiple fillers for a slot, and how the counting is performed. We argue that, when specifying an IE task, these issues should be explicitly addressed, and a number of methodological characteristics should be clearly defined. To empirically verify the practical impact of the issues mentioned above, we perform a survey of the results of different algorithms when applied to a few standard datasets. The survey shows a serious lack of consensus on these issues, which makes it difficult to draw firm conclusions on a comparative evaluation of the algorithms. Our aim is to elaborate a clear and detailed experimental methodology and propose it to the IE community. Widespread agreement on this proposal should lead to future IE comparative evaluations that are fair and reliable. To demonstrate the way the methodology is to be applied we have organized and run a comparative evaluation of ML-based IE systems (the Pascal Challenge on ML-based IE) where the principles described in this article are put into practice. In this article we describe the proposed methodology and its motivations. The Pascal evaluation is then described and its results presented.
With the advent of the Internet, billions of images are now freely available online and constitute a dense sampling of the visual world. Using a variety of non-parametric methods, we explore this world with the aid of a large dataset of 79,302,017 images collected from the Internet. Motivated by psychophysical results showing the remarkable tolerance of the human visual system to degradations in image resolution, the images in the dataset are stored as 32 x 32 color images. Each image is loosely labeled with one of the 75,062 non-abstract nouns in English, as listed in the Wordnet lexical database. Hence the image database gives a comprehensive coverage of all object categories and scenes. The semantic information from Wordnet can be used in conjunction with nearest-neighbor methods to perform object classification over a range of semantic levels minimizing the effects of labeling noise. For certain classes that are particularly prevalent in the dataset, such as people, we are able to demonstrate a recognition performance comparable to class-specific Viola-Jones style detectors.
Treebanks, or syntactically annotated corpora, are an invaluable resource for the development and evaluation of syntactic parsers, as well as for empirical research on natural language syntax. Treebanks for Swedish have a long and venerable history, represented by the pioneering work on Talbanken (Einarsson,
Reviewed by: Lexicalization and language change Jesús Fernández-Domínguez Laurel J. BrintonElizabeth Closs Traugott. 2005. Lexicalization and language change. In the series Research Surveys in Linguistics. Cambridge: Cambridge University Press. Pp. xii + 207. US $34.99 (softcover). Lexicalization has been customarily defined as “a gradual historical process, involving graphemic, phonological and semantic changes and the loss of motivation” (Lipka 2005:40), and can affect a word in its phonology, morphology, semantics, or syntax. Because it can affect the makeup of virtually any item, it stands as a central phenomenon in language change, and as such it has gathered the attention of scholars for decades. The aim of [End Page 104] Brinton and Traugott’s work is to provide a wide coverage for what has been traditionally considered under lexicalization, as well as to discuss related concepts necessary for its understanding. Lexicalization and language change develops along six chapters and progressively introduces the various conceptualizations given to the processes of language modification. Chapter 1 (pp. 1–31) sets the theoretical context of the book and introduces some basic notions, and Chapter 2 (pp. 32–61) provides a background in terms of definitions and viewpoints for lexicalization. The authors discuss next the relationship between lexicalization and grammaticalization, first in a general fashion in Chapter 3 (pp. 62–88) and then in further detail in Chapter 4 (pp. 89–110). The most relevant contents of the work are exemplified in Chapter 5 (pp. 111–140), and some conclusions and research questions are offered in Chapter 6 (pp. 141–160). Among the concepts introduced in Chapter 1, the notion of lexicon bears a special significance, as there exist various senses to it which must be clarified before attempting a definition of lexicalization (see Aronoff 1989, not mentioned by the authors). To this end, Brinton and Traugott devote several pages to outline holistic vs. componential approaches to the lexicon, to the categories of the lexicon, and to the lexicon viewed as a continuum of productivity, thus laying the conceptual background required for a proper comprehension of the book. This overview is a suitable introduction to the subject also because it is contrasted with concepts like grammar, language change, or productivity, all of which have a bearing on lexicalization and are seen by the authors as a matter of gradation. Brinton and Traugott also offer a summary of the remainder of contents, and set a number of assumptions for a study of language change “from a historical, functionalist perspective” (p. 31). Chapter 2 immerses into lexicalization proper. After a brief introduction, a central section is “Ordinary processes of word formation” (pp. 33–45), a summary of the major devices of contemporary English: compounding, derivation, conversion, back-formation, initialism, etc. Here, Brinton and Traugott rightly note that lexicalization is to be distinguished from word-formation as far as only the latter has the capacity to produce new items in a regular and predictable manner, a discussion picked up later in Chapter 4. Their review proves valuable because it offers the reader the general features of present-day word-formation in a concise and satisfactory manner, even if one can hardly agree with the inclusion of loan translation, root creation, or coinage under word-formation (see Štekauer 2005:214). A subsequent logical step is the indispensable though brief explanation of institutionalization, that is, “the spread of a usage to a community and its establishment as the norm” (p. 45), usually taken as a stage following word-formation and preceding lexicalization (see Bauer 1983:45–48; Hohenhaus 2005). A number of opinions are explained and illustrated here before turning to the core of the chapter: lexicalization as fusion (pp. 47–57) and as increase in autonomy (pp. 57–60). The authors complain of the very little attention that lexicalization as fusion has received from a historical point of view, and define it as “the development of a form from a more complex to a simpler sequence” (p. 47). The present chapter truly represents a deep and up-to-date review of the typology of the phenomenon, given that it covers lexicalization as affecting phrasal and syntactic constructions (p. 48–50), word-formation (p. 50–52), phonological...
This article examines the phenomenon of code switching in The Map of Love (1999) by the Egyptian—British writer Ahdaf Soueif. Though she chooses English as a medium for her creative expression, Soueif deploys Arabic in her narrative to represent different aspects of the linguistic and cultural norms of Egyptian society. The article's methodology is informed by Kachru's framework on contact literature and his categorization of the occurrence of literary code switching or bilingual creativity into different strategies that encompass cultural and linguistic processes. The results indicate the predominance in The Map of Love of the discourse strategies of employing lexical borrowing, culture-bound references and translational transfer. Finally, the article analyzes the functional motivation of code switching in the postcolonial context of the novel and how the use of certain creative strategies might enhance or diminish the narrative's effectiveness and readability.
The late positive potential (LPP) is a sustained positive deflection in the event-related potential that is larger following the presentation of emotional compared to neutral visual stimuli. Recent studies have indicated that the magnitude of the LPP is sensitive to emotion regulation strategies such as reappraisal, which involves generating an alternate interpretation of emotional stimuli so that they are less negative. It is unclear, however, whether reappraisal-related reductions in the LPP reflect reduced emotional processing or increased cognitive demands following reappraisal instructions. In the present study, we sought to examine whether a more or less negative description preceding the presentation of unpleasant images would similarly modulate the LPP. The LPP was recorded from 26 subjects as they viewed unpleasant and neutral International Affective Picture System images. All participants heard a brief description of the upcoming picture; prior to unpleasant images, this description was either more neutral or more negative. Following the more neutral description, the magnitude of the LPP, unpleasant ratings, and arousal ratings were all reliably reduced. These results indicate that changes in narrative are sufficient to modulate the electrocortical response to the initial viewing of emotional pictures, and are discussed in terms of recent studies on reappraisal and emotion regulation.
Objective: To carry out the native assessment of International Affective Picture System(IAPS) among Chinese older adults.Methods:Altogether 116 Chinese older adults,including 51 male and 65 female,from three communities in Dalian City,aged from 60 to 80 years,rated 60 pictures(positive:25,neutral:12,negative:23) selected from the IAPS in terms of valence,arousal and dominance with Self-Assessment Manikin(SAM).The mean affective ratings were compared to the normative ratings of USA National Institute of Mental Health(NIMH).Result: Reliability analysis indicated that the affective ratings of our sample were stable and highly internally consistent.The affective ratings of Chinese older participants were strongly correlated with the normative ratings of NIMH(r=0.92,0.54 and 0.88 respectively for valence,arousal and dominance,P0.001).But paired t test showed there were still significant differences between the two samples.Chinese aged reported relatively higher arousal and dominance than NIMH sample for all pictures [(5.33±0.93)vs.(4.83±1.25),(5.60±1.20)vs.(5.19±1.21),P0.001],but lower valence than NIMH sample[(4.99±2.28)vs.(5.28±1.85),P=0.020].Male and female Chinese older participants showed similar emotional responses to most pictures.But female Chinese older participants reported higher valence than male ones(5.05±2.33/4.93±2.24,P0.05).The 60 pictures were distributed as shape in the two-dimensional affective space(valence-arousal).The association between valence and arousal was pronounced and linear for positive pictures(r=0.71,P0.001),but unpronounced for negative pictures,(r=-0.35,P0.05).Conclusion: IAPS is highly internationally accessible just as the expectation of its designers.However,considering about great differences in many aspects such as culture,social living and age between Chinese aged and NIMH sample,they may have different affective experiences to the same emotional stimuli.Therefore it is necessary to do some revisal before the IAPS is applied to Chinese aged.
While large-scale corpora and various corpus query tools have long been recognized as essential language resources, the value of word association norms as language resources has been largely overlooked. This paper conducts some initial comparisons of the lexical relationships observed within Japanese collocation data extracted from a large corpus using the Japanese language version of the Sketch Engine (SkE) tool (Srdanović et al., 2008) and the relationships found within Japanese word association sets taken from the large-scale Japanese Word Association Database (JWAD) under ongoing construction by Joyce (2005, 2007). The comparison results indicate that while some relationships are common to both linguistic resources, many lexical relationships are only observed in one resource. These findings suggest that both resources are necessary in order to more adequately cover the diverse range of lexical relationships. Finally, the paper reflects briefly on the implementation of association-based word-search strategies into electronic dictionaries proposed by Zock and Bilac (2004) and Zock (2006).
In this paper, I argue that standard, co-descriptional glue semantics provides no clear and satisfactory role for the traditional PREDfeatures of LFG, due to the fact that the linear logic of glue semantics does the work of the Completeness and Coherence Constraints. But then I show that a reduced but significant role for PRED-features can be found in an alternative ‘Description-by-Analysis’ (DBA) formulation, proposed in Andrews (2007a). The DBA formulation is argued to be superior in various respects, and some constraints are proposed to cause the DBA approach to approximate some of the empirically justifiable aspects of the behavior of the co-descriptional formulation. The standard way to combine LFG with glue-semantics has been with a ‘co-descriptional’ architecture in which lexical entries introduce the usual grammatical features in the usual way, together with ‘meaning-constructors’ that account for the meanings, both of the PRED-feature associated with the lexical item, and any semantically intepretable grammatical features that it might introduce, either inherently or due to the inflectional morphology. Typical examples would be the following entries for the verb form went and the noun-form feet: (1) a. went:V, (↑PRED)= ‘Gomotion ’, (↑TENSE)=PAST, λx.go(x): (↑ SUBJ)e −◦ ↑p, λP.Past(P ): ↑p −◦ ↑p b. feet:N, (↑PRED)= ‘Foot’, (↑NUM)=PL, λx.Foot(x): ↑p, λP.Past(P ): ↑p −◦ ↑p Co-description was introduced and motivated in Halvorsen and Kaplan (1988) as an alternative to the earlier (and overall more often used) ‘description-byanalysis’ (DBA) architecture, in which the f-structure is the primary input to the semantics. Although the norm in glue-semantics, co-description raises a puzzle with respect to the role of PRED-features, namely, why they are there at all. The problem is that, as pointed out in Kuhn (2001), the linear logic resource management employed in glue is in itself sufficient to account for the phenomena of Completeness, Coherence, and Predicate Uniqueness, which comprise the major special properties of PRED-features. This leaves us with no clear reason why these features couldn’t just be omitted from the lexical entries of (1). Even if absence of the PRED-features caused some And, independently developed for XLE (Crouch, p.c.), although no longer used. Using p ‘proposition’ for the type of propositions rather than the usual t, and a clearly oversimplified Priorian operator treatment for tense. See for example Halvorsen (1983), Wedekind and Kaplan (1993), Frank and Semecky (2004), Crouch and King (2006), Crouch (2006). subtle problems, putting them back in would still constitute an explanatory problem, since there isn’t any principle that requires LFG lexical entries to introduce PRED-values. If the benefits of co-description were sufficiently impressive, one could presumably deal with this issue, but I will first show that the original motivation for it is insufficient, and point out that it creates various problems, one of which was noted by Andrews (2007a). Then I will describe a DBA architecture for glue, and show it it provides a role for PRED-features. But this is not the same as in pre-glue LFG, since glue will be doing the work of Completeness and Coherence (but not Predicate Uniqueness). So the last step is to propose some constraints which will cause meaning-constructors in the DBA architecture to act in a way that is similar in certain empirically justifiable respects to standard PRED-features controlling Completeness and Coherence, but avoiding the problems with co-description. 1 Problems and Non-benefits of Co-Description The main proposed benefit of co-description was that it could make available for semantic interpretation information not present in f-structure (Halvorsen and Kaplan 1988:284, 1995 version). But this ignores the fact that, thanks to the inverse of the φ projection, anything accessible from c-structure is also accessible from f-structure. Andrews (2007b), for example, proposes constraints involving c-structure in a DBA glue framework. However, it might still be the case that co-description is the best approach, either for all, or only for some, kinds of linguistic phenomena. Here I will argue that it isn’t best for what would be traditionally regarded as the interpretation of features and lexical items (by contrast, co-description seems very well suited for the properties of information-structure, c.f. Mycock (2006)). Perhaps the most immediate problem, pointed out in Andrews (2007a), is that it becomes an accident that the occurrences of features and their traditionally ascribed meanings are quite closely correlated, with only limited exceptions, such as pluralia tantum, which I’ll discuss later. There would for example be nothing obviously wrong with a variant of (1b) in which the plural meaning-constructor was present but not the plural feature-equation. But this doesn’t happen, even with the exotic plurals that English is so fond of borrowing from other languages: (2) a. These seraphim are annoyed b. This seraph is annoyed c. *This seraphim is annoyed (plural meaning, singular syntax) “Every interpretation scheme based on description-by-analysis requires that all semantically relevant information be encoded in the functional structure.” But agreement, the main motivation for having features at all, leads to a further problem with the meaning-constructors. This is that one has to decide which of the various lexical entries introducing a given feature-value occurrence is the one that is introducing the constructor. Consider an Italian example such as: (3) (le the(FEM.PL) ragazze) girl(FEM.PL) vengono come(3.PL) The girls/they are coming If the subject is present, one would presumably want the noun to introduce the plural meaning-constructor, and the verb not to (since not all NPs are in positions where there is a verb to agree with them and provide their number constructors), but if the subject is omitted, then the verb would presumably be the provider of the constructor. It is certainly not impossible to come up with grammars that will work properly, but it involves delicate choices with considerable scope for stipulation, which it would be good to reduce to the greatest extent possible. Another problem resides in the overlapping powers and responsibilities of the PRED-features, with their argument-lists, and those of the meaningconstructors that refer to grammatical functions. This is that, although the PRED-features control what governable grammatical functions can and must appear, they no longer say anything about what their semantic contributions are, since this is done by the meaning-constructors. But, left unconstrained, meaning-constructors can do all sorts of peculiar things in the way of rearranging the semantics of the grammatical functions. Below, for example, (a) interchanges the semantic role of subject and object, while (b) creates an unspecified causee agent causative: (4) a. λPxy.P (y, x): ((↑OBJ)e −◦ (↑ SUBJ)e −◦ ↑p)−◦ (↑ SUBJ)−◦ (↑OBJ)−◦ ↑p b. λPx.(∃z)(Cause(x, P (z, y))): ((↑ SUBJ)e −◦ ↑p)−◦ (↑ SUBJ)e −◦ ↑p Without some further constraints, these meaning-constructors could be introduced by inflections or grammatical particles, thereby undoing the kinds of work people have been trying to accomplish with Lexical Mapping Theory and its competitors over the last several decades. The most obvious and direct solution to the overlap problem is to drop the PRED-features entirely, since, as noted above, the resource management provided by linear logic can do all of the syntactic work of the PRED-features, and of course the meaning constructors also take over their informal role of encoding the meaning. Therefore, the natural consequence of adopting codescription is to abandon PRED-features. This might of course be the right thing to do, but I will argue in the remainder of the paper that glue-byDBA would be a good thing to try first for certain aspects of semantic interpretation, especially, morphology and the lexicon. However, note that the use of meaning-constructors introduced by the PS rules, for example by Asudeh and Crouch (2002) and Sadler and Nordlinger (2008), is not implicated in any of the problems raised here, and is consistent with what I will be proposing.
Conventional n-best reranking techniques often suffer from the limited scope of the n-best list, which rules out many potentially good alternatives. We instead propose forest reranking, a method that reranks a packed forest of exponentially many parses. Since exact inference is intractable with non-local features, we present an approximate algorithm inspired by forest rescoring that makes discriminative training practical over the whole Treebank. Our final result, an F-score of 91.7, outperforms both 50-best and 100-best reranking baselines, and is better than any previously reported systems trained on the Treebank. 1
Morphological processes in Semitic languages deliver space-delimited words which introduce multiple, distinct, syntactic units into the structure of the input sentence. These words are in turn highly ambiguous, breaking the assumption underlying most parsers that the yield of a tree for a given sentence is known in advance. Here we propose a single joint model for performing both morphological segmentation and syntactic disambiguation which bypasses the associated circularity. Using a treebank grammar, a data-driven lexicon, and a linguistically motivated unknown-tokens handling technique our model outperforms previous pipelined, integrated or factorized systems for Hebrew morphological and syntactic processing, yielding an error reduction of 12% over the best published results so far. 1
We have built a parallel treebank that includes word and phrase alignment. The alignment information was manually checked using a graphical tool that allows the annotator to view a pair of trees from parallel sentences. We found the compilation of clear alignment guidelines to be a difficult task. However, experiments with a group of students have shown that we are on the right track with up to 89% overlap between the student annotation and our own. At the same time these experiments have helped us to pin-point the weaknesses in the guidelines, many of which concerned unclear rules related to differences in grammatical forms between the languages.
We present the STYX system, which is designed as an electronic corpus-based exercise book of Czech morphology and syntax with sentences directly selected from the Prague Dependency Treebank, the largest annotated corpus of the Czech language. The exercise book offers complex sentence processing with respect to both morphological and syntactic phenomena, i. e. the exercises allow students of basic and secondary schools to practice classifying parts of speech and particular morphological categories of words and in the parsing of sentences and classifying the syntactic functions of words. The corpus-based exercise book presents a novel usage of annotated corpora outside their original context.
A web-based collaborative environment including on-line authoring tools that is managed by a central database was developed in collaboration with several countries including Peru, Bolivia, and the United States. The application involved developing a linguistics database and eLearning environment for documenting, preserving, and promoting language training for Aymara, a language indigenous to Peru and Bolivia. The database, an ontology management system called Lyra, incorporates all elements of the language (dialogues, phrase patterns, phrases, words, and morphemes) as well as cultural multimedia resources (images and sound recordings). The organization of the database enables a high level of integration among language elements and cultural resources. Authoring tools are used by experts in the Aymara language to build the linguistic database. These tools are accessible on-line as part of the collaborative environment using standard web browsers incorporating the Java plug-in. The eLearning student interface is a web-based program written in Flash. The Flash program automatically interprets and formats data objects retrieved from the database in XML format. The student interface is presented in Spanish and English. A web service architecture is used to publish the database on-line so that it can be accessed and utilized by other application programs in a variety of formats
In my thesis I have attempted to develop an integrated translation approach materialized in the form of a Dynamic Translation Model (DTM). This endeavour can be justified to the extent that Translation Studies is perceived so far as a fragmentary discipline with implicitly and explicitly opposed and apparently irreconcilable points of view: linguistics-oriented approaches and culture-and-literature-oriented approaches. The main problem arising from this lack of common ground for further developing Translation Studies is that the disciplinary boundaries are not well-established and therefore the discipline itself cannot be developed coherently. Besides, Translation Studies is still to be constructed as an autonomous and an independent discipline that has a common core of theoretical and practical problems. This lack of coherent development of the discipline is due, I think, to an epistemological mistake: to believe that one single approach can account for (that is, describe and explain) all the translational reality. I propose to distinguish a two-phase epistemological move: 1. each translation approach works on its own research interests and acknowledges that its approach deals only with one part of the whole subject matter of Translation Studies; and 2. the results obtained by each translation approach are incorporated into a holistic integrative model like the Dynamic Translation Model I propose. In order to achieve this goal I have attempted to show the key tenets of modern translation approaches, both linguistics-oriented and culture-and-literature-oriented, by quoting the main theses of the representatives of these approaches. I have then presented the most important criticisms that have been raised in relation to these diverse translation approaches, together with my own criticisms (chapters 1 and 2). Also, I have introduced the theoretical basis for an integrated approach taking Holmes’ differentiation between theoretical (product-, process-, and function-oriented) and practical approaches as a point of departure. Likewise, I have discussed the problems of integrating Translation Studies, as well as Snell-Hornby’s integrated proposal and some key aspects of literary translation relevant for my integrative endeavour (chapter 3). Finally, I have developed my proposal for a Dynamic Translation Model (chapter 4). As to the conclusions of my thesis, I can say that my holistic DTM was able to integrate functionally aspects from both linguistics-oriented and culture-and-literature-oriented approaches: historico-cultural context (Leipzig School and postcolonial studies); norms, ideology and power (Descriptive Translation Studies; G. Toury and A. Lefevere); translation commisioner (Skopos theory); sender’s communicative purpose (linguistic and pragmatic approaches: W. Koller, J. House, H. Gerzymisch-Arbogast, etc); importance of source language text (linguistic and textlinguistic approaches; stylistic approaches; B. Spillner, B. Sandig); translator’s comprehension process (hermeneutic, deconstructive, and poststructural approaches), target language receiver in the target language historico-cultural context (Descriptive Translation Studies; postcolonial and gender studies). On the other hand, the three levels of the Dynamic Translation Model help to explain the flux of translational proceses and the variables that are activated or neutralized therein. They also incorporate concepts from other disciplines such as text linguistics, pragmatics, stylistics, and the communication theory. In my integrative endeavour I also proposed new concepts and, accordingly, coined new terms: Compulsory Translational Forces (CTF) (which include both Initiator’s Translational Instructions (ITI) and Target Language Valid Translational Norms (TL-VTN), Default Equivalence Position (DEP). In the pragmatic dimension of the model special attention is paid to what I call Text Illocutionary Indicators (TII) as well as the strengthening (upgraders) and weakening (downgraders) illocutionary mechanisms in relation to the Source Language Text (SLT) and the Target Language Text (TLT). Semantic/lexical fields play a crucial role in the establishment of equivalences between SLT and TLT in the text semantic dimension, as well as what I have called Fictionalizing Stylistic Shifts in the text stylistic dimension. As to the future developments of translation research within the framework of the Dynamic Translation Model I would say that some modificationbs may be called for so that interpretation can also be accounted for. This proposal can be used profitably in the field of translation criticism. As is the case with any other integrative approach, DTM should be widely discussed and criticized in order to validate its theoretical soundness and its application in Translation Studies. This thesis is an attempt to contribute in this research direction.
Sciendo provides publishing services and solutions to academic and professional organizations and individual authors. We publish journals, books, conference proceedings and a variety of other publications.
The lexical information of verbal lexemes, such as verbs and adjectives, plays an important role in syntactic parsing, because the structure of a sentence mainly hinges on the type of verbal lexemes. The question we address in this research is how to acquire the argument structure (henceforth ARG-ST) of verbal lexemes in Korean. It is well known that manual build-up of type hierarchy usually cost too much time and resources, so an alternative method, namely automatic collection of relevant information is much more preferred. This paper proposes a procedure to automatically collect ARG-ST of Korean verbal lexemes from a Korean Treebank. Specifically, the system we develop in this paper first extracts lexical information of ARG-ST of verbal lexemes from a 0.8 million graphic word Korean Treebank in an unsupervised way, checks the hierarchical relationship among them, and builds up the type hierarchy automatically. The result is written in an HPSG-style annotation, thus making it possible to readily implement the result in an HPSG-based parser for Korean. Finally, the result is evaluated with reference to two Korean dictionaries and also with respect to a manually constructed type hierarchy.
This paper describes an ongoing effort to build a large-scale monolingual treebank of parallel/comparable Dutch text, where nodes of syntax trees are aligned and labeled according to a small set of semantic similarity relations.Such a corpus has many potential uses and applications in e.g., multi-document summarization, question-answering and paraphrase extraction.We describe the text material, preprocessing, annotation, and alignment of sentences and syntax trees, both manual and automatic.Two new annotation tools are presented, as well as results from pilot experiments on inter-annotator agreement.On the basis of this resource, new automatic alignment software and NLP applications will be developed.
We present the second version of the Penn Discourse Treebank, PDTB-2.0, describing its lexically-grounded annotations of discourse relations and their two abstract object arguments over the 1 million word Wall Street Journal corpus. We describe all aspects of the annotation, including (a) the argument structure of discourse relations, (b) the sense annotation of the relations, and (c) the attribution of discourse relations and each of their arguments. We list the differences between PDTB-1.0 and PDTB-2.0. We present representative statistics for several aspects of the annotation in the corpus. 1.
Graph-based and transition-based approaches to dependency parsing adopt very different views of the problem, each view having its own strengths and limitations. We study both approaches under the framework of beamsearch. By developing a graph-based and a transition-based dependency parser, we show that a beam-search decoder is a competitive choice for both methods. More importantly, we propose a beam-search-based parser that combines both graph-based and transitionbased parsing into a single system for training and decoding, showing that it outperforms both the pure graph-based and the pure transition-based parsers. Testing on the English and Chinese Penn Treebank data, the combined system gave state-of-the-art accuracies
We describe a parsing approach that makes use of the perceptron algorithm, in conjunction with dynamic programming methods, to recover full constituent-based parse trees. The formalism allows a rich set of parse-tree features, including PCFG-based features, bigram and trigram dependency features, and surface features. A severe challenge in applying such an approach to full syntactic parsing is the efficiency of the parsing algorithms involved. We show that efficient training is feasible, using a Tree Adjoining Grammar (TAG) based parsing formalism. A lower-order dependency parsing model is used to restrict the search space of the full model, thereby making it efficient. Experiments on the Penn WSJ treebank show that the model achieves state-of-the-art performance, for both constituent and dependency accuracy.
To date, parsers have made limited use of semantic information, but there is evidence to suggest that semantic features can enhance parse disambiguation. This paper shows that semantic classes help to obtain significant improvement in both parsing and PP attachment tasks. We devise a gold-standard sense- and parse tree-annotated dataset based on the intersection of the Penn Treebank and SemCor, and experiment with different approaches to both semantic representation and disambiguation. For the Bikel parser, we achieved a maximal error reduction rate over the baseline parser of 6.9% and 20.5%, for parsing and PP-attachment respectively, using an unsupervised WSD strategy. This demonstrates that word sense information can indeed enhance the performance of syntactic disambiguation. © 2008 Association for Computational Linguistics.
Research on the second language acquisition (SLA) of Spanish has identified grammatical structures for which an analysis of errors for second-language (L2) learners is inappropriate (Geeslin, 2003; Geeslin & Guijarro-Fuentes, 2006; Gudmestad, 2006). This is because the norms of use for such structures are changing, and prescriptive grammars do not coincide with actual language use. Thus, in order to examine such sociolinguistically-variable grammatical features in learner language, researchers have shifted to an analysis of the predictors of use of a given variant, rather than an assessment of accuracy (Geeslin, 2000). Investigations following this approach on structures such as copula choice and mood choice have been largely based on written contextualized tasks (WCT), where use is contextualized and participants indicate a preference for one of the two possible variants. The advantage of this type of task is twofold. First, in comparison with grammaticality judgment tasks, participants are not forced to select one (presumably the only) grammatical sentence from the options provided. Consequently, the WCT is more in line with the idea that variation is indeed an acceptable, and even irrefutable, part of native-like speech. Secondly, in comparison with tasks that elicit less directed production, the WCT assures that each participant will respond to the same tokens (both lexically and in terms of the contextual features that predict selection of a given variant) and that each
"Treebanks allow for the creation of a valence lexicon per side effect. The TüBa-D/Z valence lexicon has been created in lockstep with the development of the TüBa- D/Z treebank as such. For each verb encountered in the treebank, the annotators created a lexical entry that records the valence frames of the verbs contained in the sentence, unless they are already contained in the valence lexicon as result of previous annotation. The TüBa-D/Z valence lexicon currently contains a total of 8013 frames for 4896 distinct verb lemmas. Since treebank annotation is still ongoing, the lexicon will continue to grow. Such a lexicon has utility in its own right as a resource for lexicalized parsing and a variety of NLP applications. At the same time, the lexicon can serve as a source for aiding consistency of annotation and automatic detection of annotation errors"
In this article we report work on Chinese semantic role labeling, taking advantage of two recently completed corpora, the Chinese PropBank, a semantically annotated corpus of Chinese verbs, and the Chinese Nombank, a companion corpus that annotates the predicate-argument structure of nominalized predicates. Because the semantic role labels are assigned to the constituents in a parse tree, we first report experiments in which semantic role labels are automatically assigned to hand-crafted parses in the Chinese Treebank. This gives us a measure of the extent to which semantic role labels can be bootstrapped from the syntactic annotation provided in the treebank. We then report experiments using automatic parses with decreasing levels of human annotation in the input to the syntactic parser: parses that use gold-standard segmentation and POS-tagging, parses that use only gold-standard segmentation, and fully automatic parses. These experiments gauge how successful semantic role labeling for Chinese can be in more realistic situations. Our results show that when hand-crafted parses are used, semantic role labeling accuracy for Chinese is comparable to what has been reported for the state-of-the-art English semantic role labeling systems trained and tested on the English PropBank, even though the Chinese PropBank is significantly smaller in size. When an automatic parser is used, however, the accuracy of our system is significantly lower than the English state of the art. This indicates that an improvement in Chinese parsing is critical to high-performance semantic role labeling for Chinese.
THE VOICE TEACHER IS REGULARLY BESET WITH CHALLENGES in the studio regarding consonant clusters in sung German, as is the singer who approaches any vocal work in the German language. The reputation of the German language as being consonant rather than vowel oriented is commonly appreciated and justifiable. Statistical studies show that the burden of text intelligibility is carried principally by the consonants, to a greater extent than most languages. A language that can produce lexical items such as entsturzt [ent'∫tYrtst] and kraftstrotzend ['kraft∫trctsent] adopts a strongly marked position among the world's languages with respect to the involvement of consonants in its sound system. These words contain ten and fourteen phonemes respectively, of which only two or three are vowels. The remaining clusters of consonants are samples of the subject of this article. The consonant clusters normally encountered in German phonology will be inventoried and contrasted with English. The material is likely to be familiar to many readers, albeit presented in a different, perhaps more systematic perspective than is normally encountered. The subject of German consonant clusters is best dealt with in terms of phonetic, not orthographic consonants. A firm grasp of the relationship between spelling and pronunciation is naturally also essential. Two or three successive letters may represent a single phoneme, as in [arrow right] /c/ or /x/ [arrow right] /∫/ [arrow right] /k/ [arrow right] /t/ Conversely, a single written consonant may serve to indicate more than one phoneme, as in [arrow right] /ts/ This situation is familiar because it is even more pronounced in English. The word scythe contains two consonant digraphs and two letter-vowels, but phonetically only one diphthong and no clusters at all. Since the greatest challenge in consonant clusters is visual (i.e., orthographic), thinking in terms of phonetic consonants should serve to simplify the matter for a student. German, more than most other languages, has absorbed lexical items from other languages into its own vocabulary, particularly from English, French, and Italian. Thus Duden, the principal lexicographic publisher in modern Germany, devotes an entire book to Fremdworter in its series of dictionaries. The process of lexical transfer is a complex aspect of German linguistics, particularly regarding pronunciation norms. Some words, such as Situation, have been subsumed into the phonological patterning of German, while others have retained the pronunciation of the word in the language from whence it came, or have struck a middle ground, such as Orange and Weekend. This diversity of phonetic transfer gives modern spoken German a particular flavor, and reflects the country's central geographic position in Europe. There are similar examples in English, such as cul-de-sac (where the French pronunciation has been distorted) and naive (which retains the original, although English idiosyncratically employs only the feminine form). This article will confine itself to the consonant clusters that occur regularly in the standard lexis, referring to combinations resulting from foreign influences only when appropriate. It is useful to consider consonant clusters in two quite distinct groups: syllable-interior, and across syllable or word boundaries. Part I of the article will concern itself with the former; Part II (to appear in the March/April 2008 issue), with the latter. Recognition of which group an example belongs to is the first step toward establishing correct pronunciation, and in some cases is necessary to discriminate between two potentially correct pronunciations. Before outlining in tabular form the cluster environments of German and English, it will be useful to consider all the consonantal combinations that are admissible in each language. A detailed theoretical account of the phonotactic rules and constraints of each language will not be necessary for our purposes. …
Sometreebanks, such as German TIGER/NeGra, represent discontinuous elements directly, i.e. trees contain crossing edges, but the context-free grammars that are extracted from them, fail to make any use of this information. In this paper, we present amethod for extracting mildly context-sensitive grammars, i.e. simple range concatenation grammars (RCGs), from such treebanks. A measure for the degree of a treebank’s mild contextsensitivity is presented and compared to similar measures used in non-projective dependency parsing. Our work is also compared to discontinuous phrase structure grammar (DPSG).
Abstract Affective ratings of multiple religious (sub)groups (Muslims, Christians, Jews and non-believers, as well as Sunni, Alevi and Sjiit Muslims), the endorsement of Islamic minority rights and religious group identification were examined among Sunni and Alevi Turkish-Dutch participants. The findings show that both groups differ in important ways. Some Alevi participants considered themselves Muslims but others interpreted Alevi identity in a secular way. The Sunnis were quite negative towards Jews and non-believers, they more strongly endorsed Islamic minority rights and they had very high Muslim group identification. Furthermore, the Sunnis were negative towards Alevis and the Alevis were negative towards the Sunnis. Muslim group identification was positively and strongly related to feelings towards Muslims and to the endorsement of Islamic group rights.
Abstract Large linguistic databases, especially databases having a global coverage, such as the World Atlas of Language Structures, the Automated Similarity Judgment Program, and Ethnologue, are making it possible to systematically investigate many aspects of how languages change and compete for viability. Agent‐based computer simulations supplement such empirical data by analyzing the necessary and sufficient parameters for the current global distributions of languages or linguistic features. By combining empirical datasets with simulations and applying quantitative methods, it is now possible to address fundamental questions, such as ‘what are the relative rates of change in different parts of languages?’, ‘why are there a few large language families, many intermediate ones, and even more small ones?’, ‘do small languages change faster or slower than large ones?’, or ‘how does the borrowing of words relate to the borrowing of structural features?’
This paper describes first steps towards extending the METU Turkish Corpus from a sentence-level language resource to a discourse-level resource by annotating its discourse connectives and their arguments. The project is based on the same principles as the Penn Discourse TreeBank (http://www.seas.upenn.edu/~pdtb) and is supported by TUBITAK, The Scientific and Technological Research Council of Turkey. We first present the goals of the project and the METU Turkish corpus. We then describe how we decided what to take as explicit discourse connectives and the range of syntactic classes they come from. With representative examples of each class, we examine explicit connectives, their linear ordering, and types of syntactic units that can serve as their arguments. We then touch upon connectives with respect to free word order in Turkish and punctuation, as well as the important issue of how much material is needed to specify an argument. We close with a brief discussion of current plans.
Previous research found that the duration of segments decreases as children grow older. The development of suprasegmental duration, however, has not been explored. The present study investigated developmental changes in duration of the four Mandarin tones. 5-, 8-, and 12-year-old monolingual Mandarin-speaking children and young adults participated in the study. Tone durations were measured in participants’ production of monosyllabic target words elicited by picture identification tasks. The results were as follows (1) For each tone category, tone duration and variability decreased with age: 5- and 8-year-old children showed significantly longer durations than adults. Tone durations in 12-year-old children approximated adult values. (2) Despite longer durations, adultlike duration patterns across tone categories existed in all children: dipping tones were the longest, followed by rising and level tones, with falling tones being the shortest. (3) Duration differences between the rising and dipping tones became larger as children grew older. The results may be indicative of the general maturation of laryngeal control over age. Although 5- and 8-year-old children have already established lexical contrasts of tone, adultlike phonetic norms are still in the process of development. The developmental data also provide support for a hybrid account of speech production from a suprasegmental perspective.
In this paper the role of concept characteristics in lexical dialectometric research is examined in three consecutive logical steps. First, a regression analysis of data taken from a large lexical database of Limburgish dialects in Belgium and The Netherlands is conducted to illustrate that concept characteristics such as concept salience, concept vagueness and negative affect contribute to the lexical heterogeneity in the dialect data. Next, it is shown that the relationship between concept characteristics and lexical heterogeneity influences the results of conventional lexical dialectometric measurements. Finally, a dialectometric procedure is proposed which downplays this undesired influence, thus making it possible to obtain a clearer picture of the ‘truly’ regional variation. More specifically, a lexical dialectometric method is proposed in which concept characteristics form the basis of a weighting schema that determines to which extent concept specific dissimilarities can contribute to the aggregate dissimilarities between locations.
Specific phobias are the most common anxiety disorder and are characterized by avoidance behavior. Avoidance behavior impacts daily function and is proposed to impair extinction learning. However, despite its prevalence, its objective assessment remains a challenge. To this end, we developed a fully automated experimental procedure using immersive virtual reality. The procedure contained a behavioral search, forced-choice, and an approach task with varying degrees of freedom and task relevance of the stimuli. In this study, we examined the sensitivity and feasibility of these tasks to assess avoidance behavior in patients with specific phobia. We adapted the tasks by replacing the originally conditioned stimuli with a spider and a neutral animal and investigated 31 female participants composed of 15 spider-phobic and 16 non-phobic participants. As the non-phobics were quite heterogeneous in terms of their Fear of Spiders Questionnaire (FSQ) scores, we subdivided them into six "fearfuls" that had elevated FSQ scores, and 10 "non-fearfuls" that had no fear of spiders. The phobics successfully managed to complete the procedure and showed consistent avoidance behavior across all behavioral tasks. Compared to the non-fearfuls, which did not show any avoidance behavior at all, the phobics looked at the spider much more often and clearly directed their body toward it in the search task. In the approach task, they hesitated most when they were close to the spider, and their difficulty to touch the spider was reflected in a strong increase in right hand acceleration changes. The fearfuls showed avoidance behavior depending on the tasks: strongest in the search task and weakest in the approach task. Additionally, we identified subjective valence ratings of the spider as the main influence on both objective avoidance behavior and subjective well-being after exposure, mediating the effect of the FSQ. In summary, the behavioral tasks are well suited to assess avoidance behavior in phobic participants and provide detailed insights into the process of avoidance.
In order to understand how emotional state influences the listener's physiological response to speech, subjects looked at emotion-evoking pictures while 32-channel EEG evoked responses (ERPs) to an unchanging auditory stimulus ("danny") were collected. The pictures were selected from the International Affective Picture System database. They were rated by participants and differed in valence (positive, negative, neutral), but not in dominance and arousal. Effects of viewing negative emotion pictures were seen as early as 20 msec (p =.006). An analysis of the global field power highlighted a time period of interest (30.4-129.0 msec) where the effects of emotion are likely to be the most robust. At the cortical level, the responses differed significantly depending on the valence ratings the subjects provided for the visual stimuli, which divided them into the high valence intensity group and the low valence intensity group. The high valence intensity group exhibited a clear divergent bivalent effect of emotion (ERPs at Cz during viewing neutral pictures subtracted from ERPs during viewing positive or negative pictures) in the time period of interest (r(Phi) =.534, p <.01). Moreover, group differences emerged in the pattern of global activation during this time period. Although both groups demonstrated a significant effect of emotion (ANOVA, p =.004 and.006, low valence intensity and high valence intensity, respectively), the high valence intensity group exhibited a much larger effect. Whereas the low valence intensity group exhibited its smaller effect predominantly in frontal areas, the larger effect in the high valence intensity group was found globally, especially in the left temporal areas, with the largest divergent bivalent effects (ANOVA, p <.00001) in high valence intensity subjects around the midline. Thus, divergent bivalent effects were observed between 30 and 130 msec, and were dependent on the subject's subjective state, whereas the effects at 20 msec were evident only for negative emotion, independent of the subject's behavioral responses. Taken together, it appears that emotion can affect auditory function early in the sensory processing stream.
Graph-based and transition-based approaches to dependency parsing adopt very different views of the problem, each view having its own strengths and limitations. We study both approaches under the framework of beam-search. By developing a graph-based and a transition-based dependency parser, we show that a beam-search decoder is a competitive choice for both methods. More importantly, we propose a beam-search-based parser that combines both graph-based and transition-based parsing into a single system for training and decoding, showing that it outperforms both the pure graph-based and the pure transition-based parsers. Testing on the English and Chinese Penn Treebank data, the combined system gave state-of-the-art accuracies of 92.1% and 86.2%, respectively.
Approximately 35% of individuals with dementia exhibit depression and/or anxiety symptoms, often manifested by symptoms of negative affect. Exercise has been associated with improved affect but has not been demonstrated to improve affect in residents of secured dementia units in long-term care facilities. This pilot study determined whether moderate-intensity, chair-based exercise was associated with changes in negative affect in residents in secured units. The sample included 36 patients from 2 nursing homes who participated in a 12-week, 30-minute moderate-intensity group exercise program thrice weekly. Affect, measured by the Philadelphia Geriatric Center Apparent Affect Rating Scale, was assessed at weeks 3 and 12, before and after each exercise session. Paired t tests assessed the immediate effect of exercise (before/after a session) and the long-term effect of exercise (study initiation/12 wk) on patients' affect ratings. The mean age was 85 years (SD=5.5), with 86% female, and 97% white. At week 3, anxiety was significantly lower immediately after the exercise session when adjusted for level of participation (P=0.02) compared with immediately before the exercise session, indicating immediate changes in affect. Anxiety and depression were significantly reduced at week 12, when compared with week 3, after adjusting for level of participation (P=0.01; P=0.03), indicating long-term effects of the exercise intervention. The study revealed the feasibility of conducting a moderate-intensity exercise program and the potential for exercise as a nonpharmacologic intervention for reducing symptoms of negative affect and depression in this vulnerable population.
Correct identification of word meaning is a long-standing problem for lexicography, language teaching, linguistic theory, and computer processing of text. Traditional approaches typically proceed word by word, relying on evidence from introspection – and have failed. A new theory of meaning is needed. In this prototype-based approach, called the Theory of Norms and Exploitations (TNE), the first step is identifying the phraseological patterns with which each word is associated. Meanings are then associated with patterns, rather than with isolated words. Words are highly ambiguous, but patterns are mostly unambiguous. \n Patterns cannot be identified by valency alone, but require statistical analysis and semantic typing of collocates. For example, (1) blowing up a bridge and (2) blowing up a balloon activate different meanings of blow up. But how many other contexts have the same effect on the meaning of the phrasal verb? Relevant members of the lexical set for (1) include building, factory, house, hotel, etc. Such lexical sets provide a basis for machine learning and text processing. \n Authentic uses of words are classified either as normal components of a pattern or as exploitations of norms. For example, “blowing up a condom” is not normal, but exploits (2). Creative metaphors are also exploitations.
Predicting Word-Naming and Lexical Decision Times from a Semantic Space Model Brendan T. Johns (johns4@indiana.edu) Department of Psychological and Brain Sciences, 1101 E. Tenth St. Bloomington, In 47405 USA Michael N. Jones (jonesmn@indiana.edu) Department of Psychological and Brain Sciences, 1101 E. Tenth St. Bloomington, In 47405 USA Abstract organization of semantic memory. In a lexical decision task, a letter string is presented and the participant provides a speeded response of whether the string is a word or not. In a naming task, the participant’s task is to name the presented word aloud as quickly as possible. Both measures produce an index of a word’s identification latency. Orthographic and phonological factors are certainly large components of both LDT and NT, but semantics plays a significant role as well, and co-occurrence models have yet to be extended to predicting reaction time variance for these single-word identification tasks. Modeling of retrieval times is usually done by looking for the best environmental correlates of LDT and NT (Adelman & Brown, 2008). Some of the most influential models of retrieval times are based upon word frequency. Word frequency (WF) has been used to drive many different types of models, including serial-searched rank frequency models (Murray & Forster, 2004), threshold activation models (Coltheart, et al., 2001), and connectionist models (Seidenberg & McClelland, 1989). However, recent evidence suggests that word frequency may not drive retrieval times but, rather, the causal factor is a word’s contextual diversity (Adelman, Brown, & Quesada, 2006; Adelman & Brown, 2008). Contextual diversity (CD) is the number of different contexts that a word appears in, and is based on the rational analysis of memory (Anderson & Milson, 1989), particularly the principle of likely need (PLN). PLN states that the more unique contexts a word appears in, the more likely the word will be needed in any future context. Hence, a word with a high CD should be faster to retrieve under this principle. A word’s CD value is typically computed by simply counting the number of different documents in which it appears across a text corpus. This measure has been shown to be a better predictor of LDT and NT than WF (Adelman, et al., 2006). However, operationalizing CD as the number of documents in which a word occurs may not be a fair instantiation of PLN. A word that appears in many documents may have a high WF, but it should have a low CD if those documents are highly redundant, as is the case with words that belong to a popular discourse topic for which many documents exist. It is the number of different contexts and the uniqueness of contexts that determines a word’s likely need. This calls for a measure of CD that considers the semantic uniqueness of documents that a word appears in. Based on PLN, it is reasonable to assume that if a word appears in a context it has never before occurred in, We propose a method to derive predictions for single-word retrieval times from a semantic space model trained on text corpora. In Experiment 1 we present a large corpus analysis demonstrating that it is the number of unique semantic contexts a word appears in across language, rather than simply the number of contexts or the frequency of the word, that is the most salient predictor of lexical decision and naming times. In Experiment 2, we develop a co-occurrence learning model that weights new contextual uses of a word based on fit to what currently exists in the word’s memory representation, and demonstrate this model’s superiority in fitting the human data compared to models built using information about the word’s frequency or number of contexts. Finally, in Experiment 3 we find that building lexical representations using semantic distinctiveness naturally produces a better-organized semantic space to make predictions for semantic similarity between words. Keywords: Co-occurrence model; Lexical-decision; LSA; Contextual distinctiveness Introduction The last decade has seen remarkable progress with co- occurrence models of lexical semantics (e.g., Lund & Burgess, 1996; Landauer & Dumais, 1997). These models learn semantic representations for words by observing lexical co-occurrence patterns across a large text corpus, typically representing the words in a high-dimensional semantic space. This approach provides both an account of the semantic representation for words and an account of the learning mechanisms humans use to build and organize semantic memory. Co-occurrence models have seen considerable success at accounting for data in a wide variety of semantic tasks, including TOEFL synonyms (Landauer & Dumais, 1997), semantic similarity ratings and exemplar categorization (Jones & Mewhort, 2007), and free association norms (Griffiths, Steyvers, & Tenenbaum, To date, all applications of co-occurrence models have been to semantic similarity between two words or two documents. The standard prediction of semantic similarity in these models is some measure of the angle between two vectors. However, co-occurrence models should, in theory, contain sufficient information in the magnitude of their representations to make predictions about single word retrieval as well. Lexical decision time (LDT) and word naming time (NT) are both important variables that offer insight into the
Cornetto byl dvouletý projekt (STE05039), ve kterem byla vytvořena lexikalni semanticka databaze kombinujici Wordnet s informacemi typu FrameNet pro holandstinu. Kombinaci těchto lexikalnich zdrojů vznikla výrazně bohatsi databaze jazykových vztahů, ktera umožni kvalitnějsi výsledky technologii zpracovani přirozeneho jazyka, jako je desambiguace významu slov (WSD) a systemy generovani jazyka. Kromě propojeni Wordnetu s informacemi typu FrameNet je databaze take mapovana na formalni ontologii, ktera poskytuje přesne semanticke popisy.
Foreign researchers have made great achievements in annotating English text structures with Rhetorical Structure Theory(RST),which provides profound implications for Chinese text annotation.However,there are many differences between Chinese texts and English ones.In fact,Chinese sentence group theory can function as well as RST Theory in annotating Chinese texts.Firstly,the basic framework of RST is quite similar to that of Chinese sentence group theory;secondly,RST theory is less adaptable to Chinese than to English as its analysis values clauses and conjunctions,which are not so salient in Chinese;lastly,texts can be found in Qinghua Treebank whose sentence groups have been marked.