Papers reviewed and determined not to be word norm studies. Use the flag icon to report errors or suggest re-inclusion.
16504 papers
This article examines the usefulness ofvocabulary richness for authorship attributionand tests the assumption that appropriatemeasures of vocabulary richness can capture anauthor's distinctive style or identity. Afterbriefly discussing perceived and actualvocabulary richness, I show that doubling andcombining texts affects some measures incomputationally predictable but conceptuallysurprising ways. I discuss some theoretical andempirical problems with some measures anddevelop simple methods to test how wellvocabulary richness distinguishes texts bydifferent authors. These methods show thatvocabulary richness is ineffective for largegroups of texts because of the extremevariability within and among them. I concludethat vocabulary richness is of marginal valuein stylistic and authorship studies because thebasic assumption that it constitutes awordprint for authors is false.
We investigated the reliability and validity of a video-based method of measuring the magnitude of children’s emotion-modulated startle response when electromyographic (EMG) measurement is not feasible. Thirty-one children between the ages of 4 and 7 years were videotaped while watching short video clips designed to elicit happiness or fear. Embedded in the audio track of the video clips were acoustic startle probes. A coding system was developed to quantify from the video record the strength of the eye-blink startle response to the probes. EMG measurement of the eye blink was obtained simultaneously. Intercoder reliability for the video coding was high (Cohen’sκ = .90). The average within-subjects probe-by-probe correlation between the EMG- and video-based methods was .84. Group-level correlations between the methods were also strong, and there was some evidence of emotion modulation of the startle response with both the EMG- and the video-derived data. Although the video method cannot be used to assess the latency, probability, or duration of startle blinks, the findings indicate that it can serve as a valid proxy of EMG in the assessment of the magnitude of emotion-modulated startle in studies of children conducted outside of a laboratory setting, where traditional psychophysiological methods are not feasible.
In this article, we present the spatial logistics task (SLOT) platform for investigating multimodal communication between 2 human participants. Presented are the SLOT communication task and the software and hardware that has been developed to run SLOT experiments and record the participants’ multimodal behavior. SLOT offers a high level of flexibility in varying the context of the communication and is particularly useful in studies of the relationship between pen gestures and speech. We illustrate the use of the SLOT platform by discussing the results of some early experiments. The first is an experiment on negotiation with a one-way mirror between the participants, and the second is an exploratory study of automatic recognition of spontaneous pen gestures. The results of these studies demonstrate the usefulness of the SLOT platform for conducting multimodal communication research in both human-human and human-computer interactions.
Sociallexicological description of non-standard vocabulary of English and Russian military sublanguages includes notions of sociallinguistic norm, existential form of language, diglossia, national variant, literary standard, popular language, sublanguage, sociolect, lexical systems of sublanguage, military sublanguage and military sociolect, invective. Classified foundation of words stock of language is social-communicative stratification of components of literary standard and lexical popular language.
This paper reports on the use of two distinct evaluation metrics for assessing a stochastic parsing model consisting of a broad-coverage Lexical-Functional Grammar (LFG), an efficient constraint-based parser and a stochastic disambiguation model. The first evaluation metric measures matches of predicate-argument relations in LFG f-structures (henceforth the LFG annotation scheme) to a gold standard of manually annotated f-structures for a subset of the UPenn Wall Street Journal treebank. The other metric maps predicate-argument relations in LFG f-structures to dependency relations (henceforth DR annotations) as proposed by Carroll et al. (Carroll et al., 1999). For evaluation, these relations are matched against Carroll et al.'s gold standard which was manually annnotated on a subset of the Brown corpus. The parser plus stochastic disambiguator gives an F-measure of 79% (LFG) or 73% (DR) on the WSJ test set. This shows that the two evaluation schemes are similar in spirit, although accuracy is impaired systematically by mapping one annotation scheme to the other. A systematic loss of accuracy is incurred also by corpus variation: Training the stochastic disambiguation model on WSJ data and testing on Carroll et al.'s Brown corpus data yields an F-score of 74% (DR) for dependency-relation match. A variant of this measure comparable to the measure reported by Carroll et al. yields an F-measure of 76%. We examine divergences between annotation schemes aiming at a future improvement of methods for assessing parser quality.
(2002) Br J Psychiatry 180, 523; Turkington D, Kingdon D, Turner T.. Effectiveness of a brief cognitive-behavioural therapy intervention in the treatment of schizophrenia..;.:. –7. [OpenUrl][1][Abstract/FREE Full Text][2] QUESTION: In patients with schizophrenia in secondary care settings, does cognitive behavioural therapy (CBT) delivered by community psychiatric nurses (CPNs) improve symptoms? Randomised {allocation concealed*}†, unblinded,* controlled trial with 2–3 months of follow up. 6 centres in the UK (Belfast, Glasgow, Hackney, Newcastle, Southampton, and Swansea). 422 patients who were 18–65 years of age (mean age 40 y, 77% men) and were receiving treatment from psychiatric secondary care services. Exclusion criteria were need for inpatient care or intensive home treatment, primary diagnosis of drug or alcohol dependence, organic brain disease, or learning disability that could affect rating. Follow up was 84%. Patients were allocated to CBT … [1]: {openurl}?query=rft.jtitle%253DThe%2BBritish%2BJournal%2Bof%2BPsychiatry%26rft.stitle%253DBr.%2BJ.%2BPsychiatry%26rft.issn%253D0007-1250%26rft.aulast%253DTURKINGTON%26rft.auinit1%253DD.%26rft.volume%253D180%26rft.issue%253D6%26rft.spage%253D523%26rft.epage%253D527%26rft.atitle%253DEffectiveness%2Bof%2Ba%2Bbrief%2Bcognitive--behavioural%2Btherapy%2Bintervention%2Bin%2Bthe%2Btreatment%2Bof%2Bschizophrenia%26rft_id%253Dinfo%253Adoi%252F10.1192%252Fbjp.180.6.523%26rft_id%253Dinfo%253Apmid%252F12042231%26rft.genre%253Darticle%26rft_val_fmt%253Dinfo%253Aofi%252Ffmt%253Akev%253Amtx%253Ajournal%26ctx_ver%253DZ39.88-2004%26url_ver%253DZ39.88-2004%26url_ctx_fmt%253Dinfo%253Aofi%252Ffmt%253Akev%253Amtx%253Actx [2]: /lookup/ijlink?linkType=ABST&journalCode=bjprcpsych&resid=180/6/523&atom=%2Febmed%2F8%2F1%2F23.atom
this paper I will discuss a framework for semantics which allows us to record truth-conditional and compositional analyses as dependency-style corpus annotations in a direct and fine-grained fashion. This method eliminates the need for a semantic representation formalism by decomposing semantic information into simple statements about (word or morpheme) tokens. A collection of such data would form a new kind of linguistic treebank. The main purpose of this article is to show that the present approach makes it possible to combine formal semantics and corpus-oriented study of language use in new and interesting ways. The methodology of this framework, which I call Token Dependency Semantics (TDS, Dahllf [4]), is in several respects different from the common one(s) in traditional formal semantics. TDS nevertheless delivers a fairly conventional (but ontologically restrained) analysis of truth-conditional meaning
Abstract. The problem of Prepositional Phrase (PP) attachment disambiguation consists in determining if a PP is part of a noun phrase, as in He sees the room with books, or an argument of a verb, as in He fills the room with books. Volk has proposed two variants of a method that queries an Internet search engine to find the most probable attachment variant. In this paper we apply the latest variant of Volk’s method to Spanish with several differences that allow us to attain a better performance close to that of statistical methods using treebanks. 1
Natural language generation (NLG) is the task of formulating a fluent sequence of words in natural language to communicate information or ideas in applications like machine translation, human-computer dialogue, automatic summarization, and question-answering. Realization, a fundamental subtask of NLG, produces an individual sentence from a sentence plan specified in terms of linguistic relations between words and/or concepts. It involves determining the order of words, inserting function words like determiners and prepositions, performing morphological inflections, and ensuring grammaticality and agreement. An ultimate goal for natural language generation is to develop a large-scale, robust, general-purpose system. Two primary challenges are scaling up to broad coverage of syntax and producing high quality output. The irregularity of natural language makes it difficult to know how to combine linguistic primitives into fluent sentences. Also, the knowledge resources for making such a determination are time-consuming and labor-intensive to assemble, leading to a knowledge acquisition bottleneck. Evaluating whether a realizer performed appropriately is an additional challenge. There can often be more than one acceptable output, and no tools exist that can automatically assess grammaticality or fluency. This thesis takes an approach of using probabilistic models learned from text corpora to rank candidate sentences and output the most likely. It contributes (1) a symbolic mapping rule formalism and ruleset for mapping inputs to candidate outputs that achieves broad coverage through greater regularity, (2) a packed forest representation and efficient ranking algorithm that can manage the combinatorial growth in output candidates, and (3) an empirical evaluation of coverage, correctness, and the ability to handle underspecification. This evaluation is the first large-scale empirical evaluation of coverage and quality ever performed for sentence realization. The empirical evaluation is performed by automatically converting a set of 2400 hand-parsed sentences from the Penn Treebank corpus into system inputs, and then regenerating them using the system. The top-ranked output of the generator is compared to the original sentence. The results show better than 80% coverage of newspaper text and 94% precision (57% are exact matches) for almost fully-specified inputs, and the same coverage with 55% precision for minimally specified inputs.
The linguistic annotation of natural language corpora is one of the main areas of computational linguistics. Much energy has been devoted to building large syntactically annotated corpora, which are also called treebanks after the phrase-structure trees they contain. For some time now, functional information, i.e. information on whether a constituent functions as e. g. subject, object or adverbial, has also been included in the annotation. Yet, while at first glance this may not seem to be a venture too complicated, matters are not always as easy as they seem.
The paper deals with a special kind of comparative without an overt secundum comparationis, as exemplified by, say, Pale su jače kiše 'stronger rains have fallen', which is freely used in Serbian; it is called, according to the grammatical tradition, absolute comparative. Attention has been drawn, while attempts were made at revealing the crucial features of the absolute comparative, to the fact that it is marked for its "amplified extension" on the scale of gradation; this shows that the comparative in the kind of use now under consideration has not lost its nature of an instrument of comparison. The grammatical structure in question has also been characterized as displaying sui generis semantic indefiniteness: the point is that the lack of the second object of comparison makes the scope of application of a given feature on the gradation scale rather fuzzy. The first part of the paper presents the distribution of the absolute comparative within the Slavonic linguistic area; the presentation is based on the existing grammars of particular Slavonic languages (however, Bulgarian and Macedonian have not been accounted for; the reason was that these languages have been "balcanized"). The evidence supplied by grammars allows us to distinguish, within the Slavonic area, two zones: the zone of marginal use of the form in question (Russian) or its limited use (Polish), and the zone of its active use, including Slovak, Czech, Sorbian, Slovenian and Serbian. The second part of the work describes the contrast between the situation in Serbian, on the one hand, and the situation in Polish, on the other: the focus is on the distinct divergence of the two languages in terms of textual distribution, frequency of occurrence and stylistic characteristics of the investigated structure. On the basis of the materials of bilateral translations of belletristic works, as well as those of the Serbian journalistic texts (as appearing in Internet), selected types of translational equivalences of the Serbian absolute comparative in Polish texts have been discussed; these are: the basic adjective in the positive, the negated antonym of the source adjective, and the construction "co + adjective in the comparative degree". In the last part of the article some selected differences concerning the use of the absolute comparative in Serbian and Polish journalistic texts have been pointed out. As shown in the course of the analysis, the Polish journalistic style tends to express sharp appraisals and distinct evaluations. This is particularly evident in isolated elements of press, such as titles, notices, advertising slogans. As a result, the absolute comparative, with its considerable degree of indefiniteness, appears to be less appropriate here. In contrary to this, the Serbian linguistic norm admits of a milder form of utterance, it admits of formulating less categorical judgments, even in journalistic style; this enhances the use of the absolute comparative which is well anchored both in the grammatical system and in linguistic awareness of the users of Serbian.
This paper describes the use of clustering at three stages within a larger research effort to identify semantic frames used in English automatically. The first of two tasks within this effort has been the identification of sets of semantically related verb senses that invoke a common semantic frame. Within this task, clustering has been used both to build sets of verb senses with the potential of invoking a common semantic frame and then to merge sets with a high degree of overlap. The paper is organized as follows: Section 2 introduces frame semantics. Section 3 outlines the methodology used to identify sets of semantically related verb senses that invoke a common semantic frame, while section 4 presents the specific clustering algorithm used within that process. Section 5 discusses the use of this clustering algorithm for the identification of semantically related verbs in two machine-readable lexical resources: the machine-readable version of the Longman Dictionary of Contemporary English (LDOCE, 1978 edition) and WordNet, an online lexical database (http://www.cogsci.princeton.edu/-wn; version 1.7.1 has been used for the work reported here). Section 6 presents the use of clustering to merge overlapping sets of verb senses formed in previous steps. Section 7 discusses the results of these clusterings, paying particular attention to the effect of LDOCE's restricted defining vocabulary on the clustering process.
This article examines aspects of independence and integrity as a potential measure of judicial performance for use in judicial performance evaluation programmes and examines the results of a national survey of barristers and judicial officers conducted by the author. These aspects include identifying measures of judicial independence and integrity, whether ratings of independence and integrity differ between trial and appellate judges, whether judicial gender affects ratings of independence and integrity, whether judicial age or barrister experience have an effect, and the importance of judicial independence and integrity as a measure of judicial performance.
Abstract: Ontologies are becoming extremely useful tools for sophisticated software engineering. Designing applications, databases, and knowledge bases with reference to a common ontology can mean shorter development cycles, easier and faster integration with other software and content, and a more scalable product. Although ontologies are a very promising solution to some of the most pressing problems that confront software engineering, they also raise some issues and difficulties of their own. Consider, for example, the questions below: • How can a formal ontology be used effectively by those who lack extensive training in logic and mathematics? • How can an ontology be used automatically by applications (e.g. Information Retrieval and Natural Language Processing applications) that process free text? • How can we know when an ontology is complete? In this paper we will begin by describing the upperlevel ontology SUMO (Suggested Upper Merged Ontology), which has been proposed as the initial version of an eventual Standard Upper Ontology (SUO). We will then describe the popular, free, and structured WordNet lexical database. After this preliminary discussion, we will describe the methodology that we are using to align WordNet with the SUMO. We close this paper by discussing how this alignment of WordNet with SUMO will provide answers to the questions posed above. Ontologies are becoming extremely useful tools for sophisticated software engineering. Designing applications, databases, and knowledge bases with reference to a common ontology can mean shorter development cycles, easier and faster integration with other software and content, and a more scalable product. Although ontologies are a very promising solution to some of the most pressing problems that confront software engineering, they also raise some issues and difficulties of their own. Consider, for example, the questions below: • How can a formal ontology be used effectively by those who lack extensive training in logic and mathematics? • How can an ontology be used automatically by applications (e.g. Information Retrieval and Natural Language Processing applications) that process free text? • How can we know when an ontology is complete? In this paper we will begin by describing the upperlevel ontology SUMO (Suggested Upper Merged Ontology), which has been proposed as the initial version of an eventual Standard Upper Ontology (SUO). We will then describe the popular, free, and structured WordNet lexical database. After this preliminary discussion, we will describe the methodology that we are using to align WordNet with the SUMO. We close this paper by discussing how this alignment of WordNet with SUMO will provide answers to the questions posed above. keywords: natural language, ontology 1. SUMO The SUMO (Suggested Upper Merged Ontology) is an ontology that was created at Teknowledge Corporation with extensive input from the SUO mailing list, and it has been proposed as a starter document for the IEEE-sanctioned SUO Working Group [1]. The SUMO was created by merging publicly available ontological content into a single, comprehensive, and cohesive structure [2,3]. As of February 2003, the ontology contains 1000 terms and 4000 assertions. The ontology can be browsed online (http://ontology.teknowledge.com), and source files for all of the versions of the ontology can be freely downloaded (http://ontology.teknowledge.com/cgibin/cvsweb.cgi/SUO/).
XML and related W3C standards (XSLT, XML Schema, XPath, DOM, etc.) often take part in the linguistic data representation and interchange today but not as a direct way to implement efficient data manipulation software tools. The aim of the paper is to show that incorporation of XML and the standards that surround it can bring general applicability of the implemented system for various kinds of linguistic data and also easy extensibility of such systems. As a case study we present a designed and implemented system called DEB (Dictionary Editor and Browser) that is able to manage lexical data from dictionaries to lexical databases, semantic networks, and complex ontologies. The smart design also facilitates the connection to other linguistic tools such as corpus managers or morphological analysers.
In this paper, we present a modular incremental statistical model for English full parsing. Unlike other full parsing approaches in which the analysis of the sentence is a uniform process, our model separates the full parsing into shallow parsing and sentence skeleton parsing. In shallow parsing, we finish POS tagging, Base NP identification, prepositional phrase attachment and subordinate clause identification. In skeleton parsing, we use a layered feature-oriented statistical method. Modularity possesses the advantage of solving different problems in parsing with corresponding mechanisms. Feature-oriented rule is able to express the complex lingual phenomena at the key point if needed. Evaluated on Penn Treebank corpus, we obtained 89.2% precision and 89.8% recall.
Cognitive behavioral therapy (CBT) for hoarding disorder (HD) has resulted in statistically significant improvements in hoarding symptoms, but gains have been modest and most participants continue to have clinically significant symptoms at post treatment. Contingency management, an empirically-supported intervention for substance use, may be effective in overcoming barriers to effective treatment of HD, such as fluctuating motivation and insight. The objective of the current open trial was to examine the potential effectiveness of contingency management for HD in the context of a cognitive-behavioral group therapy. Twenty-two patients completing 16-week CBT groups for HD were administered monthly contingency payments based on independent evaluator-rated reductions in overall in-home clutter. Mixed effects models suggested significant reductions in hoarding symptoms as measured by the Saving Inventory-Revised (SI-R; Frost, Steketee, & Grisham, 2004) and the Clutter Image Rating Scale (CIR; Frost, Steketee, Tolin, & Renaud, 2008), with SI-R reductions resulting in a large effect size (Cohen's d =2.59) that surpassed those obtained previously in trials of CBT for HD. Mean total earning per patient was $139, and ranged from $0 to $270. These preliminary results suggest that contingency management shows promise as a cost-efficient adjunctive intervention to boost gains in CBT for HD.
One morning each of us received a phone call from Ed Hovy. Are you sitting down? he asked. He told us that as a way to combat conference overload, and to promote interaction among communities, a joint conference had been proposed to combine HLT and NAACL. A diverse oversight committee had been formed, and according to Ed, this committee had been able to agree on two people -- and only two people -- as program co-chairs, because together we represented all of the vested interests. Marti was meant to represent the standards and tastes of the NAACL and the SIGIR crowds, and Mari the speech community, and both have been working on research contracts with HLT funders. Ed told us that if either of us said no, the entire enterprise would come crashing down. There are few better ways to convince busy people to become program co-chairs. Throughout the process, Ed provided the vision for and the drive behind this conference. We salute him for making this idea a reality, and for his enthusiastic and energetic phone calls that kept everything going. This is an exciting time for research in human language technologies. After years of relative calm, the field seems suddenly to be moving by leaps and bounds. Evidence of this can be found in our conference panel on Preparing for a Surprise Language (and as embodied in the short paper Desperately Seeking Cebuano). This panel will discuss the experiences of several groups of researchers, who at the behest of DARPA, acquired and developed language resources for an entirely new language within a span of only 10 days. This experiment took place in March of 2003, and the language in question was Cebuano, a language spoken in the Philippines. Participants successfully collected a large body of lexical and textual resources and developed a range of tools, including stemmers and POS taggers. (In June, DARPA will announce a new surprise language.) The existence of a variety of language resources, combined with advances in statistical analysis and modeling techniques, is resulting in fast-paced improvements in the field. parsers can now produce syntax trees for long sentences with high accuracy and great speed. Advances are starting to be made in automated semantic analysis. Great strides are being made in the sophistication and coverage of question answering systems. Speech recognition systems have achieved suficiently high accuracy that it is now possible to do retrieval, information extraction and topic tracking on spoken documents. Large and growing collections of text and speech corpora -- and the promise of much more from the web -- have enabled many of these advances. New developments in weakly supervised and unsupervised learning algorithms are critical for taking advantage of many new data sources, and hence this was chosen as a special theme of the conference. Lexical resources such as FrameNet, WordNet, PropBank, MeSH, and the Penn TreeBank also play prominent roles in HLT advances. As a field, human language technologies research should use, as motivation and guide, an understanding of the linguistic and cognitive bases of language. The invited talk by Dr. Elissa Newport, entitled Statistical language learning: Mechanisms for language acquisition in human learners, should help enlighten the community by informing us about the latest in psycholinguistic research. We received 162 submissions for full papers, of which 37 were accepted, resulting in a highly competitive acceptance rate of 22%. For the short (late-breaking) papers track, we received 80 submissions, of which 41 were accepted (2 later withdrawn). Some of these will be presented as short talks, and others as posters. Seventeen demonstrations will be shown. We were fortunate to be able to accept 15 papers that addressed the conference theme of unsupervised and weakly supervised methods. We also encouraged papers that described techniques that cross over or combine NLP, speech and/or IR, and several of the papers demonstrate this kind of crossover. The full paper reviewing was done using a two-tier system. First, two first-tier reviewers read every paper. Then a third reviewer, known as the meta-reviewer, wrote their own review. Finally, the meta-reviewer summarized these reviews and introduced additional comments. In some cases, the meta-reviewer instigated discussion among the first-tier reviewers to work out controversial issues. The meta-reviewers also attended the program committee meeting in which all the papers were discussed and acceptances were decided. For the short papers, each short paper received at least two reviews. Those papers whose reviewers disagreed, or which received middling scores, were subsequently reviewed by a member of the program committee and the program co-chairs. Paper submission and reviewing was done online using Marti's conference reviewing software (Conga), which she updated for this conference. Marti also maintained the conference website.
Past research has found that individual differences in both attitudinal and situational variables may be associated with males’ likelihood of acquaintance rape (LAR). The present research was conducted to examine the predictive value of both attitudinal and situational factors on males’ likelihood of forcing a female acquaintance to have non-consensual sexual intercourse. In Study 1, male and female respondents (Rs) were presented with a scenario depicting a hypothetical sexual interaction between the respondent and a newly acquainted member of the opposite sex. As the encounter progressed from one sexual activity to the next, Rs made three ratings regarding their own and partner’s intent to engage in each successive activity. The scenario ended with the female refusing further activity and males’ affect ratings, adherence to attitudes conducive of rape, and LAR were measured. Males’ initial perceptions of female sexual intent (to later engage in sexual intercourse) best predicted LAR. Study 2 was conducted to examine the role of female sexual communication on perceptions of consent to sexual intercourse. Rs were presented with a scenario similar to that in Study 1, but at each stage were requested to rate the extent to which the female had consented to engage in each sexual activity. Males completed the same affect and attitudinal measures. The results of Study 2 again suggested that males’ initial perception of female consent to (later) engage in sexual intercourse best predicted LAR. The present research suggests that further investigation into the role of situational factors and males’ initial perception of sexual intent and consent in the aetiology of acquaintance rape is required.
We present a neural-network-based statistical parser, trained and tested on the Penn Treebank. The neural network is used to estimate the parameters of a generative model of left-corner parsing, and these parameters are used to search for the most probable parse. The parser's performance (88.8% F-measure) is within 1% of the best current parsers for this task, despite using a small vocabulary size (512 inputs). Crucial to this success is the neural network architecture's ability to induce a finite representation of the unbounded parse history, and the biasing of this induction in a linguistically appropriate way.
We aim at finding the minimal set of fragments that achieves maximal parse accuracy in Data Oriented Parsing (DOP). Experiments with the Penn Wall Street Journal (WSJ) treebank show that counts of almost arbitrary fragments within parse trees are important, leading to improved parse accuracy over previous models tested on this treebank. We isolate a number of dependency relations which previous models neglect but which contribute to higher accuracy. We show that the history of statistical parsing models displays a tendency towards using more and larger fragments from training data.
This article reports the outcome of a publicly funded research project titled "Redesign of the British Sign Language (BSL) Notation System with a New Font for Use in ICT," which ran from September 2000 to December 2001. The aim of the project was to redesign the British Sign Language variant of Stokoe notation (as used in the BSL/English Dictionary) for practical use in information technology systems and software, such as lexical databases, word-processing packages, and teaching and learning applications. The project�s objectives tackled design issues, not sign linguistic system-level problems. The project resulted in two new type designs for writing BSL Stokoe, united into a larger type family called William C., in memory of William C. Stokoe. I anticipate that the new type design proposals will make the provision of searchable BSL Stokoe notation in database designs a possible next step in lexicographic and other sign linguistic research projects.
Crosslinguistically vocatives are an underexplored linguistic phenomenon and in different languages they can be highly idiosyncratic and complex (Levinson, 1987, p.71). Therefore, the problem, which is discussed in this paper, is not a language-specific one, in spite of the fact that most of the languages have their own repositories for marking the role of the addressee in the communicative utterances. In our opinion this linguistic phenomenon needs its adequate treatment in HPSG because of three main reasons: 
 
 The vocative is supposed to be present on two levels: syntax and pragmatics. Therefore it needs more elaborate interpretation on the interface side, which, in HPSG, is more developed for morphology/syntax and syntax/semantics than syntax/pragmatics. Note that a challenge for the theory is the semantic weight of the vocatives with respect to the head sentence. 
 It will be useful for HPSG-oriented implementations, especially treebanks and dialogue systems. 
 On prosodic grounds the vocatives are often viewed as being 'side or extended parts' of the sentence and therefore - very close to the parenthetical constructions. From our point of view, both phenomena are pragmatic and hence, the treatment of vocative, presented here, could be generalized to cover other phenomena of pragmatic nature. 
 
 In our work the vocatives are viewed through the possibility of the integration/separation of their pragmatic, syntactic and semantic properties.
This paper investigates two elements of Maximum Entropy tagging: the use of a correction feature in the Generalised Iterative Scaling (GIS) estimation algorithm, and techniques for model smoothing. We show analytically and empirically that the correction feature, assumed to be required for the correctness of GIS, is unnecessary. We also explore the use of a Gaussian prior and a simple cutoff for smoothing. The experiments are performed with two tagsets: the standard Penn Treebank POS tagset and the larger set of lexical types from Combinatory Categorial Grammar.
Article Sprachwissen im Konflikt. Sprachliche Zweifelsfälle zwischen Linguistik und Sprachnorm. Arbeitsgruppe auf der Jahrestagung der Deutschen Gesellshaft für Sprachwissenschaft 2003 mit dem Rahmenthema Sprache. Wissen. Sprachwissenschaft, München 26.28. Februar 2003 [Languageknowledge in conflict. Borderline cases between linguistics and linguistic norms. Working group at the annual conference of the German Linguistic Society 2003 with the topic Language. Knowledge. Linguistics.] was published on March 25, 2004 in the journal Zeitschrift für germanistische Linguistik (volume 31, issue 2).
Among the most consistent findings in the warnings literature is the so-called 'familiarity effect.' Research has shown that the more familiar an individual is with a product or situation the less likely he or she is to notice, read, recall, or comply with hazard communications. The effect has been found across numerous product types and situations using various operational definitions of familiarity and measures of warning effectiveness. However, research has also shown that subjective familiarity ratings are not highly correlated with actual product experience. Thus, individuals must be capable of developing a false or exaggerated sense of familiarity. One possible source of this exaggerated familiarity is exposure to product advertising.\n\nThree experiments were conducted to investigate whether the familiarity effect can be produced from exposure to product advertising. The relationships between advertising exposure and perceived familiarity and between perceived familiarity, perceived safety and warning effectiveness were examined. Experiment 1 explored participants' attitudes and beliefs about well-known and obscure brands of household, consumer products and sought to determine how past, direct product experience influences those attitudes and beliefs. Experiments 2 and 3 examined how the number of advertising exposures and the safety-related content of advertisements influence attitudes and beliefs about the advertised products and the effectiveness of on product warnings.\n\nResults of Experiment 1 revealed that past experience can not fully explain consumers' attitudes and beliefs about household, consumer products. Experiments 2 and 3 showed that advertising influences perceived product familiarity and knowledge. While there was a trend of greater perceived safety with increased ad exposures, the effect was not significant. No effects of advertising on warning recall were found. Implications for the design of product advertisements and product packaging as well as directions for future research are discussed.
Agents seeking to discover and compose needed Web services may face knowledge sharing interoperability problems due to differing ontologies. In practice, agents may not have a global consensus ontology that will facilitate knowledge sharing and integration of required services. We investigate a method for agents to develop local consensus ontologies to aid in the communication within a multi-agent system of business-tobusiness (B2B) agents. We compare variations of syntactic and semantic similarity matching to form local consensus ontologies with and without the use of a lexical database.
This paper investigates adapting a lexicalized probabilistic context-free grammar (PCFG) to a novel domain, using maximum a posteriori (MAP) estimation. The MAP framework is general enough to include some previous model adaptation approaches, such as corpus mixing in Gildea ( Other approaches falling within this framework are more effective. In contrast to the results in Gildea ( MAP adaptation can also be based on either supervised or unsupervised adaptation data. Even when no in-domain treebank is available, unsupervised techniques provide a substantial accuracy gain over unadapted grammars, as much as nearly 5% F-measure improvement.
Abstract. This paper explores the use of initial Stochastic Context-Free Grammars (SCFG) obtained from a treebank corpus for the learning of SCFG by means of estimation algorithms. A hybrid language model is defined as a combination of a word-based n-gram, which is used to capture the local relations between words, and a category-based SCFG with a word distribution into categories, which is defined to represent the long-term relations between these categories. Experiments on the UPenn Treebank corpus are reported. These experiments have been carried out in terms of the test set perplexity and the word error rate in a speech recognition experiment. 1
The paper presents a designed and implemented tool VisDic which implements the functionality of editing and viewing different lexical resources. It was developed in the Natural Language Processing Laboratory at the Faculty of Informatics, Masaryk University. The smart design and the accent on the usage of XML standards enable to manage lexical data from dictionaries to lexical databases, semantic networks, and complex ontologies. It also facilitates the connection to other linguistic tools such as corpus managers or morphological analyzers.
Dealing with convergence in German speech islands in Russia, Brazil and the United states the article discusses the linguistic phenomena related to the notion of convergence from different vantage points including intralinguistic convergence (due to dialect-dialect contact), interlinguistic convergence (due to language-language contact), typological "convergence" (or intralinguistic change), pidginization, and cognitive processes of simplification. Most of the German speech islands are considered to be contracting - if not dying - varieties with respect to the reduction of their grammatical systems. Evidently, for a long time language contact (and sometimes variety contact) have severely increased. Linguistic norms have been weakened in terms of both norm certainty and norm loyalty thus giving way to processes similar to those common to pidgin languages. External induced changes are highly remarkable in all German speech islands. But the susceptibility for change and the ways of change are structured by systematical and typological constraints which probably turn out to be cognitive processes underlying quite "normal" linguistic change. This change is discussed as a subsequent process of "regularization" (of irregular forms), simplification (of rules) and loss of grammatical distinctions (and their compensation). The linguistic description of these interrelated processes is based on an integrated approach providing methodology from sociolinguistics, dialectology and research on language change, including the attempt to highlight the cognitive structures which furrow the line for internal simplifications under external pressure. Comparative speech island research seems to be a promising field of application for the description of the intermesh of these processes.
The primary purpose of this study was to assess the cross-cultural invariance of job performance ratings. A secondary purpose was to examine potential cross-cultural differences in correlates of performance ratings (i.e., ratee sex, age, tenure; supervisor's opportunity to observe ratee). Fast-food supervisors from Canada, South Korea, and Spain rated employees on their technical proficiency, customer service, and teamwork. Results show that these ratings demonstrate a basic level of measurement invariance, although the error variances of the ratings and pattern of construct variances and covariances were largely culture-specific. This suggests that supervisors across cultures may use and interpret the ratings similarly, but perceive differences in performance. Furthermore, age, tenure, and the supervisor's opportunity to observe the ratee were found to affect ratings differently across cultures. Overall, this study suggests that although job performance ratings are at least partially invariant across cultures, latent performance may not be, and we present some preliminary data as to why latent invariance may not exist.
The paper deals with DEB - Dictionary Editor and Browser - a new client-server system, that allows fast development of lexical databases and general ontologies. The architecture is based on XML and related W3C standards (XSLT, XML Schema, XPath, DOM, etc.). The main feature which brings the efficiency of retrieval is the extension of a standard XSLT processor with the ability to obtain additional data from the dictionary server through the mechanism of nested queries.
We have developed an example-based machine translation (EBMT) system that uses the World Wide Web for two different purposes: First, we populate the system's memory with translations gathered from rule-based MT systems located on the Web. The source strings input to these systems were extracted automatically from an extremely small subset of the rule types in the Penn-II Treebank. In subsequent stages, the source, target translation pairs obtained are automatically transformed into a series of resources that render the translation process more successful. Despite the fact that the output from on-line MT systems is often faulty, we demonstrate in a number of experiments that when used to seed the memories of an EBMT system, they can in fact prove useful in generating translations of high quality in a robust fashion. In addition, we demonstrate the relative gain of EBMT in comparison to on-line systems. Second, despite the perception that the documents available on the Web are of questionable quality, we demonstrate in contrast that such resources are extremely useful in automatically postediting translation candidates proposed by our system.
The purpose of present study was to investigate the effect of school and classroom images on adjustment to the school among children. In Study 1, two scales were constructed to assess school and classroom images in elementary school children. In the school image scale, factor analysis yielded 4 factors: positive exterior, negative exterior, dominance, and safeguard. In the classroom image scale, factor analysis yielded 4 factors: relief, crowdedness, dominance, and cheerfulness. These findings suggest that schools and classrooms allow children to project their thought, feeling, conflicts, and frames of mind. In Study 2, Multiple-regression analyses were performed on the variables of school and classroom image ratings, using school and classroom image ratings as independent variables and school moral test ratings as dependent variables. The results shows that feelings of being dominated have an effect on children adjustment to the school.
This paper presents new methods for extracting semantic knowledge from collections of annotated images. The proposed methods include novel automatic techniques for extracting semantic concepts by disambiguating the senses of words in annotations using the lexical database WordNet, using both the images and their annotations, and for discovering semantic relations among the detected concepts based on WordNet. Another contribution of this paper is the evaluation of several techniques for visual feature descriptor extraction and data clustering in the extraction of semantic concepts. Experiments show the potential of integrating the analysis of both images and annotations for improving the performance of the word-sense disambiguation process. In particular, the accuracy improves 4-15% with respect to the baselines systems for nature images.