Papers reviewed and determined not to be word norm studies. Use the flag icon to report errors or suggest re-inclusion.
16504 papers
This article examines aspects of independence and integrity as a potential measure of judicial performance for use in judicial performance evaluation programmes and examines the results of a national survey of barristers and judicial officers conducted by the author. These aspects include identifying measures of judicial independence and integrity, whether ratings of independence and integrity differ between trial and appellate judges, whether judicial gender affects ratings of independence and integrity, whether judicial age or barrister experience have an effect, and the importance of judicial independence and integrity as a measure of judicial performance.
Acronyms are a very dynamic area of the lexicon of many languages. A hybrid, modular methodology for the acquisition of acronyms is presented, which uses an existing acronym-expansion matching component, and machine learning in two separate phases for the identification of long-distance acronym definition patterns.The resulting system, using Support Vector Machines (SVM) is trained on 600 news stories from the Wall Street Journal component of the Penn Treebank corpus using a number of lexical, syntactic, and acronym-expansion matching features. Statistical cooccurrence information for acronym-expansion pairs is extracted from search engine hit counts.The system achieves Fβ=1=92.38% on 400 news stories from the same source and has good asymptotic efficiency, making it adequate for the automatic extraction of acronyms even from noisy sources, such as newspaper text.
The paper presents a designed and implemented tool VisDic which implements the functionality of editing and viewing different lexical resources. It was developed in the Natural Language Processing Laboratory at the Faculty of Informatics, Masaryk University. The smart design and the accent on the usage of XML standards enable to manage lexical data from dictionaries to lexical databases, semantic networks, and complex ontologies. It also facilitates the connection to other linguistic tools such as corpus managers or morphological analyzers.
This article reports the outcome of a publicly funded research project titled "Redesign of the British Sign Language (BSL) Notation System with a New Font for Use in ICT," which ran from September 2000 to December 2001. The aim of the project was to redesign the British Sign Language variant of Stokoe notation (as used in the BSL/English Dictionary) for practical use in information technology systems and software, such as lexical databases, word-processing packages, and teaching and learning applications. The project�s objectives tackled design issues, not sign linguistic system-level problems. The project resulted in two new type designs for writing BSL Stokoe, united into a larger type family called William C., in memory of William C. Stokoe. I anticipate that the new type design proposals will make the provision of searchable BSL Stokoe notation in database designs a possible next step in lexicographic and other sign linguistic research projects.
Article Sprachwissen im Konflikt. Sprachliche Zweifelsfälle zwischen Linguistik und Sprachnorm. Arbeitsgruppe auf der Jahrestagung der Deutschen Gesellshaft für Sprachwissenschaft 2003 mit dem Rahmenthema Sprache. Wissen. Sprachwissenschaft, München 26.28. Februar 2003 [Languageknowledge in conflict. Borderline cases between linguistics and linguistic norms. Working group at the annual conference of the German Linguistic Society 2003 with the topic Language. Knowledge. Linguistics.] was published on March 25, 2004 in the journal Zeitschrift für germanistische Linguistik (volume 31, issue 2).
Crosslinguistically vocatives are an underexplored linguistic phenomenon and in different languages they can be highly idiosyncratic and complex (Levinson, 1987, p.71). Therefore, the problem, which is discussed in this paper, is not a language-specific one, in spite of the fact that most of the languages have their own repositories for marking the role of the addressee in the communicative utterances. In our opinion this linguistic phenomenon needs its adequate treatment in HPSG because of three main reasons: 
 
 The vocative is supposed to be present on two levels: syntax and pragmatics. Therefore it needs more elaborate interpretation on the interface side, which, in HPSG, is more developed for morphology/syntax and syntax/semantics than syntax/pragmatics. Note that a challenge for the theory is the semantic weight of the vocatives with respect to the head sentence. 
 It will be useful for HPSG-oriented implementations, especially treebanks and dialogue systems. 
 On prosodic grounds the vocatives are often viewed as being 'side or extended parts' of the sentence and therefore - very close to the parenthetical constructions. From our point of view, both phenomena are pragmatic and hence, the treatment of vocative, presented here, could be generalized to cover other phenomena of pragmatic nature. 
 
 In our work the vocatives are viewed through the possibility of the integration/separation of their pragmatic, syntactic and semantic properties.
In this paper, we will present a publication system in which selectedmaterial from letter collections is presented as dialogues between twopersons.
Presents an errata to an original article by N. Shibahara and T. V. Kondo entitled 'Variables affecting naming for Japanese Kanji: A reanalysis of Yamazaki, et al. (1997)' which appeared in Perceptual and Motor Skills, (2002), 95, pp. 741-745. The first footnote is incomplete. The complete information is provided. (The following abstract of this article originally appeared in record [rid]2003-04086-006[/rid].) M. Yamazaki, A. W. Ellis, C. M. Morrison, and M. A. L. Lambon Ralph in 1997 (see record [rid]1997-05966-008[/rid]) demonstrated that written and spoken age-of-acquisitions had a stronger effect on the naming latency of single Kanji words than any other variable including familiarity. The present study was designed to reanalyze M. Yamazaki, et al.'s data, using the ratings of written and spoken age-of-acquisitions and visual and auditory familiarities taken from the Nippon Telephone and Telegram Corporation lexical database. This analysis showed that visual familiarity exerted a stronger independent effect on naming latency than two types of age-of-acquisitions. (PsycINFO Database Record (c) 2016 APA, all rights reserved)
The paper deals with a special kind of comparative without an overt secundum comparationis, as exemplified by, say, Pale su jače kiše 'stronger rains have fallen', which is freely used in Serbian; it is called, according to the grammatical tradition, absolute comparative. Attention has been drawn, while attempts were made at revealing the crucial features of the absolute comparative, to the fact that it is marked for its "amplified extension" on the scale of gradation; this shows that the comparative in the kind of use now under consideration has not lost its nature of an instrument of comparison. The grammatical structure in question has also been characterized as displaying sui generis semantic indefiniteness: the point is that the lack of the second object of comparison makes the scope of application of a given feature on the gradation scale rather fuzzy. The first part of the paper presents the distribution of the absolute comparative within the Slavonic linguistic area; the presentation is based on the existing grammars of particular Slavonic languages (however, Bulgarian and Macedonian have not been accounted for; the reason was that these languages have been "balcanized"). The evidence supplied by grammars allows us to distinguish, within the Slavonic area, two zones: the zone of marginal use of the form in question (Russian) or its limited use (Polish), and the zone of its active use, including Slovak, Czech, Sorbian, Slovenian and Serbian. The second part of the work describes the contrast between the situation in Serbian, on the one hand, and the situation in Polish, on the other: the focus is on the distinct divergence of the two languages in terms of textual distribution, frequency of occurrence and stylistic characteristics of the investigated structure. On the basis of the materials of bilateral translations of belletristic works, as well as those of the Serbian journalistic texts (as appearing in Internet), selected types of translational equivalences of the Serbian absolute comparative in Polish texts have been discussed; these are: the basic adjective in the positive, the negated antonym of the source adjective, and the construction "co + adjective in the comparative degree". In the last part of the article some selected differences concerning the use of the absolute comparative in Serbian and Polish journalistic texts have been pointed out. As shown in the course of the analysis, the Polish journalistic style tends to express sharp appraisals and distinct evaluations. This is particularly evident in isolated elements of press, such as titles, notices, advertising slogans. As a result, the absolute comparative, with its considerable degree of indefiniteness, appears to be less appropriate here. In contrary to this, the Serbian linguistic norm admits of a milder form of utterance, it admits of formulating less categorical judgments, even in journalistic style; this enhances the use of the absolute comparative which is well anchored both in the grammatical system and in linguistic awareness of the users of Serbian.
In this paper we will present work carried out lately on the 50,000 words Italian Spontaneous Speech Corpus called AVIP, under national project API, made available for free download from the website of the coordinator, the University of Naples. We will concentrate on the tuning of the parser for Italian which had been previously used to parse 100,000 words corpus of written Italian within the National Treebank initiative coordinated by ILC in Pisa. The parser receives as input the adequately transformed orthographic transcription of the dialogues making up the corpus, in which pauses, hesitations and other disfluencies have been turned into most likely corresponding punctiation marks, interjections or truncation of the word underlying the uttered segment.\nThe most interesting phenomenon we will discuss is without any doubts "overlapping", i.e. a speech event in which two people speak at the same time by uttering actual words or in some cases nonwords, when one of the speakers, usually the one which is not the current turntaker, interrupts the current speaker.\nThis phenomenon takes place at a certain point in time where it has to be anchored to the speech signal but in order to be fully parsed and subsequently semantically interpreted, it needs to be referred semantically to a following turn.
One morning each of us received a phone call from Ed Hovy. Are you sitting down? he asked. He told us that as a way to combat conference overload, and to promote interaction among communities, a joint conference had been proposed to combine HLT and NAACL. A diverse oversight committee had been formed, and according to Ed, this committee had been able to agree on two people -- and only two people -- as program co-chairs, because together we represented all of the vested interests. Marti was meant to represent the standards and tastes of the NAACL and the SIGIR crowds, and Mari the speech community, and both have been working on research contracts with HLT funders. Ed told us that if either of us said no, the entire enterprise would come crashing down. There are few better ways to convince busy people to become program co-chairs. Throughout the process, Ed provided the vision for and the drive behind this conference. We salute him for making this idea a reality, and for his enthusiastic and energetic phone calls that kept everything going. This is an exciting time for research in human language technologies. After years of relative calm, the field seems suddenly to be moving by leaps and bounds. Evidence of this can be found in our conference panel on Preparing for a Surprise Language (and as embodied in the short paper Desperately Seeking Cebuano). This panel will discuss the experiences of several groups of researchers, who at the behest of DARPA, acquired and developed language resources for an entirely new language within a span of only 10 days. This experiment took place in March of 2003, and the language in question was Cebuano, a language spoken in the Philippines. Participants successfully collected a large body of lexical and textual resources and developed a range of tools, including stemmers and POS taggers. (In June, DARPA will announce a new surprise language.) The existence of a variety of language resources, combined with advances in statistical analysis and modeling techniques, is resulting in fast-paced improvements in the field. parsers can now produce syntax trees for long sentences with high accuracy and great speed. Advances are starting to be made in automated semantic analysis. Great strides are being made in the sophistication and coverage of question answering systems. Speech recognition systems have achieved suficiently high accuracy that it is now possible to do retrieval, information extraction and topic tracking on spoken documents. Large and growing collections of text and speech corpora -- and the promise of much more from the web -- have enabled many of these advances. New developments in weakly supervised and unsupervised learning algorithms are critical for taking advantage of many new data sources, and hence this was chosen as a special theme of the conference. Lexical resources such as FrameNet, WordNet, PropBank, MeSH, and the Penn TreeBank also play prominent roles in HLT advances. As a field, human language technologies research should use, as motivation and guide, an understanding of the linguistic and cognitive bases of language. The invited talk by Dr. Elissa Newport, entitled Statistical language learning: Mechanisms for language acquisition in human learners, should help enlighten the community by informing us about the latest in psycholinguistic research. We received 162 submissions for full papers, of which 37 were accepted, resulting in a highly competitive acceptance rate of 22%. For the short (late-breaking) papers track, we received 80 submissions, of which 41 were accepted (2 later withdrawn). Some of these will be presented as short talks, and others as posters. Seventeen demonstrations will be shown. We were fortunate to be able to accept 15 papers that addressed the conference theme of unsupervised and weakly supervised methods. We also encouraged papers that described techniques that cross over or combine NLP, speech and/or IR, and several of the papers demonstrate this kind of crossover. The full paper reviewing was done using a two-tier system. First, two first-tier reviewers read every paper. Then a third reviewer, known as the meta-reviewer, wrote their own review. Finally, the meta-reviewer summarized these reviews and introduced additional comments. In some cases, the meta-reviewer instigated discussion among the first-tier reviewers to work out controversial issues. The meta-reviewers also attended the program committee meeting in which all the papers were discussed and acceptances were decided. For the short papers, each short paper received at least two reviews. Those papers whose reviewers disagreed, or which received middling scores, were subsequently reviewed by a member of the program committee and the program co-chairs. Paper submission and reviewing was done online using Marti's conference reviewing software (Conga), which she updated for this conference. Marti also maintained the conference website.
Machine translation engines draw on various types of databases. This paper is concerned with Arabic as a source or target language, and focuses on lexical databases. The non-concatenative nature of Arabic morphology, the complex structure of Arabic word-forms, and the general use of vowel-free writing present a real challenge to NLP developers. We show here how and why a stem-grounded lexical database, the items of which are associated with grammar-lexis specifications – as opposed to a root-&-pattern database –, is motivated both linguistically and with regards to efficiency, economy and modularity. Arguments in favour of databases relying on stems associated with grammar-lexis specifications (such as DIINAR.1 or the Arabic dB under development at SYSTRAN), rather than on roots and patterns, are the following: (a) The latter include huge numbers of rule-generated word-forms, which do not actually appear in the language. (b) Rule-generated lemmas – as opposed to existing ones – are widely under-specified with regards to grammar-lexis relations. (c) In a Semitic language such as Arabic, the mapping of grammar-lexis specifications that need to be associated with every lexical entry of the database is decisive. (d) These specifications can only be included in a stem-based dB. Points (a) to (d) are crucial and in the context of machine translation involving Arabic.
Cognitive behavioral therapy (CBT) for hoarding disorder (HD) has resulted in statistically significant improvements in hoarding symptoms, but gains have been modest and most participants continue to have clinically significant symptoms at post treatment. Contingency management, an empirically-supported intervention for substance use, may be effective in overcoming barriers to effective treatment of HD, such as fluctuating motivation and insight. The objective of the current open trial was to examine the potential effectiveness of contingency management for HD in the context of a cognitive-behavioral group therapy. Twenty-two patients completing 16-week CBT groups for HD were administered monthly contingency payments based on independent evaluator-rated reductions in overall in-home clutter. Mixed effects models suggested significant reductions in hoarding symptoms as measured by the Saving Inventory-Revised (SI-R; Frost, Steketee, & Grisham, 2004) and the Clutter Image Rating Scale (CIR; Frost, Steketee, Tolin, & Renaud, 2008), with SI-R reductions resulting in a large effect size (Cohen's d =2.59) that surpassed those obtained previously in trials of CBT for HD. Mean total earning per patient was $139, and ranged from $0 to $270. These preliminary results suggest that contingency management shows promise as a cost-efficient adjunctive intervention to boost gains in CBT for HD.
Codeswitching is an important socio-cultural phenomenon in many multilingual and bilingual communities. In Malaysia, where there is a diverse linguistic panorama of speech communities, codeswitching is the norm. Malaysian speakers of English often codeswitch between Malay and English. The use of native words and expressions in Malaysian English is not only restricted to spoken English, but also found in written English. This article, based on a small-scale corpus study of Malaysian English, attempts to describe the use of Malay lexical items in Malaysian English by analysing the occurrence of these native items in texts written by Malaysian speakers of English. It also aims to explain the motivation behind their usage. The study suggests that Malaysian speakers of English use native lexical items in English texts as an intentional codeswitching strategy, motivated by the effects that the use of the native items has on the interpretation of the message.
OBJECTIVE: To compare the subjective image quality of the newer generation Schick CDR detector employing complementary metal oxide semiconductor (CMOS) technology with images using the earlier generation charge-coupled device (CCD) Schick CDR detector. METHODS: All radiographic images were made using the same formalin-fixed adult cadaver maxilla with surrounding natural soft tissues in place. The X-ray generator used was a Villa Sistemi Medicali Diamatic srl AP/Explor X operated at 70 kVp and 8 mA. The source-to-detector distance was set at 38 cm and an optical bench was used to ensure reproducible beam geometry. A range of exposures was applied for both detectors. A panel of nine dentists independently observed and evaluated images made at each exposure. For both detectors, the three images ranked highest were randomized for re-evaluation in panels of six images. Each image was repeated randomly a total of 10 times. Features chosen as observation points were: (1) proximal dental caries; (2) gingival soft tissues; (3) cortical bone; (4) root canal space; (5) root apices; (6) periodontal ligament space; and (7) endodontic instrument tip clarity. Comparisons were made by use of odds ratio analysis applying a 95% confidence level. Interrater and intrarater reliabilities were computed to assess consistency in observer ratings. RESULTS: The CMOS sensor was rated as outperforming its CCD predecessor for depiction of cortical bone and root apices; the CCD detector was only rated superior for depiction of root canal space. No significant difference was found between the two detectors in perceived depiction of proximal dental caries, gingival soft tissues, periodontal ligament space or endodontic instruments. Combining rating scores from each of the tasks, CMOS and CCD detectors had a similar proportion of image ratings of excellent, acceptable and poor. CONCLUSIONS: Regarding subjective image quality, the Schick CMOS and CCD detectors were perceived to produce radiographic images of similar overall quality.
<p>The OmniPaper project has implemented three information retrieval prototypes in the area of electronic news publishing. One prototype uses SOAP as communication protocol between the central system and a number of distributed news archives. The second prototype uses an RDF metadata database, enabling direct metadata queries to the central system. Finally the Topic Map prototype uses query expansion and semantic linking for smart metadata search. The Topic Map prototype enhances thesearch experience by implementing a knowledge layer that combines the semantic content of a lexical database, consisting of concepts and keywords, with a metadata-set of newspaper articles. The linking between both is currently implemented at the level of keywords but will be developed at the level of concepts in the final prototype. The knowledge layer has been designed from a Topic Map point of view, although the XTM syntax has not been used to avoid performance issues. The consortium’s adopted view on information publishing and retrieval considers querying and navigation as two very related actions that can both be captured under the name “search for relevant information”. Navigation forces the user to followpredefined paths whereas querying enables the user to look freely for a suitable starting point. The query and navigation functionality is provided through a web engine and is build on top of the information structure of the knowledge layer.</p>
In den letzten Jahren ist die Zahl der verfgbaren linguistisch annotierten Korpora stndig gewachsen. Zu den bekanntesten gehren das Brown-Korpus, das Susanne-Korpus, die Penn-Treebank, das Negra-Korpus, das Tiger-Korpus und die im Zusam-
The linguistic annotation of natural language corpora is one of the main areas of computational linguistics. Much energy has been devoted to building large syntactically annotated corpora, which are also called treebanks after the phrase-structure trees they contain. For some time now, functional information, i.e. information on whether a constituent functions as e. g. subject, object or adverbial, has also been included in the annotation. Yet, while at first glance this may not seem to be a venture too complicated, matters are not always as easy as they seem.
Code switching is a practice constrained by grammatical principles and shaped by environmental, social and personal influences (Milroy and Wei 1995). There are several factors crucial to understanding of code switching like the community in which it takes place or mode of the bilingual speaker. Some communities accept code switching within a single context as the norm for communicative interactions whereas others maintain a strict distinction between the languages (Heller 1995). It is thus imperative to study code switching in a proper linguistic and cultural context. Language mode is an important factor to be considered in any study on bilingual aphasia. Language mode is the state of activation of bilingual’s languages and language processing mechanisms at a given time (Grosjean 2000). A bilingual can be on a continuum depending on the situation he is in. At one end they may be in monolingual mode where there would be ideally no mixing and at the other end they find themselves in a bilingual mode mixing languages freely (Grosjean 1982). The movement of a bilingual along the continuum results in varying language behaviors. Earlier research paid little attention to language mode but as Grosjean (2000) highlights, it needs to be controlled in any bilingual experiment by evaluating monolingual and bilingual modes on different days with different interlocutors. In the present study an attempt has been made to do so. Bilingual aphasic speakers like normal bilinguals, need to alternate and use context appropriate languages. Sometimes the deficit in linguistic competence may affect this ability to alternate the linguistic codes (Munoz, Marquardt and Copeland 1998). Bilingual aphasics have been seen to combine languages in a variety of ways. They may use several languages together in same utterance (Gloning & Gloning, 1965; Mosner & Pilsch, 1971) or produce the correct name of an object in an unsolicited language (Gloning & Gloning, 1965; Weisenburg & Mcbride, 1935;) even when it is impossible for the same patient to produce the correct name in that language upon request. Junque, Vendrell, Vendrell-Beret and Tobena (1989) and Paradis (1995) suggest that mixing of languages is frequently observed recovery pattern among bilingual aphasics. One of the earliest detailed reports on language mixing was by Perecman (1984) of a 80-year-old male who suffered extensive bilateral temporal hematomes resulting from a car accident. Data was analyzed for different levels (phonological, morphological, lexical-semantic and syntactic) of code switching. She concluded that language boundaries are poorly delineated in polyglot aphasic’s mental grammar and remarked that utterance level mixing and spontaneous translation are abnormal behaviors seen in bilingual aphasics. Grosjean (1985) contradicts these findings by specifying that utterance level mixing is not unique to bilingual aphasics as suggested by Perecman (1984). He pointed out that the interlocutor in the above study was a multilingual who mixed languages and this in turn could have triggered language mixing in subject as a communicative strategy. He identified factors such as language mode, pre morbid language use and test constraints as strategic in any study dealing with language mixing. Hyltenstam (1995) analyzed samples of language mixing from 31 cases of bilingual aphasia reported in literature using Poplack’s syntactic constraints and the MLF (matrix language frame, Myers-Scotton 1993) model. He found that it is reasonable to believe that the code switching of aphasic speakers is structured according to same conversational constraints as in normal speakers. Munoz, Marquardt and Copeland (1998) pointed to methodological shortfalls that comprised data interpretation such as little information about pre morbid language use, presence of bilingual interlocutors, limited samples and lack of controls. In order to overcome these, Munoz, Marquardt and Copeland (1998), compared the code switching patterns of aphasic and neurologically normal bilingual speakers of English and Spanish using Matrix language frame (MLF) model. Communicative difficulties resulting from code switching with
En este artículo estudiamos el problema de la estimación de gramáticas incontextuales \nestocásticas en formato general y su uso en un modelo de lenguaje híbrido. \nEn este trabajo se propone la estimación de una gramática incontextual estocástica usando \nuna nueva versión del algoritmo de Earley que permite manejar muestras parentizadas. El \nmodelo de lenguaje híbrido es definido como una combinación lineal de un modelo de ngramas \nbasado en palabras, que se utiliza para capturar las relaciones locales entre palabras, \ny una gramática estocástica, basada en categorías junto con una distribución de palabras en \ncategorías, que se utiliza para representar las relaciones a largo término entre estas categorías. Se han realizado experimentos usando el corpus UPenn Treebank. La evaluación \nde los modelos se ha realizado desde el punto de vista de la perplejidad de un conjunto \nde test, y desde el punto de vista de la tasa de errores por palabra en un experimento de \nreconocimiento automático del habla.
One morning each of us received a phone call from Ed Hovy. Are you sitting down? he asked. He told us that as a way to combat conference overload, and to promote interaction among communities, a joint conference had been proposed to combine HLT and NAACL. A diverse oversight committee had been formed, and according to Ed, this committee had been able to agree on two people -- and only two people -- as program co-chairs, because together we represented all of the vested interests. Marti was meant to represent the standards and tastes of the NAACL and the SIGIR crowds, and Mari the speech community, and both have been working on research contracts with HLT funders. Ed told us that if either of us said no, the entire enterprise would come crashing down. There are few better ways to convince busy people to become program co-chairs. Throughout the process, Ed provided the vision for and the drive behind this conference. We salute him for making this idea a reality, and for his enthusiastic and energetic phone calls that kept everything going. This is an exciting time for research in human language technologies. After years of relative calm, the field seems suddenly to be moving by leaps and bounds. Evidence of this can be found in our conference panel on Preparing for a Surprise Language (and as embodied in the short paper Desperately Seeking Cebuano). This panel will discuss the experiences of several groups of researchers, who at the behest of DARPA, acquired and developed language resources for an entirely new language within a span of only 10 days. This experiment took place in March of 2003, and the language in question was Cebuano, a language spoken in the Philippines. Participants successfully collected a large body of lexical and textual resources and developed a range of tools, including stemmers and POS taggers. (In June, DARPA will announce a new surprise language.) The existence of a variety of language resources, combined with advances in statistical analysis and modeling techniques, is resulting in fast-paced improvements in the field. parsers can now produce syntax trees for long sentences with high accuracy and great speed. Advances are starting to be made in automated semantic analysis. Great strides are being made in the sophistication and coverage of question answering systems. Speech recognition systems have achieved suficiently high accuracy that it is now possible to do retrieval, information extraction and topic tracking on spoken documents. Large and growing collections of text and speech corpora -- and the promise of much more from the web -- have enabled many of these advances. New developments in weakly supervised and unsupervised learning algorithms are critical for taking advantage of many new data sources, and hence this was chosen as a special theme of the conference. Lexical resources such as FrameNet, WordNet, PropBank, MeSH, and the Penn TreeBank also play prominent roles in HLT advances. As a field, human language technologies research should use, as motivation and guide, an understanding of the linguistic and cognitive bases of language. The invited talk by Dr. Elissa Newport, entitled Statistical language learning: Mechanisms for language acquisition in human learners, should help enlighten the community by informing us about the latest in psycholinguistic research. We received 162 submissions for full papers, of which 37 were accepted, resulting in a highly competitive acceptance rate of 22%. For the short (late-breaking) papers track, we received 80 submissions, of which 41 were accepted (2 later withdrawn). Some of these will be presented as short talks, and others as posters. Seventeen demonstrations will be shown. We were fortunate to be able to accept 15 papers that addressed the conference theme of unsupervised and weakly supervised methods. We also encouraged papers that described techniques that cross over or combine NLP, speech and/or IR, and several of the papers demonstrate this kind of crossover. The full paper reviewing was done using a two-tier system. First, two first-tier reviewers read every paper. Then a third reviewer, known as the meta-reviewer, wrote their own review. Finally, the meta-reviewer summarized these reviews and introduced additional comments. In some cases, the meta-reviewer instigated discussion among the first-tier reviewers to work out controversial issues. The meta-reviewers also attended the program committee meeting in which all the papers were discussed and acceptances were decided. For the short papers, each short paper received at least two reviews. Those papers whose reviewers disagreed, or which received middling scores, were subsequently reviewed by a member of the program committee and the program co-chairs. Paper submission and reviewing was done online using Marti's conference reviewing software (Conga), which she updated for this conference. Marti also maintained the conference website.
Abstract This paper describes a framework for building story traces (compact global views of a narrative) and story projections (selections of key elements of a narrative) and their applications in text understanding and classification. Word and sense properties are extracted from documents using the WordNet lexical database enhanced with Prolog inference rules and a number of lexical transformations. Inference rules are based on navigation in various WordNet relation chains (hypernyms, meronyms, entailment and causality links, etc.) and derived relations expressed as Boolean combinations of node and edge properties used to direct the navigation. The resulting abstract story traces provide a compact view of the underlying narrative's key content elements and a means for automated indexing and classification of text collections. Ontology driven projections act as a kind of “semantic lenses” and provide a means to select a subset of a narrative whose key sense elements are subsumed by a set of concepts, predicates and properties expressing the focus of interest of a user. Finally, we discuss applications of these techniques in text understanding, classification of text collections and answering questions about a text.
This paper introduces a novel Support Vector Machines (SVMs) based voting algorithm for reranking, which provides a way to solve the sequential models indirectly. We have presented a risk formulation under the PAC framework for this voting algorithm. We have applied this algorithm to the parse reranking problem, and achieved labeled recall and precision of 89.4%/89.8% on WSJ section 23 of Penn Treebank.
Some studies with children have shown that there is no semantic priming at short stimulus onset asynchrony (SOA) in lexical decision and naming tasks for homographs. The predictions of spreading activation theories might explain this missing effect. There may be differences in children's and adults' memory structures. We have explored this hypothesis. The development of memory structure representations for homographs was measured by a Pathfinder algorithm. In Experiment 1, the three dependent variables were: the number of links in the network, closeness measures (C), and distances between nodes. Results revealed developmental differences in network structure representations in adults and children. In Experiment 2, results revealed that these differences were not due to the cohort effect. In Experiment 3, the relationship between associative strength, as measured by associative norms, and distances, as measured by Pathfinder algorithm, was explored. The results of these three experiments and empirical research from semantic priming experiments show that these differences in memory structure representations could be one of the sources of the missing semantic priming effect in children.
This paper deals with the problem of how to interrelate theory-specific treebanks and how to transform one treebank format to another. Currently, two approaches to achieve these goals can be differentiated. The first creates a mapping algorithm between treebank formats. Categories of a source format are transformed into a target format via a given set of general or language-specific mapping rules. The second relates treebanks via a transformation to a general model of linguistic categories, for example based on the EAGLES recommendations for syntactic annotations of corpora, or relying on the HPSG framework. This paper proposes a new methodology as a solution for these desiderata.
Perceptions of social closeness and familiarity were assessed among 44 monozygotic (MZA) and 33 dizygotic (DZA) reunited twin pairs, and several individual twins and triplets. Significantly greater MZA than DZA closeness and familiarity were found. Closeness and familiarity ratings for co-twins exceeded those for nonbiological siblings with whom twins were raised. Correlations between perceptions of physical resemblance and social closeness and familiarity were positive and statistically significant. However, most correlations between social relatedness and contact time were non-significant. Associations between social relatedness and similarities in selected behavioral traits were also examined. The findings support various theoretical perspectives anticipating greater affiliation among close relatives than distant relatives.
We investigate the performance of the Structured Language Model when one of its components is modeled by a connectionist model. Using a connectionist model and a distributed representation of the items in the history makes the component able to use much longer contexts than possible with currently used interpolated or backoff models, both because of the inherent capability of the connectionist model to fight the data sparseness problem, and because of the only sub-linear growth in the model size when increasing the context length. Experiments show significant improvement in perplexity and moderate reduction in word error rate over the baseline SLM results on the UPENN treebank and Wall Street Journal (WSJ) corpora respectively. The results also show that the probability distribution obtained by our model is much less correlated to regular N-grams than the baseline SLM model.
This paper investigates adapting a lexicalized probabilistic context-free grammar (PCFG) to a novel domain, using maximum a posteriori (MAP) estimation. The MAP framework is general enough to include some previous model adaptation approaches, such as corpus mixing in Gildea ( Other approaches falling within this framework are more effective. In contrast to the results in Gildea ( MAP adaptation can also be based on either supervised or unsupervised adaptation data. Even when no in-domain treebank is available, unsupervised techniques provide a substantial accuracy gain over unadapted grammars, as much as nearly 5% F-measure improvement.
The article is devoted to the problem of regulation and unification of law terminology in modern Ukrainian \nlegislation. Analyzing with this purpose the text of the Criminal code of Ukraine, author pays attention to cases \nof non-compliance with lexical, grammar, stylistic norms, replication of Russian syntactical constructions in \nseveral items of Code. Author also gives recommendations of the correct usage of terms and words.
We present a neural network method for inducing representations of parse histories and using these history representations to estimate the probabilities needed by a statistical left-corner parser. The resulting statistical parser achieves performance (89.1% F-measure) on the Penn Treebank which is only 0.6% below the best current parser for this task, despite using a smaller vocabulary size and less prior linguistic knowledge. Crucial to this success is the use of structurally determined soft biases in inducing the representation of the parse history, and no use of hard independence assumptions.
The sphere of language has become a privileged domain in which to interrogate the causes and effects of social injustice. (1) What is the right language to resist rape? Why is the language women use during rape frequently considered to be the wrong language? Such questions have a particular urgency in the context of women who have been raped by a man that they know. Why is their language the subject of particular scrutiny when they come before the law? How is it that the things these women say during rape can be used to turn violence into consensual sex? What relation between rape and language subtends the possibility of this transformation? Sharon Marcus, in the most influential reflection on the relation between rape and language, characterises feminist engagements with this question in the following manner: Whose words count in a rape trial? Whose 'no' can ever mean 'no'? How do rape trials condone men's misinterpretations of women's words? How do rape trials consolidate men's subjective accounts in objective 'norms of truth' and deprive women's subject accounts of cognitive value? Feminists have also insisted on the importance of naming rape as violence and of collectively narrating stories of rape. (2) This passage offers the central coordinates of the conventional feminist understanding of language in the context of rape. In particular, the primary coordinate here is words: words in general, which can count or not count, specific words like 'no' and category words which have the power to name. On this view, the problem with language in the context of rape is a problem of words, which are deprived of their power to designate the world according to women's experience. Australian criminologist Patricia Easteal articulates this position as she reflects on the relation between rape and After attending a meeting with other legal officers, she remarks that: [E]verytime... anyone present mentioned a judge or a lawyer, the pronouns 'he' and 'him' were used. I would contend that the use of language accurately reflects the reality that men continue to hold most of these positions in the criminal justice system. It goes deeper though. It mirrors the power that males have and use to maintain a stranglehold on the institutions and structures of Australian culture. And, the language in turn affects or even directs how we see the rest of our reality, including sexual assault and the law. (3) In this passage the problem of language is presented through a problem with words: words as a general category and specific grammatical categories like pronouns (4) which accumulate into the broad concept of masculine language. Here, language is understood as an aggregate of words (lexical items) which contain a self-evident signifying power. That is, words (connoting a reality) emit and effect their meanings, in a manner shorn of grammar, genre or context. (5) In other words, language is presented here in a theological formation: that is, in terms of its powers of naming. One of those things that language names (from a view) is sexual assault, and this act of naming will have a determining effect. Strategies for contesting this effect of words are not offered here, but would presumably include the use of non-gender specific pronouns and words that would reflect realities distinct to those inscribed by male power. (6) Dale Spender offers a famous solution to this linguistic-political problem when she argues that: 'Women need a word which renames male violence and misogyny and which asserts their blameless nature, a word which places the responsibility for rape where it belongs--on the dominant group'. (7) Spender's conclusion here implies that words can function as the symbols of political arguments: in this instance, new words about rape, invented by women, could (through condensation) denote already completed argumentative positions that clearly designate responsibility for sexual violence. …
This paper investigates two elements of Maximum Entropy tagging: the use of a correction feature in the Generalised Iterative Scaling (GIS) estimation algorithm, and techniques for model smoothing. We show analytically and empirically that the correction feature, assumed to be required for the correctness of GIS, is unnecessary. We also explore the use of a Gaussian prior and a simple cutoff for smoothing. The experiments are performed with two tagsets: the standard Penn Treebank POS tagset and the larger set of lexical types from Combinatory Categorial Grammar.
Abstract: Ontologies are becoming extremely useful tools for sophisticated software engineering. Designing applications, databases, and knowledge bases with reference to a common ontology can mean shorter development cycles, easier and faster integration with other software and content, and a more scalable product. Although ontologies are a very promising solution to some of the most pressing problems that confront software engineering, they also raise some issues and difficulties of their own. Consider, for example, the questions below: • How can a formal ontology be used effectively by those who lack extensive training in logic and mathematics? • How can an ontology be used automatically by applications (e.g. Information Retrieval and Natural Language Processing applications) that process free text? • How can we know when an ontology is complete? In this paper we will begin by describing the upperlevel ontology SUMO (Suggested Upper Merged Ontology), which has been proposed as the initial version of an eventual Standard Upper Ontology (SUO). We will then describe the popular, free, and structured WordNet lexical database. After this preliminary discussion, we will describe the methodology that we are using to align WordNet with the SUMO. We close this paper by discussing how this alignment of WordNet with SUMO will provide answers to the questions posed above. Ontologies are becoming extremely useful tools for sophisticated software engineering. Designing applications, databases, and knowledge bases with reference to a common ontology can mean shorter development cycles, easier and faster integration with other software and content, and a more scalable product. Although ontologies are a very promising solution to some of the most pressing problems that confront software engineering, they also raise some issues and difficulties of their own. Consider, for example, the questions below: • How can a formal ontology be used effectively by those who lack extensive training in logic and mathematics? • How can an ontology be used automatically by applications (e.g. Information Retrieval and Natural Language Processing applications) that process free text? • How can we know when an ontology is complete? In this paper we will begin by describing the upperlevel ontology SUMO (Suggested Upper Merged Ontology), which has been proposed as the initial version of an eventual Standard Upper Ontology (SUO). We will then describe the popular, free, and structured WordNet lexical database. After this preliminary discussion, we will describe the methodology that we are using to align WordNet with the SUMO. We close this paper by discussing how this alignment of WordNet with SUMO will provide answers to the questions posed above. keywords: natural language, ontology 1. SUMO The SUMO (Suggested Upper Merged Ontology) is an ontology that was created at Teknowledge Corporation with extensive input from the SUO mailing list, and it has been proposed as a starter document for the IEEE-sanctioned SUO Working Group [1]. The SUMO was created by merging publicly available ontological content into a single, comprehensive, and cohesive structure [2,3]. As of February 2003, the ontology contains 1000 terms and 4000 assertions. The ontology can be browsed online (http://ontology.teknowledge.com), and source files for all of the versions of the ontology can be freely downloaded (http://ontology.teknowledge.com/cgibin/cvsweb.cgi/SUO/).
Analyzing from angle of logics and linguistics, it can be seen that certain corresponding relations exist between norm, fact and In aspect of norm, law wears structure of the former/the while in aspect of fact it wears structure of legal fact/legal relation. Between former of norm and fact, and between latter of norm and relation, there is an respective corresponding At same time, between former lexical item and constitutive requirements of fact, and between latter lexical item and component elements of relation, a corresponding relation exists respectively. What's more, certain relation exists even between transverse logical necessary relation of former and latter, and that of fact and Such kind of correspondence is decided by presumption and regularity, characters of norms per se. Norms possess presumption or regularity, so law can be distinguished from other isomerism elements. Therefore, it's possible to forcibly kink two facts that are absolutely irrelevant in daily life.
This paper describes a fast algorithm that selects features for conditional maximum entropy modeling. Berger et al. (1996) presents an incremental feature selection (IFS) algorithm, which computes the approximate gains for all candidate features at each selection stage, and is very time-consuming for any problems with large feature spaces. In this new algorithm, instead, we only compute the approximate gains for the top-ranked features based on the models obtained from previous stages. Experiments on WSJ data in Penn Treebank are conducted to show that the new algorithm greatly speeds up the feature selection process while maintaining the same quality of selected features. One variant of this new algorithm with look-ahead functionality is also tested to further confirm the good quality of the selected features. The new algorithm is easy to implement, and given a feature space of size F, it only uses O(F) more space than the original IFS algorithm.