Papers reviewed and determined not to be word norm studies. Use the flag icon to report errors or suggest re-inclusion.
16504 papers
Spoken language proficiency is intuitively related to effective and efficient communication in spoken interactions. However, it is difficult to derive a reliable estimate of spoken language proficiency by situated elicitation and evaluation of a person’s communicative behavior. This paper describes the task structure and scoring logic of a group of fully automatic spoken language proficiency tests (for English, Spanish and Dutch) that are delivered via telephone or Internet. Test items are presented in spoken form and require a spoken response. Each test is automatically-scored and primarily based on short, decontextualized tasks that elicit integrated listening and speaking performances. The tests present several types of tasks to candidates, including sentence repetition, question answering, sentence construction, and story retelling. The spoken responses are scored according to the lexical content of the response and a set of acoustic base measures on segments, words and phrases, which are scaled with IRT methods or parametrically combined to optimize fit to human listener judgments. Most responses are isolated spoken phrases and sentences that are scored according to their linguistic content, their latency, and their fluency and pronunciation. The item development procedures and item norming are described.
In academic courses in which one task for the students is to understand empirical methodology and the nature of scientific inquiry, the ability of students to create and implement their own experiments allows them to take intellectual ownership of, and greatly facilitates, the learning process. The Psychology Experiment Authoring Kit (PEAK) is a novel spreadsheet-based interface allowing students and researchers with rudimentary spreadsheet skills to create cognitive and cognitive neuroscience experiments in minutes. Students fill in a spreadsheet listing of independent variables and stimuli, insert columns that represent experimental objects such as slides (presenting text, pictures, and sounds) and feedback displays to create complete experiments, all within a single spreadsheet. The application then executes experiments with centisecond precision. Formal usability testing was done in two stages: (1) detailed coding of 10 individual subjects in one-on-one experimenter/subject videotaped sessions and (2) classroom testing of 64 undergraduates. In both individual and classroom testing, the students learned to effectively use PEAK within 2 h, and were able to create a lexical decision experiment in under 10 min. Findings from the individual testing in Stage 1 resulted in significant changes to documentation and training materials and identification of bugs to be corrected. Stage 2 testing identified additional bugs to be corrected and new features to be considered to facilitate student understanding of the experiment model. Such testing will improve the approach with each semester. The students were typically able to create their own projects in 2 h.
Current trends in language technology require treebanks that do not stop at the level of constituent structure, but include deeper and richer levels of analysis, including appropriate meaning structures. Capturing sufficient detail at different levels of linguistic description is too complex a task to be practically achievable by manual annotation or shallow parsing; rather it requires sophisticated tools that help secure the consistency of parallel but different structures. We are constructing a multilevel treebanking tool that incorporates a deep parser and grammar for Norwegian. Thus, we are tightly linking our treebank to grammar development so as to achieve a sound embedding in grammatical theory and yield more useful results for applications.
The aim of this paper is to present a lexical database of English collocations used in scientific language, which is being built in three Spanish universities (Barcelona, Illes Balears and Leon) and is mainly intended for the Spanish-speaking scientific community. The shortage of specialized dictionaries providing contextual information on the grammatical and collocational patterns in specific registers prompted the onset of this project. Our database is based on the analysis of a corpus of written texts in the areas of biology, biochemistry, and biomedicine, and provides the grammatical, semantic, and collocational information necessary for the correct and precise use of each term in scientific discourse. The paper describes the steps followed in the creation of the data base and it includes the case study of one of its entries.
These two letters and two inventories preserved in the rich heritage of Anton Hodinka in the manuscript depository of the Hungarian Academy of Sciences Library present an exciting picture of the everyday life of the 18th century. Nevertheless, I find these documents valuable not because of this fact but due to their vocabulary which reflects the Rusyn language adequately. These original sources are the splendid illustrations of the Rusyn language wordstock used in everyday life of that period. Therefore I have not spared myself to copy, study and publish the manuscripts in question because I should like to contribute to enriching the Rusyn language history. As a matter of fact the Rusyn language of the 18th century reflects the synthesis of three elements: the Church Slavonic liturgy language, the Old Ukrainian language and the living folk language. The formation and unification of the literary language norm, which was not regulated by grammars and dictionaries, was greatly influenced by the bishop's office documents due to the great authority and prestige of the church in the region. The three above-mentioned elements of the Rusyn literary language of the 18th century can be revealed in all language layers (phonetical, morphological, syntactical, lexical, semantical). I shall give several examples on the elements of the Rusyn folk language.
The aim of this paper is to investigate a case of transfer within the context of language death. By examining data from Jersey Norman French (known to its speakers as Jèrriais) it illustrates the difficulty in determining linguistic norms for this relatively undocumented variety and suggests possible strategies to overcome this problem. The study compares systematically the occurrence of overt and covert transfer in the speech of a sample of fifty native speakers of Jèrriais via the analysis of a number of linguistic variables. The extent to which transfer-induced changes are themselves becoming established as norms within this speech community will also be considered.
The brain basis of action words may be neuron ensembles binding language- and action-related information that are dispersed over both language- and action-related cortical areas. This predicts fast spreading of neuronal activity from language areas to specific sensorimotor areas when action words semantically related to different parts of the body are being perceived. To test this, fast neurophysiological imaging was applied to reveal spatiotemporal activity patterns elicited by words with different action-related meaning. Spoken words referring to actions involving the face or leg were presented while subjects engaged in a distraction task and their brain activity was recorded using high-density magnetoencephalography. Shortly after the words could be recognized as unique lexical items, objective source localization using minimum norm current estimates revealed activation in superior temporal (130 msec) and inferior frontocentral areas (142-146 msec). Face-word stimuli activated inferior frontocentral areas more strongly than leg words, whereas the reverse was found at superior central sites (170 msec), thus reflecting the cortical somatotopy of motor actions signified by the words. Significant correlations were found between local source strengths in the frontocentral cortex calculated for all participants and their semantic ratings of the stimulus words, thus further establishing a close relationship between word meaning access and neurophysiology. These results show that meaning access in action word recognition is an early automatic process ref lected by spatiotemporal signatures of word-evoked activity. Word-related distributed neuronal assemblies with specific cortical topographies can explain the observed spatiotemporal dynamics reflecting word meaning access.
Although some progress has been made on the quality of Machine Translation in recent years, there is still a significant potential for quality improvement. There has also been a shift in paradigm of machine translation, from “classical” rule-based systems like METAL or LMT1 towards example-based or statistical MT.2 It seems to be time now to evaluate the progress and compare the results of these efforts, and draw conclusions for further improvements of MT quality.
Four experiments examined the role of meaning frequency (dominance) and associative strength (measured by associative norms) in the processing of ambiguous words in isolation. Participants made lexical decisions to targets words that were associates of the more frequent (dominant) or less frequent (subordinate) meaning of a homograph prime. The first two experiments investigated the role of associative strength at long SOAs (Stimulus Onset Asynchrony) (750 ms.), showing that meaning is facilitated by the targets' associative strength and not by their dominance. The last two experiments traced the role associative strength at short SOAs (250 ms), showing that the manipulation of the associative strength has no effect in the semantic priming. The conclusions are: on the one hand, semantic priming for homographs is due to associative strength manipulations at long SOAs. On the other hand, the manipulation of the associative strength has no effect when automatic processes (short SOAs) are engaged for homographs.
Reviewed by: Henry James and Queer Modernity Jonathan Warren Eric Haralson. Henry James and Queer Modernity. New York: Cambridge UP, 2003. 265 pp. $60.00 (cloth). Eric Haralson's brilliantly reasoned, witty, and erudite study discovers the emergence of a variously inflected, incrementally evolving, and broadly comprehensible queerness operative in or, better, unavoidably definitive of, Henry James. It is a critical history of the inevitability of James's queerness. One would be wrong to imagine that the inevitability is merely governed by indisputable facts of sexuality. With the example of Haralson's fine synthesis of the excellent Jamesian sexuality scholarship of the past decade, the best of James studies is clearly well past such simplistic essentialism in diagnosing queerness. His book rightly turns away from any "misguided critical and popular obsession with genital proof of James's homosexuality" (221 n. 21) and smartly jettisons limited arguments based on lexical dating of "queer" as a synonym for homosexual, denying Jamesian diction the possibility of such reference. The great generosity of Haralson's work, especially to Jamesians and scholars and critics of international modernism, twentieth-century American literature, and queer studies is its methodological example, a specific alternative to unsophisticated impasses by way of historically assiduous close reading and marvelously entertaining instruction. Haralson's book unfolds as a series of such readings, chronologically patterned in six chapters, plus an introduction and "coda." The first four chapters track James's own evocation of "the broad, complex cultural process—a process uneven, shadowy, and multiply sited—by which 'queer' came to include 'homosexual' among its meanings" (5): from the inchoate sexual significations of "protogay aesthetes" in Roderick Hudson and The Europeans to James's deepening interest in alternative styles of masculinity, male friendship, and queer camaraderie in The Tragic Muse and "The Author of 'Beltraffio'" to "The Turn of the Screw," for Haralson a "monitory fable about the contagion of boyhood [End Page 105] homosexuality" (24) in which the governess frantically envisions the very limits of Victorian insight into "the most unnamable of things" (83). The Ambassadors marks the "apogee" of James's challenge to gender and sexuality norms (102). Haralson decodes the queerness of Lambert Strether's affinities for Little Bilham, the sexual implications of Strether's celebrated outpouring of "tutorial effusion" (104), his bachelor respectability as an inclination toward "rare youth[s]" (AM 133), and his envious fascination with women who can elicit outbursts of male virility. By the 1950s, Hattie Jacques and the international ado of Lamebrain Stretcher and Chapstick Nuisance from Asshole, Mass. were camp allusions, fully available to queer mockery, tribute, and the like (103). Haralson's last chapters track how James could become Hattie Jacques, focusing on how he signifies for Cather, Stein, Hemingway, and their contemporaries. Via rich readings of novels, correspondence, criticism, and attributed remarks, which my hints can only schematize, James emerges as a paragon of "sensuous yet ethically earnest and sufficiently masculine aestheticism" (25) for Cather (in contradistinction to the catastrophe of Oscar Wilde), as a paranoiac phantasm of male effeminacy for Hemingway, and, for Stein, as a heroic general of extraordinary distinction, unwilling to "forfeit... a self resistantly 'queer' to normative pressures, and possessing the integrity of its difference" (210). Reassuringly aware of queerness's inexorable defiance of the mere platitude of statement, Haralson distinguishes James, in his own writing, as an increasingly savvy celebrant of a queerness "at once powerful and elusive" (1), not to mention typically powerful because elusive. Cautiously resistant to readings that would presume to impose alien matrices of logic and identity, Haralson discerns in James's writing—and in those all-important modernist readings that imagine "Henry James" into literary and broader cultural existence (as exemplar or admonitory arbiter or effete monstrosity or fag, for example)—the emergence of such structures in and for their time. As its account of queerness's liveliness and multiple functions unfolds, Haralson's book is superbly cognizant of the specific phases of its vexed legibility, particularly as queerness orients and productively disorients normative surveillance and regulation during the last decades of the nineteenth century and the first few of the twentieth. Haralson is exactly right when he observes, "Feeling or reading the...
There is growing evidence that Internet-mediated psychological tests can have satisfactory psychometric properties and can measure the same constructs as traditional versions. However, equivalence cannot be taken for granted. The prospective memory questionnaire (PMQ; Hannon, Adams, Harrington, Fries-Dias, & Gibson, 1995) was used in an on-line study exploring links between drug use and memory (Rodgers et al., 2003). The PMQ has four factor-analytically derived subscales. In a large (N763) sample tested via the Internet, only two factors could be recovered; the other two subscales were essentially meaningless. This demonstration of nonequivalence underlines the importance of on-line test validation. Without examination of its psychometric properties, one cannot be sure that a test administered via the Internet actually measures the intended construct.
In this paper, we discuss the role that temporal information plays in natural language text, specifically in the context of question answering systems. We define a descriptive framework with which we can examine the temporally sensitive aspects of natural language queries. We then investigate broadly what properties a general specification language would need, in order to mark up temporal and event information in text. We present a language, TimeML, which attempts to capture the richness of temporal and event related information in language, while demonstrating how it can play an important part in the development of more robust question answering systems.
This paper presents a method for quantitatively estimating intonation variation in Mandarin speech. Intonation variation is relative to identical lexical tone structures, and its estimation is performed on two sets of fundamental frequency ( ) contours: one for norms and the other as variants. This is done by transforming target values in pairs from the norms to the variants in which the prosodic contribution to these contours is analyzed as sequences of targets, all of which are confined to the basic elements of the underlying lexical tone structures. The tone transformations are constrained under an assumption of the structural formulation of contours proposed previously. When the norms take the base values of the four lexical tones measured from isolated words in a neutral mood and voice, this method solves acoustic correlations of tone and intonation from the observed contours. The method was implemented on a computer, and its capability of estimating intonation variation was shown through the analysis and synthesis of contours.
In this article, I explore the ways in which ethnic identity is expressed by following the formulaic socio-linguistic norm, the very method of which defies the authenticity of identity itself, thereby asserting the identity's multi-facetedness as sustained in performative linguistic practice. I look at multi-sited socio-linguistic interactions among Koreans in Japan, who claim their primary identity to be that of North Korea's overseas citizens even though none of them have North Korean passport or nationality. Their identity, in other words, is based on ideological commitment, which is in reality supported by their ongoing linguistic practice. A close look at their socio-linguistic life reveals their ethnicity's dual or multiple ontology, which challenges among other things the currently dominant assertion of Japanese self in the western academic discourse.
This paper addresses the current needs for so-called emotion in speech, but points out that the issue is better described as the expression of relationships and attitudes rather than the currently held raw (or big-six) emotional states. From an analysis of more than three years of daily conversational speech, we find the direct expression of emotion to be extremely rare, and contend that when speech technologists say that what we need now is more ‘emotion’ in speech, what they really mean is that the current technologies are too text-based, and that more expression of speaker attitude, affect, and discourse relationships is required.
"They... Speak Better English Than the English Do":Colonialism and the Origins of National Linguistic Standardization in America Paul K. Longmore (bio) Recent scholarship has traced efforts to fashion an American national language through standardization of forms and usage. Christopher Looby, in Voicing America: Language, Literary Form, and the Origins of the United States, David Simpson, in The Politics of American English, 1776–1850, and Kenneth Cmiel, in Democratic Eloquence: The Fight Over Popular Speech in Nineteenth-Century America, all examine public debates about these matters and, in particular, the labors of linguistic reformers to shape the national tongue. But these important studies focus mainly on the revolutionary, early national, and antebellum periods, and although Cmiel recounts the impact of late eighteenth-century British prescriptivists on postrevolutionary American thinking he does not extensively consider colonial efforts to regulate the language.1 In fact, attempts to shape written and spoken American English according to ideas of correctness, propriety, and, most important, a national standard began before American Independence. But those ideas and that standard were British rather than American. The effort reflected colonial desire to copy metropolitan English linguistic norms in order to attain cultural legitimacy within the British Empire. Postrevolutionary exertions perpetuated attitudes and activities that began in the late colonial period. This essay examines the colonial origins of the movement to standardize and nationalize American English. The central fact of colonials' experience is that they act as agents of an expansionist imperial society. As one result, dominant colonial groups are acutely aware of the metropolitan standard of the language they share with the homeland. In developing an extraterritorial variety of that language, they often labor to match the metropolitan standard. Transplanted speakers of various dialects of a common tongue encounter one another in [End Page 279] new geographical and social environments. Contact often produces dialect mixing and leveling and a compromise dialect called a koine. Koineization largely involves unconscious modification of speech forms. But the attentiveness of many colonials to a metropolitan standard indicates that colonial koines arise from not just spontaneous changes but conscious shaping. As users of the koine "nativize" their common tongue, they continuously render normative judgments about alternative usages. Prescribing what is correct, they seek to standardize the extraterritorial version of the language (Siegel 8; Haas; Stein). North American British colonials, especially those in the elite and middling ranks, took as their model the written and spoken English of the imperial center. Like elite and middling Britons, higher-status colonials used this "proper" and "true" English to distinguish themselves from people below them in the social hierarchy. Nonetheless and again like socially ambitious Britons, many colonials wielded linguistic correctness as a tool of social mobility. Colonials of all ranks emulated metropolitan Standard English in order to elevate their standing within the Empire. In the long run in a pattern typical of colonies of settlement, their efforts unintentionally helped to create a common language that provided one basis for American nationhood. Colonials' adoption of the metropolitan standard of English and their manner of applying it appear in three kinds of evidence: contemporary observers' evaluations of colonial speech; higher-status colonials' descriptions of British immigrants' non-standard English speech; and colonials' formal efforts to educate themselves in metropolitan Standard English. Eighteenth-century observers praised Anglophone colonials for matching metropolitan linguistic norms. They focused on pronunciation and accent, vocabulary and phraseology. William Eddis, secretary to Maryland's royal governor (1769–1777), avowed, "[T]he pronunciation of the generality of the people has an accuracy and elegance that cannot fail of gratifying the most judicious ear" (33). Jonathan Boucher, a tutor and Anglican priest in the Chesapeake (1759–1775), asserted that colonials displayed "the purest Pronunciation of the English Tongue that is anywhere to be met with" (30). "Accuracy," "elegance," and "purity" referred to both colonials' emulation of metropolitan standard pronunciation and the absence from their speech of British regional accents. Lord Adam Gordon, [End Page 280] a Scot, made the same point about word usage and grammar. Describing mid-1760s Philadelphia, he admitted that "the propriety of Language here surprized me much, the English tongue being spoken by all ranks, in a degree of purity and perfection, surpassing...
In this paper we present a quantitative and qualitative analysis of annotation in the Hinoki treebank of Japanese, and investigate a method of speeding annotation by using part-of-speech tags. The Hinoki treebank is a Redwoods-style treebank of Japanese dictionary definition sentences. 5,000 sentences are annotated by three different annotators and the agreement evaluated. An average agreement of 65.4% was found using strict agreement, and 83.5% using labeled precision. Exploiting POS tags allowed the annotators to choose the best parse with 19.5% fewer decisions.
Most recent statistical parsers fall into one of two groups. The largest group consists of parsers which are based on some variation of a probabilistic context-free grammar, use joint probability models, and use tabular methods to find the most probable parse. Parsers in the second group are based on probabilistic push-down automata, use conditional probability models, and use some form of state-space search to find the most probable parse. This thesis is a study of natural language parsing as a control problem. This view leads to parsers of the second type. We show that search can be done very efficiently for such parsers. The control approach leads to a particular interpretation of the history-based parsing tradition, in which history is equated with state. The corresponding probability model is called a Markov parsing model, which can be used both for syntactic disambiguation and for search. The resulting parsers are simple, fast, have excellent coverage, and are reasonably accurate. Using treebanks (collections of text, which are expert-annotated with syntactic structure), we learn controllers for parsers that can be applied with little or no search. We call these greedy or nearly-greedy policies. Thus we are studying parsers which are constrained to operate efficiently.
Dismal is a spreadsheet that works within GNU Emacs, a widely available programmable editor. Dismal has three features of particular interest to those who study behavior: (1) the ability to manipulate and align sequential data, (2) an open architecture that allows users to expand it to meet their particular needs, and (3) an instrumented and accessible interface for studies of human-computer interaction (HCI). Example uses of each of these capabilities are provided, including cognitive models that have had their sequential behavior aligned with subject’s protocols, extensions useful for teaching and doing HCI design, and studies in which keystroke logs from the timing package in Dismal have been used.
Abstract. We propose to apply classical development methodologies to the design and implementation of Lexical Databases(LDB), which embody conceptual and linguistic knowledge. We represent the conceptual knowledge as an ontology, and the linguistic knowledge, which depends on each language, in lexicons. Our approach is based on a single language-independent ontology. Besides, we study some conceptual and linguistic requirements; in particular, meaning classifications in the ontology, focusing on taxonomies. We have followed a classical software development methodology for implementing lexical information systems in order to reach robust, maintainable, and integrateable relational databases (RDB) for storing the conceptual and linguistic knowledge. 1
To ease the interpretation of higher order factor analysis, the direct relationships between variables and higher order factors may be calculated by the Schmid-Leiman solution (SLS; Schmid & Leiman, 1957). This simple transformation of higher order factor analysis orthogonalizes first-order and higher order factors and thereby allows the interpretation of the relative impact of factor levels on variables. The Schmid-Leiman solution may also be used to facilitate theorizing and scale development. The rationale for the procedure is presented, supplemented by syntax codes for SPSS and SAS, since the transformation is not part of most statistical programs. Syntax codes may also be downloaded from www.psychonomic.org/archive/.
A visual presentation procedure is introduced that presents target words followed by a dynamic mask until recognition. This form of stimulus degradation prolongs the word recognition process. Differences in word recognition latencies—which are usually quite small—are magnified, and thus can be more easily observed. The results of two experiments on the Internet with a total of 141 participants establish the task’s ability to magnify differences in word recognition latencies stemming from word familiarity (Experiment 1) and word prototypicality (Experiment 2). Both factors interact with stimulus degradation, but at different presentation intervals; these results are discussed as evidence for comparing models of word recognition. The new procedure can be used for assessing individual differences, such as implicit motives and self-focused attention. Further applications are discussed.
Broadly diffused languages, like Spanish, enjoy diverse variations. In the international space common to Spanish, such as that covered by radio and television, the most sought after variants are those with the largest audience, although criteria are often subjective. In order to make decisions in this matter, the article offers proposals based on dispersion criteria (e.g. the number of countries that use a given norm) and population criteria (number of speakers). To exemplify this, the phonetic norms most frequently heard internationally are described, and cases of lexical variation are presented. The conclusion is that the media, promote national and international standardization turning it is to their advantage.
The European Language Resources Association (ELRA) was founded in 1995 with the mission of providing language resources (LR) to European research institutions and companies. In this paper we describe the background, the mission and the major activities since then.
This first issue of Language Resources and Evaluation is dedicated to the memory of Antonio Zampolli, whom few would dispute is the one person who has led the way in promoting and establishing the development of language resources (LR) of all kinds for the past four decades. In this inaugural issue, we have attempted to bring together articles by major figures in the field in order to provide an overview of the history, state of the art, and the future of the creation, annotation, exploitation, evaluation, and distribution of LR. Hopefully, this collection of articles will serve not only as a tribute to Antonio, but also as a framework out of which this journal – which almost certainly would not have existed were it not for him – can grow.
We present a new method to describe the contextual meaning of a key word in a corpus. The vocabulary of the sentences containing this word is compared to that of the entire corpus in order to highlight the words which are significantly overutilized in the neighbourhood of this key word (they are associated in the author’s mind) and the ones which are significantly underutilized (they are mutually exclusive). This method provides an interesting tool for lexicography and literary studies as is shown by applying it to the word amour (love) in the work of Pierre Corneille, the most famous French playwright of the 17th century.
The role of language resources and language technology evaluation is now recognized as being crucial for the development of written and spoken language processing systems. Given the increasing challenge of multilingualism in Europe, the development of language technologies requires a more internationally distributed effort. This paper first describes several recent and on-going activities in France aimed at the development of language resources and evaluation. We then outline a new project intended to enhance collaboration, cooperation, and resource sharing among the international language processing research community.
This paper presents a Chinese parsing method which takes data-oriented parsing technique as the basic framework and utilizes the similarity-based probability estimate technique. Through the initial selection process, the fragment-combination forms of the input sentence are acquired on the constructed knowledge source including treebank, fragment-bank and fragment-combination-bank. Then by using the similarity-based probability estimate technique, the combination parsing process can be completed successfully. To prove the method efficiency, the knowledge source is constructed on the real-world Chinese corpus, and the other corpus is used as the test set. The experiment results show that every test parameter is satisfied.
This paper presents a high performance method to identify English proper nouns (PNs) based on maximum entropy model (MaxEnt). Most traditional PNs recognition systems use lexical resources such as name list, as new names are constantly coming into existence, these are necessarily incomplete. Therefore machine learning methods are used to identify PNs automatically. In the framework of MaxEnt model, semantic and lexical information of surrounding words and word itself acting as atomic features comprises feature templates and forms feature without requiring extra expert knowledge. The test on WSJ of Penn Treebank II shows that this method guarantees high precision and recall, and at the same time it can reduce the quantity of features dramatically, downsize system space consumption, and decrease the time of training and testing, so as to improve the efficiency considerably. The method in this paper can be transformed to identify other specific noun easily because the principle of methods is universal.
Linguistic translation refers to a theory according to which the translator finds equivalents in the target language for all the linguistic elements of the source text without any radical change or additional explanation. This strategy of translation is usually used to translate those texts in which the style, in addition to the meaning, is also of special importance and plays a significant role in conveying the exact meaning. Some theorists regard this translation a kind or form-based translation which may sometimes, especially in complex sentences, result in obscurity, awkwardness, and unintelligibility. In cultural adaptation, used in finding cultural equivalence, the purpose is not translating the individual words, as cultural equivalence is different from lexical equivalence. What is lost from the meaning in adaptation is usually more than that which is either "lost" or remains awkward and unitelligible in linguistic translation. Therefore, linguistic translation, when observes the norms and standards of the target language in communicating the meaning in a clear and natural way, can be a better strategy for translating religious, literary, and technical texts.
Many of the available image databases have keyword annotations associated with the images. In spite of the availability of good quality low-level visual features that reflect well the physical content, image retrieval based on visual features alone is subject to semantic gap. Text annotations are related to image context or semantic interpretation of the visual content and are not necessarely directly linked to the visual appearance of the images. Keywords and visual features thus provide complementary information. Using both sources of information is an advantage in many applications and recent work in this area reflects this interest. In this paper, we address the challenge of semantic gap reduction using a hybrid visual and conceptual representation of the content within an active relevance feedback context. We introduce a new feature vector, based on the keyword annotations available for the images, which makes use of conceptual information extracted from an external lexical database, information represented by a set of "core concepts". Our experiments show that the use of the proposed hybrid conceptual and visual feature vector dramatically improves the quality of the relevance feedback results.
With the growing access to heterogeneous and independent data repositories, determining the semantic difference of two ontologies is critical in information retrieval, information integration and semantic web applications. In this paper, we propose an ontology comparison tool based on a novel senses refinement algorithm, which builds a senses set to accurately represent the semantics of the input ontology. The senses refinement algorithm automatically extracts senses from the electronic lexical database WordNet (locally installed or online), removes unnecessary senses based on the relationship among the entity classes of the ontology, and specifies relations and constraints of the concepts in the refined senses set. The senses refinement converts the measurement of ontology difference into simple set operations based on set theory, thus ensures the efficiency and accuracy of the ontology comparison. Our experimental studies show that the proposed senses refinement algorithm outperforms the naive senses set construction algorithm in terms of efficiency and accuracy.
The development of technologies for monitoring the welfare of crewmembers is a critical requirement for extended spaceflight. Behavior analytic methodologies provide a framework for studying the performance of individuals and groups, and brief computerized tests have been used successfully to examine the impairing effects of sleep, drug, and nutrition manipulations on human behavior. The purpose of the present study was to evaluate the feasibility and sensitivity of repeated performance testing during spaceflight. Four National Aeronautics and Space Administration crewmembers were trained to complete computerized questionnaires and performance tasks at repeated regular intervals before and after a 10-day shuttle mission and at times that interfered minimally with other mission activities during spaceflight. Two types of performance, Digit-Symbol Substitution trial completion rates and response times during the most complex Number Recognition trials, were altered slightly during spaceflight. All other dimensions of the performance tasks remained essentially unchanged over the course of the study. Verbal ratings of Fatigue increased slightly during spaceflight and decreased during the postflight test sessions. Arousal ratings increased during spaceflight and decreased postflight. No other consistent changes in rating-scale measures were observed over the course of the study. Crewmembers completed all mission requirements in an efficient manner with no indication of clinically significant behavioral impairment during the 10-day spaceflight. These results support the feasibility and utility of computerized task performances and questionnaire rating scales for repeated measurement of behavior during spaceflight.
We investigated the performance efficacy of beam search parsing and deep parsing techniques in probabilistic HPSG parsing using the Penn treebank. We first tested the beam thresholding and iterative parsing developed for PCFG parsing with an HPSG. Next, we tested three techniques originally developed for deep parsing: quick check, large constituent inhibition, and hybrid parsing with a CFG chunk parser. The contributions of the large constituent inhibition and global thresholding were not significant, while the quick check and chunk parser greatly contributed to total parsing performance. The precision, recall and average parsing time for the Penn treebank (Section 23) were 87.85%, 86.85%, and 360 ms, respectively.
The web has caused an explosion of documents, requiring the need for an automated text categorization system. This paper explores the notion of semantic feature selection by employing WordNet [Introduction to WordNet: An On-line Lexical Database], a lexical database. The proposed semantic approach employs noun synonyms and word senses for feature selection to select terms that are semantically representative of a category of documents. The categorical sense disambiguation extends the use of WordNet, which has been typically used for text retrieval and word sense disambiguation [A WordNet-based Algorithm for Word Sense Disambiguation]. Our experiments on the Reuters-21578 dataset have shown that automated semantic feature selection is able to perform better than well known statistical feature selection methods, Information Gain and Chi-Square as a feature selection method.
Linguistic politeness is intimately connected with social norms. Estoniansociety has gone through considerable change over the last ten years. It hasregained independence and, at the same time, switched from a planned toa market economy as well as from dictatorship to democracy. A decade ismost probably not long enough for linguistic norms to change drastically:as we know, the structure of a language often takes much longer to change.Politeness, however, may to some extent be subject to deliberate influence,as witnessed, for example, by the reform of Swedish du (you, sg.) where therecommendations of some left-wing organisations on the usage of mutualdu (T) have won general social acceptance. It is, thus, not unlikely that change is taking place in Estonian politeness at present....
Norms may also be understood as social realization of correctness notions and linguistic norms as performance instructions. This article reports on a case study of the application of Toury' s norm theory, particularly his operative translation norms, in subtitle translation. The authors believe that is a norm-governed communicative activity between two or more languages, and that subtitle translation, governed by linguistic and textual norms, should seek invisibility of subtitling as the ultimate goal.
Preservice and experienced teachers ( N=58, from 7 universities) wrote lesson plans for a hypothetical beginning band lesson, using one page from a band method book as source material. Lesson plans were analyzed for word count, level of detail, and for strategies that appeared most frequently. Experienced teachers used fewer words than undergraduates but revealed the same number of strategies and level of detail, on average. There were institutional differences in the variety of strategies incorporated, indicating certain institutions may value a wider range of strategies and activities in beginning band classes. Participants also compared their written plans to a published lesson plan and rated their familiarity with various approaches, giving another view on strategies considered most common. Familiarity ratings were similar when comparing preservice and experienced teachers and when comparing institutions. Degrees of prevalence of specific strategies, such as decontextualization of material, repetition, and modeling are discussed. May 7, 2004 January 18, 2005.
Linguistic research and language technology development employ large repositories of ordered trees. XML, a standard ordered tree model, and XPath, its associated language, are natural choices for storing and querying linguistic data. However, several important expressive features required for linguistic queries are missing in XPath. In this paper, we motivate and illustrate these features with a variety of linguistic queries. Then we define extensions to XPath which support linguistic tree queries. We provide a relational representation for trees, and define an SQL translation for queries. Experiments demonstrate that the query system is significantly faster than other linguistic tree query systems for a wide range of queries. 1