Papers reviewed and determined not to be word norm studies. Use the flag icon to report errors or suggest re-inclusion.
18265 papers
Although some progress has been made on the quality of Machine Translation in recent years, there is still a significant potential for quality improvement. There has also been a shift in paradigm of machine translation, from “classical” rule-based systems like METAL or LMT1 towards example-based or statistical MT.2 It seems to be time now to evaluate the progress and compare the results of these efforts, and draw conclusions for further improvements of MT quality.
Translation tests are widely used for high school term tests and entrance examinations as well as university entrance examinations. Although a considerable number of papers point out the possibility of low reliability for scoring, little is known about the factors raters play in the reliability of scoring (Watanabe, 1994). This study examines how the professional backgrounds of raters affect rating criteria. The results indicated that novice raters tended to over-estimate examinees' comprehension whereas experienced raters were more likely to focus on the correctness of the Japanese sentence. In addition, it turned out that the difficulty of sentences affected the scoring of both experienced and novice raters. After administering a sorting task, the difficulty of the sentences showed that the perception of sentence difficulty did not correspond to the difficulty of examinees' translation. The paper closes by suggesting several pedagogical implications for administering translation tests. Of particular importance is that test developers should consider not only the complexity of sentence structures and vocabulary, but also the examinees' topic familiarity of the sentences to be translated.
We present a system that automatically identifies Attribution, an intra-sentential relation in the RST Treebank. The system uses uses syntactic information from Penn Treebank parse trees. It identifies Attributions as structures in which a verb takes an SBAR complement, and achieves a f-score of.92. This supports our claim that the Attribution relation should be eliminated from a discourse treebank, since it represents information that is already present in the Penn Treebank, in a different form. More generally, we suggest that intra-sentential relations in the RST Treebank might all be eliminable in this way. 1
Epidemiological evidence indicates that nearly 100,000 Gulf War veterans (GWVs) returned from the Persian Gulf reporting myriad symptoms with no apparent medical explanation. Predictive and confirmatory factor analytic procedures verified musculoskeletal pain as one of three categories describing the unexplained symptoms of GWVs. Over a decade has passed and the pain complaints of GWVs have received little scientific attention, beyond establishing that a problem exists. PURPOSE To comprehensively assess pain sensitivity and to determine the effect of exercise in GWVs with medically unexplained musculoskeletal pain versus healthy veteran controls. METHODS Thirty GWVs (n=12 unexplained muscle pain; n=18 healthy) completed a series of psychophysical pain assessments designed to determine pain sensitivity to heat and pressure stimuli. Testing included: 1) measures of pressure pain threshold and tolerance (in seconds), 2) suprathreshold ratings for pressure pain intensity and affect (0–20 scale), 3) heat pain thresholds (degrees C), and 4) pain intensity and affect ratings in response to temperatures ranging from 44 degrees C to 50 degrees C (Descriptor Differential Scaling). Testing occurred prior to and following 30 minutes of cycling exercise set at 70% of peak oxygen capacity. RESULTS There were no group differences in age, height, weight or self-reported physical activity. There were no group differences for either pressure or heat pain thresholds or pressure pain ratings. GWVs with unexplained muscle pain consistently rated heat stimuli as more intense and unpleasant compared to controls both pre [Intensity: F(1,27)=4.1, p=0.05; Affect: F(1,27)=5.4, p=0.03] and post [Intensity: F(1,27)=7.9, p=0.009; Affect: F(1,27)=7.4, p=0.01] exercise. GWVs with unexplained pain rated exercise as more painful (t=3.2, p=0.003), but not as requiring more effort than controls. Exercise generally had a large effect on heat pain ratings for sick veterans but had no effect on heat pain ratings for healthy veteran controls. CONCLUSIONS GWVs with medically unexplained musculoskeletal pain are more sensitive to experimental heat pain stimuli than healthy veteran controls, with the differences being larger following 30 minutes of intense exercise. Surprisingly, healthy GWVs did not exhibit the expected analgesic response following exercise. Our data suggest that a multidimensional and multimodal approach are important for determination of psychophysical differences and exercise effects in GWVs with medically unexplained musculoskeletal pain. Supported by DVA #561-00215
We describe briefly the redevelopment of Space Fortress (SF), a research tool widely used to study training of complex tasks involving both cognitive and motor skills, to be executed on currentgeneration systems with significantly extended capabilities, and then compare the performance of human participants on an original PC version of Space Fortress (SF) with the revised Space Fortress (RSF). Participants trained on SF or RSF for 10 sets of eight 3-min practice trials and two 3-min test trials. They then took tests involving retention, resistance to secondary task interference, and transfer to a different control system. They then switched from SF to RSF or from RSF to SF for 2 sets of final tests and completed rating scales comparing RSF and SF. Slight differences were predicted on the basis of a scoring error in the original version of SF used and on slightly more precise joystick control in RSF. The predictions were supported. The SF group started better but did worse when they transferred to RSF. Despite the disadvantage of having to be cautious in generalizing from RSF to SF, we conclude that RSF has many advantages, which include accommodating new PC hardware and new training techniques. A monograph that presents the methodology used in creating RSF, details on its performance and validation, and directions on how to download free copies of the system may be downloaded from www .psychonomic.org/archive/.
The gradient descent optimization method has been a de facto standard learning algorithm in computational models of category learning. However, it can be considered as a normative (vs. descriptive) model of human learning processes. In particular, there are three concerns associated with the learning algorithm& #x2014;namely, complexity, regularity, and context independency. In response to these limitations, the present study introduces an alternative, hypothesis-testing& #x2014;like learning algorithm on the basis of a stochastic optimization method. The new learning model, termed SCODEL, provides qualitatively simple interpretations for its implied category-learning processes. Moreover, SCODEL is the first modeling attempt to depict individually unique and context-dependent learning processes. Four simulation studies were conducted and showed that the present model has the competence to operate as several different types of learners in various plausibly real-life situations.
This article is devoted to the problem of quantifying noun groups in German. After a thorough description of the phenom ena, the results of corpus-based investigations are described. Moreover, some examples are given that underline the necessity of integrating some kind of information other than grammar sensu stricto into the treebank. We argue that a more sophisticated and fine-grained annotation in the treebank would have very positve effects on stochastic parsers trained on the treebank and on grammars induced from the treebank, and it would make the treebank more valuable as a source of data for theoretical linguistic investigations. The information gained from corpus research and the analyses that are proposed are realized in the framework of SILVA, a parsing and extraction tool for German text corpora.
本研究では語彙概念構造を利用した動詞の辞書データベースを構築する.語彙概念構造は構文上での語の振る舞いを基に意味構造を記述するもので,こうした辞書データは言語処理において特に言語生成,例えば翻訳や言い換え,要約といった応用に必須となる言語資源である.提案する辞書データは語彙概念構造のみを記述するのではなく,語彙概念構造の分析で明らかにする語の振る舞い,ならびに分析の基となる統語テストを記述するため,分類根拠がわかりやすく透明性の高い辞書データとなることが期待できる.
This paper presents an integrated interactive user interface for teaching grammatical analysis on the Internet (Visual Interactive Syntax Learning), developed at the University of Southern Denmark, offering a unified system of analysis for 22 different languages, 7 of which are supported by live grammatical analysis of running text. For reasons of robustness, efficiency and correctness, the system's internal tools are based on the Constraint Grammar Formalism (Karlsson 1990), but users are free to choose from a variety of notional filters, supporting different descriptional paradigms, with a current teaching focus on syntactic tree structures, language independent grammatical categories and the form-function dichotomy. VISL's core NLP-programs use hybrid multi-level parsers (Bick 2003), while graphical teaching applications and corpus searching tools are implemented as platform independent Java-programs and Perlegi's. Though lexica and parsing rules are developed individually for each language, a common CG and treebank data format facilitates source data transfer into grammar teaching games, structural or color based visualisation, and linguistic revision of corpus data
The structural theory of metaphor (STM) uses techniques from possible worlds semantics to generate and interpret metaphors. STM is presented in detail in The Logic of Metaphor: Analogous Parts of Possible Worlds (Steinhart, 2001). STM is based on Kittay’s semantic field theory of metaphor (1987) and ultimately on Black’s interactionist theory (1962, 1979). STM uses an intensional calculus to specify truth-conditions for many grammatical forms of metaphor. The truth-conditional analysis in STM is inspired in part by Miller (1979) and Hintikka & Sandu (1994). STM is by no means a toy theory. It has been successfully tested on dozens of large texts taken from real authors. Its methods can be applied in very large linguistic databases like WordNet or MindNet.
Different words are usually assumed to be semantically independent in most existing similarity measures, which is not often true in practice. The semantic relatedness between words cannot be conveniently employed in the existing measures. We propose a novel similarity measure based on the earth mover's distance (EMD). In the proposed measure, the semantic distances between words are computed based on the electronic lexical database-WordNet and then the EMD is employed to calculate the document similarity with a many-to-many matching between words. Experiments and results demonstrate the effectiveness of the proposed similarity measure.
Face mask is now a common feature in our social environment. Although face covering reduces our ability to recognize other's face identity and facial expressions, little is known about its impact on the formation of first impressions from faces. In two online experiments, we presented unfamiliar faces displaying neutral expressions with and without face masks, and participants rated the perceived approachableness, trustworthiness, attractiveness, and dominance from each face on a 9-point scale. Their anxiety levels were measured by the State-Trait Anxiety Inventory and Social Interaction Anxiety Scale. In comparison with mask-off condition, wearing face masks (mask-on) significantly increased the perceived approachableness and trustworthiness ratings, but showed little impact on increasing attractiveness or decreasing dominance ratings. Furthermore, both trait and state anxiety scores were negatively correlated with approachableness and trustworthiness ratings in both mask-off and mask-on conditions. Social anxiety scores, on the other hand, were negatively correlated with approachableness but not with trustworthiness ratings. It seems that the presence of a face mask can alter our first impressions of strangers. Although the ratings for approachableness, trustworthiness, attractiveness, and dominance were positively correlated, they appeared to be distinct constructs that were differentially influenced by face coverings and participants' anxiety types and levels.
The paper presents a spoken document summarization scheme using a topic-related corpus and semantic dependency grammar. The summarization score considers speech recognition confidence, word significance, word trigram, semantic dependency grammar (SDG) and probabilistic context free grammar (PCFG). In addition, a topic-related corpus consisting of keywords as well as articles is used to estimate the word significance score using latent semantic indexing (LSI). Semantic relations between words are determined by SDG using HowNet and Sinica Treebank. A dynamic programming algorithm is applied to decide the summarization ratio and look for the best summarization result according to summarization scores. Experimental results indicate that the proposed approach effectively extracts important words with semantic dependency and gives a promising speech summary.
The goal of this article is to account for the resolution of vowel sequences across word boundaries in Catalan. Specifically, the paper accounts for the distribution of hiatuses and syllable contraction cases between two lexical words in this language. The article argues that V1 (the last vowel of the first word) does not undergo any change if it is followed by a vowel V2 bearing nuclear stress (or phrasal stress) prominence. The blocking of V1 glide formation will be seen in relation with the systematic maintenance of schwa in this position (canti ara [i »a] ‘you sing.imp now’, canto ara [u »a] ‘I sing now’, tallo ungles [u »u] ‘I cut nails’, canta ara [ »a] ‘he/she sings now’). Blocking of glide formation or schwa deletion is thus not due to rhythmic reasons (stress clash), as some previous studies have contended, but rather to the presence of a nuclear stress prominence on V2. This phenomenon will be interpreted as the instantiation of an alignment constraint which aligns the word-initial nuclear stressed foot to the left edge of the prosodic word. This alignment constraint triggers a ‘prosodic isolation’ phenomenon which prevents vowel gliding or deletion from applying. Finally, the paper also accounts for vowel sandhi in contexts where V2 is not stressed: in these contexts, syllable contraction is the norm. * Earlier versions of this work were presented at PaPI 2003 (Phonetics and Phonology in Iberia, Lisbon), at the Toulouse International Conference “From representations to constraints” (Toulouse, July 2003) and at the XVth International Congress of Phonetic Sciences (Barcelona, August 2003). We are grateful to those who attended these meetings for interesting observations and comments, and especially to Eulalia Bonet, Sonia Colina, Sonia Frota, Jose Ignacio Hualde, Michael Kenstowicz, John Kingston, Maria Rosa Lloret, Joan Mascaro, John McCarthy, Daniel Recasens, Elizabeth Selkirk, Donca Steriade, Hubert Truckenbrodt, Marina Vigario, and Max Wheeler for discussion of some parts of the material included in the article. Thanks are also due to Nuria Riera for transcribing the vowel contacts present in 5 spontaneous conversations of the Corpus Oral de Catala and to Marta Paya and Lluis Payrato for kindly providing us with a copy of this database before its publication. Finally, we thank Teresa Barenys, Julia Cufi, Teresa Espinal, Anna Gavarro, Nuria Marti, Jaume Sola, and Xavier Vall, who patiently responded to our questionnaire. All remaining errors are of course ours. This research was funded by grants 2002XT-00032 and 2001SGR 00150 from the Generalitat de Catalunya and BFF2003-06590 and BFF2003-09453-C02-C02 from the Ministry of Science and Technology of Spain.
The goal of the present study was to examine how young children form impressions of peers with implications for understanding and ameliorating peer rejection. We focused on four questions: whether young children make dispositional attributions from observed behavior; how children combine complex behavioral information into a coherent evaluation; how other traits are inferred; and how modifiable are existing impressions. The participants were 51 kindergarteners, 53 second graders, and 104 college students. All subjects watched videoclips of a same-sex child actor performing three mean or kind behaviors, three smart or not smart behaviors, and three shy or not shy behaviors. Two of the three behavioral dimensions had the same valence given by a valence condition while the third dimension had opposite valence. Some subjects also had prior expectancies of one trait of predominant valence. We hypothesized that in the expectancy condition, only the expected trait will be represented in memory, while all three traits will be represented if there is no expectancy. After providing ratings of how much they liked the actor and how various traits described him or her, two more video clips of the same actor were presented. These new behaviors were incongruent in valence with the expected trait or, in the no-expectancy condition, with one of the traits with predominant valence. Ratings of liking and the same personality traits were then collected again. The results showed that expectancy did not suppress the encoding of the unexpected traits but it made the ratings of the expected trait more extreme. Importantly, all age groups made dispositional attributions of all three presented traits that accurately reflected the valence of the observed behaviors, although 5-6 year-olds gave less negative ratings to traits with negative valence. The global evaluation of the actor was predicted by all presented trait ratings, and was evaluatively consistent with the inferred trait ratings. The presentation of incongruent behaviors resulted in substantial modifications of all impressions that had predominant valence. Kindergarteners modified their ratings the least, especially if their impressions were positive. Implications for understanding peer rejection are discussed.
This paper presents a statistical approach to unknown word type prediction for a deep HPSG grammar. Our motivation is to enhance robustness in deep processing. With a predictor which predicts lexical types for unknown words according to the context, new lexical entries can be generated on the fly. The predictor is a maximum entropy based classifier trained on a HPSG treebank. By exploring various feature templates and the feedback from parse disambiguation results, the predictor achieves precision over 60%. The models are general enough to be applied to other constraint-based grammar formalisms. 1
Recent empirical experiments on surface realizers have shown that grammars for generation can be effectively evaluated using large corpora. Evaluation metrics are usually reported as single averages across all possible types of errors and syntactic forms. But the causes of these errors are diverse, and the extent to which the accuracy of generation over individual syntactic phenomena is unknown. This article explores the types of errors, both computational and linguistic, inherent in the evaluation of a surface realizer when using large corpora. We analyze data from an earlier wide coverage experiment on the FUF/SURGE surface realizer with the Penn TreeBank in order to empirically classify the sources of errors and describe their frequency and distribution. This both provides a baseline for future evaluations and allows designers of NLG applications needing off-the-shelf surface realizers to choose on a quantitative basis. 1
The purpose of this paper is to automatically generate Chinese chunk bracketing by a bottom-to-top mapping (BTM) model with a BTM dataset. The BTM model is designed as a supporting model with parsers. We define a word-layer matrix to generate the BTM dataset from Chinese Treebank. Our model matches auto-learned patterns and templates against segmented and POS-tagged Chinese sentences. A sentence that can be matched with some patterns or templates is called a matching sentence. The experimental results have shown that the chunk bracketing of the BTM model on the matching sentences is high and stable. By applying the BTM model to the matching sentences and the Ngram model to the non-matching sentences, the experiment results show the F-measure of an N-gram model can be improved. 1
Reviewed by: Word sense disambiguation: The case for combinations of knowledge sources by Mark Stevenson Cornelia Tschichold Word sense disambiguation: The case for combinations of knowledge sources. By Mark Stevenson. (CSLI studies in computational linguistics.) Stanford: CSLI Publications, 2003. Pp. 175. ISBN 1575863901. $25. Disambiguating words is easy for human beings, but difficult for computers. Computational linguistics has developed methods to reliably find the correct part of speech for the large majority of words in running text, but the disambiguation of polysemous words and homonyms (bat as animal, sports tool, or blink of the eye) is a more complex process. This difference is due mainly to the lack of sufficiently complete and formalized data about word senses. Stevenson shows how progress can be achieved by reusing existing lexical databases and combining them in an optimal way. Ch. 1 introduces the problem of polysemy and points out the potential areas of application for word sense disambiguation (WSD). Ch. 2 gives some historical background on the area, intended for readers unfamiliar with the field. In Ch. 3, lexicographic problems associated with polysemous words and attempts at arriving at suitable databases (such as Word-Net) are discussed. As it does not seem likely that machines can take over any significant part of the lexicographic work involved in the production of semantic databases, the re-use of machine-readable dictionaries appears to be the only viable solution for the immediate future. S refutes a number of criticisms that have been made against the use of a machine-readable dictionary for WSD, mainly due to their lack of alternatives, and proposes methods for at least partially remedying the known shortcomings. Ch. 4 describes the knowledge sources that can be used for WSD, that is, syntactic, semantic, and pragmatic information, and the conditions needed to combine them. WordNet and the Longman dictionary of contemporary English (LDOCE) are identified as two potentially useful on-line lexicographic databases. In Ch. 5, the computational similarities and differences of part-of-speech tagging and WSD are explained. In Ch. 6, S explains how his system combining the various knowledge sources was implemented: the preprocessing stage filters out proper names, tokenizes the input text, and identifies the part of speech for each word (using a Brill-type tagger). This is followed by a shallow syntactic analysis and finally the lexical look-up stage. At the disambiguation stage, the part-of-speech tags are used to filter out any (syntactically) incompatible senses, before a number of partial (semantic) taggers are brought into play. The first of these uses LDOCE senses, with any subsenses grouped where possible; the second uses categories of synonyms, and the last selectional restrictions. Known collocations are also taken into account. The implementation involved a memory-based machine learning system that was first trained on annotated data and then used to combine all the knowledge sources for WSD. Chs. 7 and 8 deal with evaluation of the author’s and other known systems for WSD. S illustrates the unsatisfactory state of evaluation tools and procedures in the area of WSD, before demonstrating that his system achieves better results thanks to the combination of a number of available lexical resources. [End Page 1022] The book is a readable introduction and description of the problems WSD poses for computational linguistics, making a clear case for a hybrid approach that uses knowledge-based and corpus-based sources of information to identify the sense of ambiguous words. Cornelia Tschichold University of Wales Swansea, Great Britain Copyright © 2005 Linguistic Society of America
There has been recent interest in looking at what is required for a tree query language for linguis-tic corpora. One approach is to start from exist-ing formal machinery, such as tree grammars and automata, to see what kind of machine is an ap-propriate underlying one for the query language. The goal of the paper is then to examine what is an appropriate machine for a linguistic tree query language, with a view to future work dening a query language based on it. In this paper we review work relating XPath to regular tree gram-mars, and as the paper's rst contribution show how regular tree grammars can also be a basis for extensions proposed for XPath for common lin-guistic corpus querying. As the paper's second contribution we demonstrate that, on the other hand, regular tree grammars cannot describe a number of structures of interest; we then show that, instead, a slightly more powerful machine is appropriate, and indicate how linguistic tree query languages might be augmented to include this extra power. 1
The aim of this paper is to present a lexical database of English collocations used in scientific language, which is being built in three Spanish universities (Barcelona, Illes Balears and Leon) and is mainly intended for the Spanish-speaking scientific community. The shortage of specialized dictionaries providing contextual information on the grammatical and collocational patterns in specific registers prompted the onset of this project. Our database is based on the analysis of a corpus of written texts in the areas of biology, biochemistry, and biomedicine, and provides the grammatical, semantic, and collocational information necessary for the correct and precise use of each term in scientific discourse. The paper describes the steps followed in the creation of the data base and it includes the case study of one of its entries.
Dismal is a spreadsheet that works within GNU Emacs, a widely available programmable editor. Dismal has three features of particular interest to those who study behavior: (1) the ability to manipulate and align sequential data, (2) an open architecture that allows users to expand it to meet their particular needs, and (3) an instrumented and accessible interface for studies of human-computer interaction (HCI). Example uses of each of these capabilities are provided, including cognitive models that have had their sequential behavior aligned with subject’s protocols, extensions useful for teaching and doing HCI design, and studies in which keystroke logs from the timing package in Dismal have been used.
To ease the interpretation of higher order factor analysis, the direct relationships between variables and higher order factors may be calculated by the Schmid-Leiman solution (SLS; Schmid & Leiman, 1957). This simple transformation of higher order factor analysis orthogonalizes first-order and higher order factors and thereby allows the interpretation of the relative impact of factor levels on variables. The Schmid-Leiman solution may also be used to facilitate theorizing and scale development. The rationale for the procedure is presented, supplemented by syntax codes for SPSS and SAS, since the transformation is not part of most statistical programs. Syntax codes may also be downloaded from www.psychonomic.org/archive/.
This paper presents a high performance method to identify English proper nouns (PNs) based on maximum entropy model (MaxEnt). Most traditional PNs recognition systems use lexical resources such as name list, as new names are constantly coming into existence, these are necessarily incomplete. Therefore machine learning methods are used to identify PNs automatically. In the framework of MaxEnt model, semantic and lexical information of surrounding words and word itself acting as atomic features comprises feature templates and forms feature without requiring extra expert knowledge. The test on WSJ of Penn Treebank II shows that this method guarantees high precision and recall, and at the same time it can reduce the quantity of features dramatically, downsize system space consumption, and decrease the time of training and testing, so as to improve the efficiency considerably. The method in this paper can be transformed to identify other specific noun easily because the principle of methods is universal.
We present a strictly lexical parsing model where all the parameters are based on the words. This model does not rely on part-of-speech tags or grammatical categories. It maximizes the conditional probability of the parse tree given the sentence. This is in contrast with most previous models that compute the joint probability of the parse tree and the sentence. Although the maximization of joint and conditional probabilities are theoretically equivalent, the conditional model allows us to use distributional word similarity to generalize the observed frequency counts in the training corpus. Our experiments with the Chinese Treebank show that the accuracy of the conditional model is 13.6% higher than the joint model and that the strictly lexicalized conditional model outperforms the corresponding unlexicalized model based on part-of-speech tags.
Cette thèse aborde le problème de la structuration de bases lexicales multilingues (BDLM) en lexies et axies, à partir de ressources existantes. Ce travail est motivé par l'inadéquation des techniques existantes utilisées isolément, pour la structuration de BDLM. Pour résoudre ce problème, la stratégie proposée est de composer des techniques existantes de désambiguïsation pour structurer semi-automatiquement des bases lexicales multilingues à lexies et acceptions interlingues. De plus, cette thèse propose une catégorisation des critères d'évaluation de la qualité des BDLM, ainsi que les mesures correspondantes. Cette stratégie a été implémentée dans Jeminie, un système logiciel adaptable qui permet d'implémenter à la fois des méthodes de structuration de BDLM et des mesures de qualité, sous la forme de modules logiciels réutilisables. Des compositions arbitraires de ces modules peuvent être définies par un lexicologue dans un langage de haut niveau d'abstraction, ce qui permet d'adapter facilement la structuration et l'évaluation de qualité en fonction des objectifs du lexicologue et des ressources disponibles sans nécessiter de connaissances en programmation. L'intérêt de cette approche a été validé expérimentalement: la qualité des BDLM obtenues est meilleure par combinaison de techniques qu'avec chaque technique antérieure utilisée seule.
We present a treebank conversion method by which we construct an RMRS bank for HPSG parser evaluation from the TIGER Dependency Bank. Our method effectively performs automatic RMRS semantics construction from functional dependencies, following the semantic algebra of (Copestake et al., 2001). We present the semantics construction mechanism, and focus on some special phenomena. Automatic conversion is followed by manual validation. First evaluation results yield high precision of the automatic semantics construction rules. 1
This is a presentation of the historical development of human-computer interfaces in educational computer science, which makes it necessary for Natural Language Processing to understand and interpret the learner”s interventions. It is also an analysis of Lexical Databases and their application in education.
This paper proposes a new document representation method to text categorization. It applies category-based semantic field (CBSF) theory for text categorization to gain a more efficient representation of documents. The lexical chain is introduced to compute CBSF and Hownet* used as a lexical database. In particular, the title of each document functions as a clue to forecast the potential CBSF of the test document. Combined with classifier, this approach is examined in text categorization and the result indicates that it performs better than conventional methods with features chosen on the basis of bag-of-words (BOW) system, on the same task.
Norms may also be understood as social realization of correctness notions and linguistic norms as performance instructions. This article reports on a case study of the application of Toury' s norm theory, particularly his operative translation norms, in subtitle translation. The authors believe that is a norm-governed communicative activity between two or more languages, and that subtitle translation, governed by linguistic and textual norms, should seek invisibility of subtitling as the ultimate goal.
OBJECTIVES: To test for simvastatin-induced changes in affect and affective processes in elderly volunteers. DESIGN: Randomized, clinical trial. SETTING: The Geriatric Behavioral Psychopharmacology Laboratory at the University of Pennsylvania. PARTICIPANTS: Eighty older volunteers, average age 70, with high normal/mildly elevated serum cholesterol. INTERVENTION: Simvastatin up to 20 mg/d or placebo for 15 weeks. MEASUREMENTS: Daily diary records of positive and negative affects and of events and biweekly measures of depressive symptoms. Affect ratings were obtained using the Lawton positive and negative affect scales; independent raters coded the valences of events. RESULTS: Thirty-one of 39 subjects assigned to placebo and 33 of 41 receiving simvastatin completed the study. During biweekly assessments, four subjects on simvastatin and one on placebo experienced depressive symptoms, as manifest by Center for Epidemiological Studies Depression scale scores greater than 16 (exact P=.36). Diary data demonstrated significant effects on affective processes. For positive affect, there was a significant medication-by-time interaction that reflected decreases in positive affect in subjects receiving simvastatin, greatest in those patients whose final total cholesterol levels were below 148 mg/dL. For negative affect, there were significant medication-by-event, and medication-by-event-by-time interactions, reflecting a time-limited increase in the apparent effect of negative events. CONCLUSION: Simvastatin has statistically significant effects on affect and affective processes in elderly volunteers. The decrease in positive affect may be significant clinically and relevant to the quality of life of many patients.