Papers reviewed and determined not to be word norm studies. Use the flag icon to report errors or suggest re-inclusion.
16504 papers
Abstract An analysis was made of 22 supervision groups in two psychotherapy training programmes at different levels. Its main focus concerned role patterns based on self-image ratings and changes over time. The results showed no significant differences between the two categories of supervisees, whereas the differences between the supervisors and the supervisees, independent of level of training, were highly significant. The results indicate that it is just as difficult to find one's voice and role in a supervision group at an advanced as at a basic level. For the supervisors the result was interpreted in terms of their roles in relation to the supervisees and the aim of the supervision.
This study investigated the effect of a weak magnetic field (50 microT, 20 Hz sinusoidal, 5 s duration) on concurrent perceptions of visual stimuli. Subjects were seated between Helmholtz coils and gave post-exposure ratings for the affective content and arousing nature of presented images. They were blind as to the presence or absence of a simultaneously presented field. Skin conductance and arousal ratings did not show significant differences between experimental and control conditions, but the affective content rating did (P = 0.041), with the images viewed under field exposure being rated as having a more positive affect. Such measures might thus be useful as additional indicators of magnetic field detection. A post-hoc analysis of skin conductance profiles showed that 48% of subjects exhibited a lowering of skin conductance during field exposure, 34% exhibited no apparent reaction, and 17% exhibited an increase. Overall ratings given by each of the groups appeared to relate to these physiological profiles.
Traditional theories of finance posit that the pricing of securities in financial markets should be done according to the quality of their underlying technical fundamentals. However, research on financial markets has tended to indicate that factors other than technical fundamentals are often used by market participants to gauge the value of securities. This phenomenon may be quite prevalent in markets for initial public offerings (IPSs), where securities lack a financial history. The imagery and affect associated with securities can be a powerful basis upon which to judge their worth. Advanced business students in a securities analysis course were asked to evaluate a number of industry groups represented on the New York Stock Exchange in terms of a set of judgmental variables. After providing imagery and affective evaluations for each industry group, the participants judged the likelihood that they would invest in companies associated with each industry. Imagery and affective ratings were highly correlated with one another and with the likelihood of investing. Judgments of performance correlated poorly to moderately with actual market performance as measured by weighted average returns for the industry groups studied. The results suggest that imagery and affect are part of a coherent psychological framework for evaluating classes of securities, but that framework may have low validity for predicting performance.
This paper describes a hybrid proposal to combine n-grams and Stochastic Context-Free Grammars (SCFGs) for language modeling. A classical n-gram model is used to capture the local relations between words, while a stochastic grammatical model is considered to represent the long-term relations between syntactical structures. In order to define this grammatical model, which will be used on large-vocabulary complex tasks, a category-based SCFG and a probabilistic model of word distribution in the categories have been proposed. Methods for learning these stochastic models for complex tasks are described, and algorithms for computing the word transition probabilities are also presented. Finally, experiments using the Penn Treebank corpus improved by 30% the test set perplexity with regard to the classical n-gram models.
We present some novel machine learning techniques for the identification of subcategorization information for verbs in Czech. We compare three different statistical techniques applied to this problem. We show how the learning algorithm can be used to discover previously unknown subcategorization frames from the Czech Prague Dependency Treebank. The algorithm can then be used to label dependents of a verb in the Czech treebank as either arguments or adjuncts. Using our techniques, we are able to achieve 88 % accu-racy on unseen parsed text. 1.
Three studies focused on the development and enhancement of narrative skills within a preschool classroom. The purpose of Study 1 was to collect local norms on narrative development. Fifty-two preschool African American English speakers representing 3-, 4-, and 5-year-old age groups, narrated a familiar storybook. Some children in each age group evidenced use of nine story element types. Developmental changes were characterized by growth in types as well as tokens of story elements. Study 2 demonstrated that preschoolers’ narratives can be influenced by the narratives of their peers. Paired children narrated a familiar storybook to each other. The stories of paired children were significantly more similar in form (shared story element types) and content (shared lexical types) than those of unpaired children. Study 3 provided a preliminary test of an intervention designed to exploit the effect of peer models for long-term gain in narrative abilities. Two tutees practiced book narration following the clinician-prompted models of their peer tutors. As a result, the tutees demonstrated an expanded repertoire of story elements and an increased frequency of use of story element types in both trained and untrained stories. Their rate of growth in story element use was superior to that of their classmates who had not participated in the intervention. The benefit of peers for achieving instructional congruence in cases of clinicianclient mismatch is emphasized.
This article reports the results of apreliminary analysis of translation equivalents infour languages from different language families,extracted from an on-line parallel corpus of GeorgeOrwell's Nineteen Eighty-Four. The goal ofthe study is to determine the degree to whichtranslation equivalents for different meanings of apolysemous word in English are lexicalized differentlyacross a variety of languages, and to determinewhether this information can be used to structure orcreate a set of sense distinctions useful in naturallanguage processing applications. A coherenceindex is computed that measures the tendency fordifferent senses of the same English word to belexicalized differently, and from this data aclustering algorithm is used to create sensehierarchies.
This work combines a set of available techniques – whichcould be further extended – to perform noun sense disambiguation. We use several unsupervised techniques (Rigau et al., 1997) that draw knowledge from a variety of sources. In addition, we also apply a supervised technique in order to show that supervised and unsupervised methods can be combined to obtain better results. This paper tries to prove that using an appropriate method to combine those heuristics we can disambiguate words in free running text with reasonable precision.
Age of acquisition (AoA) has been reported to be a predictor of the speed of reading words aloud (word naming) and lexical decision, with early-acquired words being responded to faster than later-acquired words in both tasks. All previous studies of AoA effects have, however, relied upon adult estimates of word learning age the validity of which it is easy to cast doubt upon. Using objective age of acquisition norms derived from children's naming data, this study shows that AoA effects do not depend upon the use of adult ratings. In addition to effects of real AoA, influences of word frequency and orthographic neighbourhood size were obtained in both word naming and lexical decision. Imageability affected lexical decision but not word naming, while the characteristics of the word's initial phoneme affected word naming but not lexical decision.
SENSEVAL set itself the task of evaluating automaticword sense disambiguation programs (see Kilgarriff andRosenzweig, this volume, for an overview of theframework and results). In order to do this, it wasnecessary to provide a `gold standard' dataset of `correct' answers. This paper will describe thelexicographic part of the process involved in creatingthat dataset. The primary objective was for a group oflexicographers to manually examine keywords in a largenumber of corpus contexts, and assign to each contexta sense-tag for the keyword, taken from the Hectordictionary. Corpus contexts also had to be manuallypart-of-speech (POS) tagged. Various observationsmade and insights gained by the lexicographers duringthis process will be presented, including a critiqueof the resources and the methodology.
We discuss the advantages of lexicalized tree-adjoining grammar as an alternative to lexicalized PCFG for statistical parsing, describing the induction of a probabilistic LTAG model from the Penn Treebank and evaluating its parsing performance. We find that this induction method is an improvement over the EM-based method of (Hwa, 1998), and that the induced model yields results comparable to lexicalized PCFG.
This paper describes the evaluation of a WSD method withinSENSEVAL. This method is based on Semantic Classification Trees (SCTs)and short context dependencies between nouns and verbs. The trainingprocedure creates a binary tree for each word to be disambiguated. SCTsare easy to implement and yield some promising results. The integrationof linguistic knowledge could lead to substantial improvement.
This paper presents results for a maximum-entropy-based part of speech tagger, which achieves superior performance principally by enriching the information sources used for tagging. In particular, we get improved results by incorporating these features: (i) more extensive treatment of capitalization for unknown words; (ii) features for the disambiguation of the tense forms of verbs; (iii) features for disambiguating particles from prepositions and adverbs. The best resulting accuracy for the tagger on the Penn Treebank is 96.86% overall, and 86.91% on previously unseen words.
We present some novel machine learning techniques for the identification of subcategorization information for verbs in Czech. We compare three different statistical techniques applied to this problem. We show how the learning algorithm can be used to discover previously unknown subcategorization frames from the Czech Prague Dependency Treebank. The algorithm can then be used to label dependents of a verb in the Czech treebank as either arguments or adjuncts. Using our techniques, we are able to achieve 88% precision on unseen parsed text.
This article considers approaches which rerank the output of an existing probabilistic parser. The base parser produces a set of candidate parses for each input sentence, with associated probabilities that define an initial ranking of these parses. A second model then attempts to improve upon this initial ranking, using additional features of the tree as evidence. The strength of our approach is that it allows a tree to be represented as an arbitrary set of features, without concerns about how these features interact or overlap and without the need to define a derivation or a generative model which takes these features into account. We introduce a new method for the reranking task, based on the boosting approach to ranking problems described in Freund et al. (1998). We apply the boosting method to parsing the Wall Street Journal treebank. The method combined the log-likelihood under a baseline model (that of Collins [1999]) with evidence from an additional 500,000 features over parse trees that were not included in the original model. The new model achieved 89.75 % F-measure, a 13 % relative decrease in F-measure error over the baseline model’s score of 88.2%. The article also introduces a new algorithm for the boosting approach which takes advantage of the sparsity of the feature space in the parsing data. Experiments show significant efficiency gains for the new algorithm over the obvious implementation of the boosting approach. We argue that the method is an appealing alternative—in terms of both simplicity and efficiency—to work on feature selection methods within log-linear (maximum-entropy) models. Although the experiments in this article are on natural language parsing (NLP), the approach should be applicable to many other NLP problems which are naturally framed as ranking tasks, for example, speech recognition, machine translation, or natural language generation.
BOOK NOTICES 209 Linguistic databases. Ed. by John Nerbonne. (CSLI lecture notes 77.) Stanford, CA: CSLI, 1998. Pp. xxi, 243. The papers in this collection were originally presented at the 'Linguistic Databases' conference, University of Groningen, 23-24 March, 1995. Because ofthe almostproverbial rapidity with which information technology develops, the collection as a whole is dated already, but there is still much of interest to be found. Not all papers read at the conference are in this volume, but the papers cover a wide range of subjects, mostly practical in nature, not theoretical. After a clear and readable introduction by Nerbonne, the papers are presented in no particularorder, though the editor groups the papers in five main areas: syntactic corpora and databases, phonetic databases, applications in linguistic theory, applications, and extending basic technologies. The papers themselves are not presented according to this grouping, however, and at first sight the book appears rather disorganized. The wide variety of subjects can be deduced from the titles of the papers presented: 'Test suites for natural language processing', 'From annotated corpora to databases: The SgmlQL language', 'Markup of a test suite with SGML', 'An open systems approach for an acoustic-phonetic continuous speech database: The S_tools database-management system ', "The reading database of syllable structure',? database application for the generation of phonetic atlas maps', 'Swiss French polyphone and polyvar: Telephone speech databases to model inter- and intra-speaker variability', 'Investigating argument structure: The Russian nominalization database', "The use of a psycholinguistic database in the simplification of text for aphasie readers', "The computer learner corpus: A testbed for electronic EFL tools', 'Linking WordNet to a corpus query system', 'Multilingual data processing in the CELLAR environment '. The issue whether to use open free systems or closed proprietary systems is addressed in several papers. Some papers present applications developed both in open and closed systems. This is one area where developments have been going very fast, and nowadays freely available databases are often as capable as their commercial counterparts. Some of the applications presented in this collection are available from the Internet, and url's are often given. The collection can serve as a good introduction to the field for relative outsiders as ample references and links are given. The papers themselves vary greatly in subject matter so not all will be of interest to every reader. My particular favorite was 'From annotated corpora to databases: the SgmlQL language '. [BOUDEWUN REMPT.j Understanding phonology. By Carlos Gussenhoven and Haike Jacobs. (Understanding language series.) London: Arnold, 1998. Pp. xii, 286. This textbook is intended as an introduction to phonology aimed at 'students with little or no prior knowledge of linguistics' (back cover). As in many other textbooks, it uses exercises as a learning tool. Two types ofexercises are proposed. The ones identified by a key, 'intended as an expository aid' (xi), are provided with a solution in an appendix (though it is not always so much a clear cut answer as a guide for reflection, which is, to my view, a lot better). The ones identified by a dot are intended as practice material, and no solution is offered. I thought the idea of having two types of exercises a good one since it gives the reader the opportunity both for individual work and for discussion with others. Also, whenever it may apply, an optimality theoretic analysis is offered to describe a phonological process. Ch. 1, "The production of speech', is a basic introduction to phonology, phonetics, and phonation. Ch. 2,'Some typology: Sameness and difference', cleverly covers the universal and language specific aspects of phonological structures and typology. Ch. 3,'Making the form fit', addresses phonological grammar and adaptation by presenting the nativization of loan words in both the rules and the constraints approaches. Ch. 4, 'Underlying and surface representations', Ch. 5, 'Distinctive features', and Ch. 6, 'Ordered rules', deal with the basic notions of generative phonology within the SPE type formalism and introduce the reader to the school of linear phonology. Ch. 7,? case study: The diminutive suffix in Dutch', shows how these notions are applied. In Ch. 8, 'Levels of representation', Gussenhoven and Jacobs present an intermediate level of representation between the underlying representation and...
Three state-of-the-art statistical parsers are combined to produce more accurate parses, as well as new bounds on achievable Treebank parsing accuracy. Two general approaches are presented and two combination techniques are described for each approach. Both parametric and non-parametric models are explored. The resulting parsers surpass the best previously published performance results for the Penn Treebank.
Abstract Traditional theories of finance posit that the pricing of securities in financial markets should be done according to the quality of their underlying technical fundamentals. However, research on financial markets has tended to indicate that factors other than technical fundamentals are often used by market participants to gauge the value of securities. This phenomenon may be quite prevalent in markets for initial public offerings (IPOs), where securities lack a financial history. The imagery and affect associated with securities can be a powerful basis upon which to judge their worth. Advanced business students in a securities analysis course were asked to evaluate a number of industry groups represented on the New York Stock Exchange in terms of a set of judgmental variables. After providing imagery and affective evaluations for each industry group, the participants judged the likelihood that they would invest in companies associated with each industry. Imagery and affective ratings were highly correlated with one another and with the likelihood of investing. Judgments of performance correlated poorly to moderately with actual market performance as measured by weighted average returns for the industry groups studied. The results suggest that imagery and affect are part of a coherent psychological framework for evaluating classes of securities, but that framework may have low validity for predicting performance.
This paper proposes a new error-driven HMM-based text chunk tagger with context-dependent lexicon. Compared with standard HMM-based tagger, this tagger uses a new Hidden Markov Modelling approach which incorporates more contextual information into a lexical entry. Moreover, an error-driven learning approach is adopted to decrease the memory requirement by keeping only positive lexical entries and makes it possible to further incorporate more context-dependent lexical entries. Experiments show that this technique achieves overall precision and recall rates of 93.40% and 93.95% for all chunk types, 93.60% and 94.64% for noun phrases, and 94.64% and 94.75% for verb phrases when trained on PENN WSJ TreeBank section 00-19 and tested on section 20-24, while 25-fold validation experiments of PENN WSJ TreeBank show overall precision and recall rates of 96.40% and 96.47% for all chunk types, 96.49% and 96.99% for noun phrases, and 97.13% and 97.36% for verb phrases.
In this contribution we discuss how a fuzzy querying interface can support the generation of linguistic database summaries - a special technique of data mining. Links between our approach to linguistic summaries and the well-known technique of association rules is shown. The implementation of linguistic summaries generation using the authors’ FQUERY for Access package is presented.
This article focuses on the user-friendliness of lexical information sources. Whereas our previous study on user-friendliness (Euralex 1998) emphasized the context-sensitive needs of dictionary users, our present study goes one step further and suggests that it would be possible to compile interactive lexical databases that would be both context- and user-sensitive. We approach the function of lexical databases from two perspectives: from their role as primary information sources and from their role as lexical interfaces to other knowledge bases. Our approach is generally based on frame-semantics.We apply semantic frames to capture the different ways of conceptualization used when searching a knowledge base for social and health care services.
It is generally recognized that the common nonterminal labels for syntactic constituents (NP, VP, etc.) do not exhaust the syntactic and semantic information one would like about parts of a syntactic tree. For example, the Penn Treebank gives each constituent zero or more 'function tags' indicating semantic roles and other related information not easily encapsulated in the simple constituent labels. We present a statistical algorithm for assigning these function tags that, on text already parsed to a simplelabel level, achieves an F-measure of 87%, which rises to 99% when considering 'no tag' as a valid choice.
In this paper we present the results of a quantitative evaluation of the discrepancies between the Italian and English lexica in terms of lexical gaps. This evaluation has been carried out in the context of MultiWordNet, an ongoing project that aims at building a multilingual lexical database. The quantitative evaluation of the English-to-Italian lexical gaps shows that the English and Italian lexica are highly comparable and gives empirical support to the MultiWordNet model. 1.
No language in the world is homogeneous, or ever will be. Whereas earlier forms of English were characterised by extreme variation on all levels and Middle English is in fact best described as a loose conglomerate of unstable varieties, we usually lack any more detailed insight into what functions this variation had for the individual speaker. The social correlates so well known from modern sociolinguistics, such as age, sex, education, religion, can normally not be applied to the existing texts, nor can even the geographical range of recorded forms be determined with any degree of certainty. Finally, if modern dialect or other non-standard features are contrasted with (as the term non-standard implies) an accepted standard form of a language, this method would necessarily fail with Middle English even if we knew more about it than we do and, in view of the state of surviving documents, ever will. It is safe to assume that for its speakers the linguistic heterogeneity of Middle English was ordered in some way, but it was so only for continually shifting speech communities, whose number and individual geographical spread we know very little about. The scene changed dramatically in the fifteenth century: the emergence of a new standard language began to re-institute a linguistic norm for written supraregional English. This development was a natural consequence of the acceptance of English in public domains, and was speeded up by the change-over to English as the Chancery language in 1430.
International audience
This paper describes the methodology that is being used to augment the Penn Treebank annotation with sense tags and other types of semantic information. Inspired by the results of SENSEVAL, and the high inter-annotator agreement that was achieved there, similar methods were used for a pilot study of 5000 words of running text from the Penn Treebank. Using the same techniques of allowing the annotators to discuss difficult tagging cases and to revise WordNet entries if necessary, comparable inter-annotator rates have been achieved. The criteria for determining appropriate revisions and ensuring clear sense distinctions are described. We are also using hand correction of automatic predicate argument structure information to provide additional thematic role labeling. 1.
In this paper, we present a method for comparing Lexicalized Tree Adjoining Grammars extracted from annotated corpora for three languages: English, Chinese and Korean. This method makes it possible to do a quantitative comparison between the syntactic structures of each language, thereby providing a way of testing the Universal Grammar Hypothesis, the foundation of modern linguistic theories.
Bagging and boosting, two effective machine learning techniques, are applied to natural language parsing. Experiments using these techniques with a trainable statistical parser are described. The best resulting system provides roughly as large of a gain in F-measure as doubling the corpus size. Error analysis of the result of the boosting technique reveals some inconsistent annotations in the Penn Treebank, suggesting a semi-automatic method for finding inconsistent treebank annotations.
This article focuses on ongoing work done for Portuguese concerning the phenomenon of lexical co-occurrence known as collocation (cf. Cruse, 1986, inter al.). Instances of the syntactic variety formed by noun plus adjective have been especially observed. Collocational instances are not lexical entries, and thus should not be stored in the lexicon as multiword lexical units. Their processing can be conceived through relations linking the lexical components. Mechanisms for dealing with the collocation-hood of the expressions are required to be included in the systems, topographically, in their lexical modules. Lexical databases like wordnets, with a general architecture typically structured on semantic relations, make room for the specification of this phenomenon. This can be handled through the definition of ad-hoc relations expressing the different semantic effects the adjectival modification bring to nominal phrases, collocationally. 1
In this paper, we present a neural-networks-based knowledge discovery and data mining (KDDM) methodology based on granular computing, neural computing, fuzzy computing, linguistic computing, and pattern recognition. The major issues include 1) how to make neural networks process both numerical and linguistic data in a data base, 2) how to convert fuzzy linguistic data into related numerical features, 3) how to use neural networks to do numerical-linguistic data fusion, 4) how to use neural networks to discover granular knowledge from numerical-linguistic data bases, and 5) how to use discovered granular knowledge to predict missing data. In order to answer the above concerns, a granular neural network (GNN) is designed to deal with numerical-linguistic data fusion and granular knowledge discovery in numerical-linguistic databases. From a data granulation point of view, the GNN can process granular data in a database. From a data fusion point of view, the GNN makes decisions based on different kinds of granular data. From a KDDM point of view, the GNN is able to learn internal granular relations between numerical-linguistic inputs and outputs, and predict new relations in a database. The GNN is also capable of greatly compressing low-level granular data to high-level granular knowledge with some compression error and a data compression rate. To do KDDM in huge data bases, parallel GNN and distributed GNN will be investigated in the future.
The present study investigated the relationship between daily diary affect ratings and ambulatory cardiovascular activity in 117 male Vietnam combat veterans (61 with posttraumatic stress disorder [PTSD] and 56 without PTSD). Participants completed 12-14 hr of ambulatory monitoring and daily diary affect ratings. Compared with veterans without PTSD, veterans with PTSD reported higher negative affect and lower positive affect in daily diary ratings. No differences were detected for mean laboratory initial recordings or mean ambulatory heart rate (HR), systolic blood pressure (SBP), or diastolic blood pressure (DBP). However, compared with veterans without PTSD, veterans with PTSD demonstrated higher SBP and DBP variability and a higher proportion of HR activity (compared with initial recording values) during daily activity. There was a significant Time of Day x Group interaction for mean HR, with a trend for PTSD participants to maintain HR levels during evening hours.
A BILINGUAL LEXICAL DATABASE FOR FRAME SEMANTICS Get access Thierry Fontenelle Thierry Fontenelle 19 Rue du Merschgrund (L-8373 Hobscheid, Luxembourg)University of Liège(B-4000 Liege, Belgium) (fontenel@pt_lu) Search for other works by this author on: Oxford Academic Google Scholar International Journal of Lexicography, Volume 13, Issue 4, December 2000, Pages 232–248, https://doi.org/10.1093/ijl/13.4.232 Published: 01 December 2000
Bagging and boosting, two effective machine learning techniques, are applied to natural language parsing. Experiments using these techniques with a trainable statistical parser are described. The best resulting system provides roughly as large of a gain in F-measure as doubling the corpus size. Error analysis of the result of the boosting technique reveals some inconsistent annotations in the Penn Treebank, suggesting a semi-automatic method for finding inconsistent treebank annotations.
1.1 Notion of word....................................... 4 1.2 Tests of wordhood..................................... 5 1.3 Compatibility with other guidelines............................ 6
The value of language resources is greatly enhanced if they share a common markup with an explicit minimal semantics. Achieving this goal for lexical databases is difficult, as large-scale resources can realistically only be obtained by up-translation from pre-existing dictionaries, each with its own proprietary structure. This paper describes the approach we have taken in the Concede project, which aims to develop compatible lexical databases for six Central and Eastern European languages. Starting with sample entries from original presentation-oriented electronic representations of dictionaries, we transformed the data into an intermediate TEI-compatible representation to provide a common baseline for evaluating and comparing the dictionaries. We then developed a more restrictive encoding, formalised as an XML DTD with a clearly-defined semantic interpretation. We present this DTD and discuss a sample conversion from TEI, together with an application which hyperlinks a HTML represent...
We present a method for automatically detecting errors in a manually marked corpus using anomaly detection. Anomaly detection is a method for determining which elements of a large data set do not conform to the whole. This method fits a probability distribution over the data and applies a statistical test to detect anomalous elements. In the corpus error detection problem, anomalous elements are typically marking errors. We present the results of applying this method to the tagged portion of the Penn Treebank corpus.
Computer-driven systems for constructing composite faces of suspects (E-fit; Mac-a-Mug) have largely replaced mechanical systems (Photofit; the Identikit) in police use, yet little is known of their comparative effectiveness in rendering an accurate likeness. Participants (N = 24) constructed 2 of 4 familiar or unfamiliar faces, for one of which they used Photofit and for the other, E-fit. A likeness of each face was made first under target-absent conditions and then with photographs of the target present. The accuracy of the resulting composites was assessed by familiarity ratings, names elicited, and matching accuracy. The computer-driven system showed consistent superiority only when a familiar face was constructed in the presence of photographs; when participants worked from memory, E-fit was no better than Photofit. The implications of these findings for theories of face retrieval and the operational use of composites are discussed.
This paper describes the design criteria and annotation guidelines of Sinica Treebank. The three design criteria are: Maximal Resource Sharing, Minimal Structural Complexity, and Optimal Semantic Information. One of the important design decisions following these criteria is the encoding of thematic role information. An on-line interface facilitating empirical studies of Chinese phrase structure is also described.
The present study used the picture perception paradigm to examine the extent to which three well-documented psychophysiological measures demonstrate consistency across time in response to emotional stimuli. The three measures were the eye-blink startle response and the activation in two facial muscle regions (zygomatic and corrugator). Twenty-seven young women were assessed on two occasions, 2 weeks apart. Whereas activation in the corrugator and zygomatic muscle regions demonstrated the predicted patterns at both assessments (with some attenuation in the zygomatic muscle regions), the startle response had limited consistency across the two assessments. The startle response revealed the predicted linear pattern of valence modulation during the first assessment. During the second assessment, startle magnitude response was a quadratic function of valence ratings and a linear function of arousal ratings. The unexpected pattern of startle response during the second session appeared to be related to the content of the pleasant slides, with action slides generating quadratic valence modulation and erotic slides continuing to exhibit the expected linear valence modulation.
Article choice can pose difficult problems in applications such as machine translation and automated summarization. In this paper, we investigate the use of corpus data to collect statistical generalizations about article use in English in order to be able to generate articles automatically to supplement a symbolic generator. We use data from the Penn Treebank as input to a memory-based learner (TiMBL 3.0; We discuss competitive results obtained using a variety of lexical, syntactic and semantic features that play an important role in automated article generation.