Papers reviewed and determined not to be word norm studies. Use the flag icon to report errors or suggest re-inclusion.
18265 papers
This chapter studies the norms from a merely formal viewpoint, in order to concentrate exclusively on its logical-linguistic structure. The problem of legal norms making reference to theories of language and its different functions is a course of study which is used more in Italy. The chapter focuses on language in general, though spoken and written language as the most prominent form. From the formal-linguistic perspective, the legal norm is a proposition, which is to say it is a sequence of words endowed with significance. This differs from statements, as a lexical-syntactic group of linguistic signs with which a proposition is expressed. From the point of view or prism of the functions of language, it can be said that the norm is a or preceptive proposition. The chapter discusses on the various possible types of propositions, paying special attention to the so-called prescriptive or preceptive propositions.Keywords: legal norm; logical-linguistic structure; prescriptive proposition
Background: Implicit racial bias denotes socio-cognitive attitudes towards other-race groups that are exempt from conscious awareness. In parallel, other-race faces are more difficult to differentiate relative to own-race faces - the ''Other- Race Effect.'' To examine the relationship between these two biases, we trained Caucasian subjects to better individuate other-race faces and measured implicit racial bias for those faces both before and after training. Methodology/Principal Findings: Two groups of Caucasian subjects were exposed equally to the same African American faces in a training protocol run over 5 sessions. In the individuation condition, subjects learned to discriminate between African American faces. In the categorization condition, subjects learned to categorize faces as African American or not. For both conditions, both pre- and post-training we measured the Other-Race Effect using old-new recognition and implicit racial biases using a novel implicit social measure - )
Se référant à divers ouvrages et auteurs de tous temps, styles ou nationalités, cette étude sur la représentation du substrat dialectal et étranger dans la littérature française et anglo-américaine, et sur sa traduction, cherche principalement à comprendre la démarche des auteurs recourant à la retranscription phonétique, pour ensuite mieux appréhender celle des traducteurs. Analysant en premier lieu les diverses motivations qui poussent ces écrivains à bouleverser les normes grammaticales et orthographiques pour transfigurer dans l’écrit, l’oral et la parole, puis s’interrogeant quant à la validité d’une méthode à appliquer à ces créations linguistiques, cet exposé tente de répondre, notamment par un examen énumératif des procédés matérialisant l’accent dialectal ou étranger, aux questions d’ordre lexical, grammatical ou morphosyntaxique qu’infère cette intrusion de la langue parlée dans le texte. Elucidant enfant les outils traductologiques mis en place dans ces littératures, ce travail propose un traducteur futur de pleinement s’en inspirer, et ainsi faire de l’achoppement, un argument à la créativité.
We employ a single-trial correlational MEG analysis technique to investigate early processing in the visual recognition of morphologically complex words. Three classes of affixed words were presented in a lexical decision task: free stems (e.g., taxable), bound roots (e.g., tolerable), and unique root words (e.g., vulnerable, the root of which does not appear elsewhere). Analysis was focused on brain responses within 100-200 msec poststimulus onset in the previously identified letter string and visual word-form areas. MEG data were analyzed using cortically constrained minimum-norm estimation. Correlations were computed between activity at functionally defined ROIs and continuous measures of the words' morphological properties. ROIs were identified across subjects on a reference brain and then morphed back onto each individual subject's brain (n = 9). We find evidence of decomposition for both free stems and bound roots at the M170 stage in processing. The M170 response is shown to be sensitive to morphological properties such as affix frequency and the conditional probability of encountering each word given its stem. These morphological properties are contrasted with orthographic form features (letter string frequency, transition probability from one string to the next), which exert effects on earlier stages in processing ( approximately 130 msec). We find that effects of decomposition at the M170 can, in fact, be attributed to morphological properties of complex words, rather than to purely orthographic and form-related properties. Our data support a model of word recognition in which decomposition is attempted, and possibly utilized, for complex words containing bound roots as well as free word-stems.
We present the first data-driven dependency parser for Romanian, which has been developed using the MaltParser system and trained and evaluated on a dependency treebank for Romanian developed within the RORIC-LING project. The parser achieves a labeled attachment score of 88.6 % (unlabeled 92.0%) when evaluated on held-out data from the treebank. We present a partial error analysis, focusing on accuracy for different parts of speech and dependencies of different length.
This paper studies the nature of the BEI-construction in Cantonese, with Mandarin as the standard language of comparison. Although the BEI-construction has been much studied in Mandarin, the same in not true for Cantonese. Although this construction has traditionally been termed a "passive", I will show that it can have a different range of semantic interpretations in Cantonese. I argue that BEI is not confined to passive, but is used under certain circumstances to form a causative construction as well. The differences in behaviour between passive-BEI and causative-BEI can be seen in tests with anaphoric binding. I conclude that while the passive structure is mono-clausal, the causative structure must be bi-clausal. The Cantonese BEI-constructions have an obligatory agent-phrase which cannot be dropped. This differs from Mandarin and the challenge is to find an account for this phenomenon, especially if we are to claim that this construction is a passive. The optionality of the agent phrase is characteristic of passives and yet Cantonese deviates from this norm. I argue that passive in Cantonese is a syntactic process and predict that only transitive verbs may participate in this construction. I utilize the universal v-VP structure on transitive verbs, proposed by Chomsky (1995), to guarantee that the external theta role must be retained. I also examine the much debated status of BEI which is used in the BEI-construction. Although this construction can be used to derive both a passives and a causatives, it does not necessarily mean that two separate BEIs must be posited. I conclude that BEI can be treated as a category-neutral element which can interact in both causative and passive structures. To support this proposal I appeal to the functional versus lexical distinction of categories and projections.
Rapid eye movement (REM) sleep and dreaming may be implicated in cross-night adaptation to emotionally negative events. To evaluate the impact of REM sleep deprivation (REMD) and the presence of dream emotions on a possible emotional adaptation (EA) function, 35 healthy subjects randomly assigned to REMD (n = 17; mean age 26.4 +/- 4.3 years) and control (n = 18; mean age 23.7 +/- 4.4 years) groups underwent a partial REMD and control nights in the laboratory, respectively. In the evening preceding and morning following REMD, subjects rated neutral and negative pictures on scales of valence and arousal and EA scores were calculated. Subjects also rated dream emotions using the same scales and a 10-item emotions list. REMD was relatively successful in decreasing REM% on the experimental night, although a mean split procedure was applied to better differentiate subjects high and low in REM%. High and low groups differed - but in a direction contrary to expectations. Subjects high in REMD% showed greater adaptation to negative pictures on arousal ratings than did those low in REMD% (P < 0.05), even after statistically controlling sleep efficiency and awakening times. Subjects above the median on EA(valence) had less intense overall dream negativity (P < 0.005) and dream sadness (P < 0.004) than subjects below the median. A correlation between the emotional intensities of the morning dream and the morning picture ratings supports a possible emotional carry-over effect. REM sleep may enhance morning reactivity to negative emotional stimuli. Further, REM sleep and dreaming may be implicated in different dimensions of cross-night adaptation to negative emotions.
English is spoken worldwide by both native (L1) and nonnative (L2) speakers. It is therefore imperative to establish how easily L1 and L2 speakers understand each other. We know that L1 listeners adapt to foreign-accented speech very rapidly (Clarke & Garrett, 2004), and L2 listeners find L2 speakers (from matched and mismatched L1 backgrounds) as intelligible as native speakers (Bent & Bradlow, 2003). But foreign-accented speech can deviate widely from L1 pronunciation norms, for example when adult L2 learners experience difficulties in producing L2 phonemes that are not part of their native repertoire (Strange, 1995). For instance, Italian L2 learners of English often lengthen the lax English vowel /I/, making it sound more like the tense vowel /i/ (Flege et al., 1999). This blurs the distinction between words such as bin and bean. Unless listeners are able to adapt to this kind of pronunciation variance, it would hinder word recognition by both L1 and L2 listeners (e.g., /bin/ could mean either bin or bean). In this study we investigate whether Italian-accented English interferes with on-line word recognition for native English listeners and for nonnative English listeners, both those where the L1 matches the speaker accent (i.e., Italian listeners) and those with an L1 mismatch (i.e., Dutch listeners). Second, we test whether there is perceptual adaptation to the Italian-accented speech during the experiment in each of the three listener groups. Participants in all groups took part in the same cross-modal priming experiment. They heard spoken primes and made lexical decisions to printed targets, presented at the acoustic offset of the prime. The primes, spoken by a native Italian, consisted of 80 English words, half with /I/ in their standard pronunciation but mispronounced with an /i/ (e.g., trick spoken as treek), and half with /i/ in their standard pronunciation and pronounced correctly (e.g., treat). These words also appeared as targets, following either a related prime (which was either identical, e.g., treat-treat, or mispronounced, e.g., treek-trick) or an unrelated prime. All three listener groups showed identity priming (i.e., faster decisions to treat after hearing treat than after an unrelated prime), both overall and in each of the two halves of the experiment. In addition, the Italian listeners showed mispronunciation priming (i.e., faster decisions to trick after hearing treek than after an unrelated prime) in both halves of the experiment, while the English and Dutch listeners showed mispronunciation priming only in the second half of the experiment. These results suggest that Italian listeners, prior to the experiment, have learned to deal with Italian-accented English, and that English and Dutch listeners, during the experiment, can rapidly adapt to Italian-accented English. For listeners already familiar with a particular accent (e.g., through their own pronunciation), it appears that they have already learned how to interpret words with mispronounced vowels. Listeners who are less familiar with a foreign accent can quickly adapt to the way a particular speaker with that accent talks, even if that speaker is not talking in the listeners’ native language.
In respect to the sexual differences that exist in language,this paper summarizes the features of female language.In terms of linguistic structure,such differences are shown at the phonetic,lexical and syntactic levels.In terms of extralinguistic structure,they can be shown by women's strategy and purpose of speaking,the quantity of their speech and their topics of conversation.Such differences not only imply women's special characters and physiological and psychological features,but also embody the prevailing social norms and cultural psychology.
Background: One of the most debated issues in the cognitive neuroscience of language is whether distinct semantic domains are differentially represented in the brain. Clinical studies described several anomic dissociations with no clear neuroanatomical correlate. Neuroimaging studies have shown that memory retrieval is more demanding for proper than common nouns in that the former are purely arbitrary referential expressions. In this study a semantic relatedness paradigm was devised to investigate neural processing of proper and common nouns. Methodology/Principal Findings: 780 words (arranged in pairs of Italian nouns/adjectives and the first/last names of well known persons) were presented. Half pairs were semantically related (''Woody Allen'' or ''social security''), while the others were not (''Sigmund Parodi'' or ''judicial cream''). All items were balanced for length, frequency, familiarity and semantic relatedness. Participants were to decide about the semantic relatedness of t)
The article discusses a study which investigated whether subtitles, which provide lexical information, support perceptual learning about foreign speech. As explained, subtitles indicate which words are being spoken, which could improve lexically-guided learning about foreign speech sounds. As part of the study, Dutch participants were asked to watch videos containing unfamiliar regionally-accented English, with or without subtitles. Both English and Dutch subtitles were used which made it possible to compare the effects of subtitles in the language spoken in the videos with the effects of subtitles in the observers' native language.
This article introduces the topic of “Multilingual language resources and interoperability”. We start with a taxonomy and parameters for classifying language resources. Later we provide examples and issues of interoperatability, and resource architectures to solve such issues. Finally we discuss aspects of linguistic formalisms and interoperability.
Background: Studies of experts' problem-solving abilities have shown that experts can attend to the deep structure of a problem whereas novices attend to the surface structure. Although this effect has been replicated in many domains, there has been little investigation into such effects in medicine in general or patient management in particular. Methodology/Principal Findings: We designed a 10-item forced-choice triad task in which subjects chose which one of two hypothetical patients best matched a target patient. The target and its potential matches were related in terms of surface features (e.g., two patients of a similar age and gender) and deep features (e.g., two diabetic patients with similar management strategies: a patient with arthritis and a blind patient would both have difficulty with self-injected insulin). We hypothesized that experts would have greater knowledge of management categories and would be more likely to choose deep matches. We contacted 130 novices (medical)
This paper examines the ideologies and practices surrounding respect at a Korean American heritage language school in California. It illustrates the interaction between locally circulating metadiscourses about children’s dispositions, intentions, and identities and the enforcement of classroom norms of respect. In some cases, teachers accommodated to children’s linguistic norms though a metadiscourse that reframed the indexicality of potentially disrespectful behavior. In other cases, forms of bodily demeanor were naturalized as indexical of children’s deliberate communication of disrespect. Teachers’ classroom narratives presented theories of affective accommodation and affective display, where respect for a teacher’s feelings was supposed to be given priority over respect for a child’s feelings, but children did not always comply with these theories. By illustrating how teachers’ metapragmatic ideologies about children’s identities as Korean Americans, contexts of language acquisition, and linguistic needs mediate the interactional construction of (dis)respect, this paper demonstrates the hybrid/multidirectional nature of language socialization. (PsycINFO Database Record (c) 2016 APA, all rights reserved)
Methods for discriminant analysis were compared with respect to classification accuracy under nonnormality through Monte Carlo simulation. The methods compared were linear discriminant analyses based both on raw scores and on ranks; linear logistic discrimination; and mixture discriminant analysis. Linear discriminant analysis and linear logistic discrimination were suboptimal in a number of scenarios with skewed predictors. Linear discriminant analysis based on ranks yielded the highest rates of classification accuracy in only a limited number of situations and did not produce a practically important advantage over competing methods. Mixture discriminant analysis, with a relatively small number of components in each group, attained relatively high rates of classification accuracy and was most useful for conditions in which skewed predictors had relatively small values of kurtosis.
Word Sense Disambiguation is the most critical issue in natural language processing. Although it has been addressed by many researchers, no satisfactory results are reported. Rule based systems alone can not handle this issue due to ambiguous nature of the natural language. Knowledge-based systems are therefore essential to find the intended sense of a word form. Machine readable dictionaries have been widely used in word sense disambiguation. The problem with this approach is that the dictionary entries for the target words are very short. WordNet is the most developed and widely used lexical database for English. The entries are always updated and many tools are available to access the database on all sorts of platforms. The WordNet database can be converted in MySQL format and we have modified it as per our requirement. Sense's definitions of the specific word, "Synset" definitions, the "Hypernymy" relation, and definitions of the context features (words in the same sentence) are retrieved from the WordNet database and used as an input of our Disambiguation algorithm.
We present a valency lexicon for Latin verbs extracted from the Index Thomisticus Treebank, a syntactically annotated corpus of Medieval Latin texts by Thomas Aquinas. In our corpus-based approach, the lexicon reflects the empirical evidence of the source data. Verbal arguments are induced directly from annotated data. The lexicon contains 432 Latin verbs with 270 valency frames. The lexicon is useful for NLP applications and is able to support annotation. 1
Dolgozatomban, amint arra a címből is lehet következtetni, az 1996 és 2005 között adatolható magyarországi börtönszlenget mutatom be, azt a csoportnyelvet, amelynek az átfogó tanulmányozása hazánkban eddig még nem történt meg. Munkámban büntetés-végrehajtási intézeteink fogvatartottjainak belső, informális nyelvhasználatának általános kérdéseivel foglalkozom, és a mai magyar börtönszleng szó- és kifejezéskészletének szótárba foglalásán túl kísérletet teszek a vizsgált csoportnyelv nyelvi-szociolingvisztikai leírására. \n \nCélkitűzésemet, a magyar börtönszleng átfogó tanulmányozását az indokolta, hogy a kutatás első éveiben olyan mennyiségű és minőségű, a nyelvtudományban eddig még nem tárgyalt adatokra bukkantam, amelyek érdemesnek mutatkoztak arra, hogy egy mélyebb, megtervezett szlengkutatás irányuljon erre a területre. \n \nElsősorban célom volt a magyar börtöszlenget feltárni, bemutatni keletkezését, funkcióját, működését, a szlenghasználó közösségben betöltött szerepét. Célom volt továbbá rámutatni nyelvi előzményeire, összevetni a már létező bűnözői nyelvi adatbázissal, vagyis a tolvajnyelv elemeivel, egyben definiálni helyét a magyar szlengkutatás területén. Kutatásom során mindvégig azt tartottam szem előtt, hogy hol van az ember a szlengben, így célom volt annak leírása is, mikor, milyen körülmények között motiváltak a vizsgált csoport tagjai szlenghasználatra. Ennek kiderítéséhez a nyelvi adatok feltárásán és rendszerezésén túl, a zárt közeg csoportjainak vizsgálatára is ki kellett terjeszteni a kutatást, megfigyelve a csoportszerveződés lehetőségeit és okait a börtöntársadalomban. \n \nAs already the title has suggested, my dissertation presents Hungarian prison slang as attested from 1996 to 2005. This group language has never seen an overall study in Hungary up to now. I deal with general questions of the internal and informal language use of the prisoners of penal institutions in my work and, together with rendering the words and expressions into a dictionary, I attempt at the linguistic and sociolinguistic description of the group language examined. \n \nMy objective, the overall study of Hungarian slang, is justified by the data being of such quantity and quality and having never been dealt with in linguistics that seemed worth to be examined by a deeper and planned slang research. \n \nMy main objective was to explore Hungarian prison slang, to present its origins, functions and operation as well as its role within slang user communities. My aim also was to show its linguistic predecessors, to compare it to an existing linguistic database of criminals, that is, to the elements of cant, and, together with it, to define its place in the area of Hungarian slang research. During my studies, I always kept man’s role in slang in mind so my objective was the description of the conditions among which the members of a researched group are motivated for slang usage. In order to learn about it, I had to extend my research to the examination of the groups of closed space, observing the possibilities and reasons within prison community.
The lexical development system OntoNet is introduced, which includes a browser and an editor for the WordNet 3.0 database. The aim of the OntoNet project is to provide a comfortable and up-to-date access to the lexical database for the modification of WordNet or the development of new wordnets.
In recent years, the specter of litigants turning to religious or customary sources of law as authoritative guides to regulate their behavior, alongside or in lieu of secular norms, has risen to the forefront of politics in many countries worldwide. In this essay, we draw upon citizenship theory and comparative constitutional jurisprudence to identify two different categories of judicial response to religious-based claims for recognition, accommodation, and exemption: 1) 'diversity as inclusion;' and 2) 'non-state law as competition.' As long as legal claims for accommodation are not seen by courts as challenging the lexical superiority of the constitutional religion itself ('diversity as inclusion'), they stand a fair chance of success. Contrast that with the unyielding reluctance of legislatures and judiciaries to accept as binding or even cognizable any potentially competing legal order that originates in sacred or customary sources of identity and authority. This pattern of clamping down and refusing to accept any alternative sources of regulation becomes particularly visible where the legal challenge at issue is interpreted as raising doubts regarding which set of norms and institutions, or what set of high priests, should have the final word in authoritatively resolving legal disputes within a given society ('non-state law as competition'). This is a challenge that no secular legal order, no matter how tolerant and otherwise open to providing exemptions and accommodations to religious believers, can accept with indifference. For what perceived to be at stake here is the very authority and source of legitimacy of the accepted civil religion. We demonstrate these claims by focusing on recent jurisprudence from Canada and South Africa, two polities that represent the most difficult cases for our argument; if there is any place we would expect to find recognition by secular countries of religious or customary sources of law and authority, it would be in these diverse societies that have made an explicit constitutional commitment to promote their citizens’ freedom to preserve and enhance their multitude of backgrounds and distinctive cultural, linguistic and religious heritages as part of their 'mosaic' (Canada) or 'rainbow nation' (South Africa) conceptions of citizenship. Although operating in different contexts, the South African Constitutional Court and the Supreme Court of Canada seem to have made every effort to subject traditional legal regimes to general principles of constitutional law. By so doing, they have erected a new wall of separation that places noncompliance with the values of the civil religion beyond the pale of accepted accommodation, offering to those who espouse them the potential to either bring these alternative legal domains under the general rule of constitutional law or encounter the wrath of state fiat.
The basic concept of semantic Web,ontology and semantic annotation are described.Then the semantic annotation technology and tool today are introduced and analyzed,and a way of automatic semantic annotation based on HTML documents that contain rich semantic data on the Web is presented.This method couples structural analysis of documents with semantic analysis incorporating domain ontologies and lexical database Hownet,discovers the semantic partition tree corresponding to documents,and annotates HTML documents with semantic lables.The experiment is based on the HTML documents of electronic products,the result shows the method is feasible.
The aim of the project entitled "Computerized Historical Linguistic Database of Latin Inscriptions of the Imperial Age" (http://lldb.elte.hu/) is to develop and digitally publish a fundamental computerized historical linguistic database that incorporates and treats the Vulgar Latin material of the Latin inscriptions from a specific group of the European provinces of the Roman Empire in the first phase. This will, on the one hand, allow for a more thorough study of the regional changes and the diversity of the Latin language of the Imperial Age. On the other hand, it could also serve as a basis for subsequent international co-operation, in the course of which further work on the computerized historical linguistic database may be executed. This paper intends to present the past and the present, as well as the future possibilities of this Database.
We describe the Hindi Discourse Relation Bank project, aimed at developing a large corpus annotated with discourse relations. We adopt the lexically grounded approach of the Penn Discourse Treebank, and describe our classification of Hindi discourse connectives, our modifications to the sense classification of discourse relations, and some crosslinguistic comparisons based on some initial annotations carried out so far.
Corpora present a basic informational source for varied NLP applications. Their construction becomes now days necessary. This paper presents a tool created to help syntactic tagging an Arabic Treebank. It is based on an already constructed grammar called ArabTAG which constitutes a representational method of Arabic grammatical rules using TAG formalism. This tagging tool is a node among a set of other pre-treatment steps as: morpho-syntactic analysis, grammatical tagging and sentence segmentation. It gets as input a sentence and helps to give it the appropriate syntactic tree in an incremental manner.
We describe for dependency parsing an annotation adaptation strategy, which can automatically transfer the knowledge from a source corpus with a different annotation standard to the desired target parser, with the supervision by a target corpus annotated in the desired standard. Furthermore, instead of a hand-annotated one, a projected treebank derived from a bilingual corpus is used as the source corpus. This benefits the resource-scarce languages which haven't different handannotated treebanks. Experiments show that the target parser gains significant improvement over the baseline parser trained on the target corpus only, when the target corpus is smaller.
The first edition of a speech corpus with a speech reconstruction layer (edited transcript). The project of speech reconstruction of Czech and English has been started at UFAL together with the PIRE project in 2005, and has gradually grown from ideas to (first) annotation specification, annotation software and actual annotation. It is part of the Prague Dependency Treebank family of annotated corpus resources and tools, to which it adds the spoken language layer(s).
The present paper outlines an ongoing project of annotation of the extended nominal coreference and the bridging anaphora in the Prague Dependency Treebank. We describe the annotation scheme with respect to the linguistic classification of coreferential and bridging relations and focus also on details of the annotation process from the technical point of view. We present methods of helping the annotators -- by a pre-annotation and by several useful features implemented in the annotation tool. Our method of the inter-annotator agreement is focused on the improvement of the annotation guidelines; we present results of three subsequent measurements of the agreement.
영한 기계번역에서 영어 단어의 품사결정은 번역할 문장에 사용된 어휘의 품사 모호성을 해소하기 위해 필요하다. 어휘의 품사 모호성은 구문 분석을 복잡하게 하고 정확한 번역을 생성하는 것을 어렵게 한다. 본 논문에서는 이러한 문제점을 해결하기 위해 어휘 분석 이후 구문 분석 이전에 품사 모호성을 해소하려 하였으며 품사 모호성을 해소하기 위한 CatAmRes 모델을 제안하고 다른 품사태깅 방법과 성능 비교를 하였다. CatAmRes는 Penn Treebank 말뭉치를 이용하여 Bayesian Network를 학습하여 얻은 확률 분포와 말뭉치에서 나타나는 통계 정보를 이용하여 영어 단어의 품사를 결정을 한다. 본 논문에서 제안한 영어 품사결정 모델 CatAmRes는 결정할 품사의 적정도 값을 계산하는 Calculator와 계산된 적정도 값에 근거하여 품사를 결정하는 POSDeterminer로 구성된다. 실험에서는 CatAmRes의 동작과 성능을 테스트 하기 위해 WSJ, Brown, IBM 영역의 말뭉치에서 추출한 테스트 데이터를 이용하여 품사결정의 정확도를 평가하였다.
The aim of the present paper was to study heart rate changes during a video stimulation depicting two actors (male and female) producing dynamic facial expressions of happiness, sadness, and a neutral expression. We measured ballistocardiographic emotion-related heart rate responses with an unobtrusive measurement device called the EMFi chair. Ratings of subjective responses to the video stimuli were also collected. The results showed that the video stimuli evoked significantly different ratings of emotional valence and arousal. Heart rate decelerated in response to all stimuli and the deceleration was the strongest during negative stimulation. Furthermore, stimuli from the male actor evoked significantly larger arousal ratings and heart rate responses than the stimuli from the female actor. The results also showed differential responding between female and male participants. The present results support the hypothesis that heart rate decelerates in response to films depicting dynamic negative facial expressions. The present results also support the idea that the EMFi chair can be used to perceive emotional responses from people while they are interacting with technology.
This paper presents Thai syntactic resource: Thai CG treebank, a categorial approach of language resources. Since there are very few Thai syntactic resources, we designed to create treebank based on CG formalism. Thai corpus was parsed with existing CG syntactic dictionary and LALR parser. The correct parsed trees were collected as preliminary CG treebank. It consists of 50,346 trees from 27,239 utterances. Trees can be split into three grammatical types. There are 12,876 sentential trees, 13,728 noun phrasal trees, and 18,342 verb phrasal trees. There are 17,847 utterances that obtain one tree, and an average tree per an utterance is 1.85.
OBJECTIVES: The present study in hypomanic and manic patients explored how amygdala responses to affective stimuli depend on the valence of the stimuli presented. METHODS: We compared 10 patients with 10 matched healthy control subjects. We measured blood oxygen level-dependent (BOLD) responses in the amygdala while subjects passively viewed photographs taken from the International Affective Picture System. After the fMRI session, subjects saw the pictures again and subjectively rated the emotional valence and intensity of each picture. RESULTS: Compared to healthy individuals, hypomanic or manic patients showed higher valence ratings in positive pictures and associated larger BOLD responses in the left amygdala during positive versus neutral picture viewing. This enhanced amygdala activation was correlated with Young Mania Rating Scale scores and with euphoric as opposed to irritable symptom presentation. CONCLUSIONS: Increased valence ratings and amygdala responses to positive affective stimuli may reflect a positive processing bias contributing to elevated mood states characteristic for euphoric mania.
Most treebank work in the past has focused on European and Asian languages. The Wikipedia Treebank page lists treebanks (or treebank projects) for about 20 modern European languages (ranging from Basque to Swedish), five Asian languages (Chinese, Japanese, Hindi, Korean, Thai), two ancient languages (Greek and Latin), plus Arabic and Hebrew. Almost no treebanking work has been done on African or American indigenous languages.1 In the past we have explored parallel treebanks for English, German and Swedish [7]. Now we would like to explore to what extent our tools and guidelines will work when we include a very different language, Quechua, for which only few NLP resources exist. Since Quechua is spoken in Latin America, Spanish as parallel language is a natural choice. We have first compiled a parallel corpus Quechua Spanish. We have then stepwise analyzed and annotated the Quechua and the Spanish texts. For Spanish we have used the treebanking guidelines developed by [8]. As for Quechua there were no such guidelines so that we had to experiment with finding the appropriate grammar formalism and develop our own guidelines. In this paper we describe the characteristics of Quechua and our steps towards its morphological and syntactic annotation. We argue for Role and Reference Grammar as a suitable grammar formalism. We briefly describe how we annotated the parallel Spanish texts and demonstrate how we plan to align the Quechua with the Spanish trees.
Transition-based approaches have shown competitive performance on constituent and dependency parsing of Chinese. State-of-the-art accuracies have been achieved by a deterministic shift-reduce parsing model on parsing the Chinese Treebank 2 data (Wang et al., 2006). In this paper, we propose a global discriminative model based on the shift-reduce parsing process, combined with a beam-search decoder, obtaining competitive accuracies on CTB2. We also report the performance of the parser on CTB5 data, obtaining the highest scores in the literature for a dependency-based evaluation.
Aiming at exploring the possibility of increasing the parsing accuracy by linguistic means,an experiment of Chinese dependency parsing is conducted by using MaltParser and a self-built treebank.Through the detailed analysis for the parsing results,the possible suggestion about improving the performance of the parser is provided and it is used as the guidance to modify the annotation scheme of the treebank.Experimental results show that the accuracy of unlabeled dependency attachment score increases 5.5%,and the accuracy of labeled score raises 7.5%.
We consider linguistic database summaries in the sense of Yager (1982), in an implementable form proposed by Kacprzyk & Yager (2001) and Kacprzyk, Yager & Zadrozny (2000), exemplified by, for a personnel database, “most employees are young and well paid” (with some degree of truth) and their extensions as a very general tool for a human consistent summarization of large data sets. We advocate the use of the concept of a protoform (prototypical form), vividly advocated by Zadeh and shown by Kacprzyk & Zadrozny (2005) as a general form of a linguistic data summary. Then, we present an extension of our interactive approach to fuzzy linguistic summaries, based on fuzzy logic and fuzzy database queries with linguistic quantifiers. We show how fuzzy queries are related to linguistic summaries, and that one can introduce a hierarchy of protoforms, or abstract summaries in the sense of latest Zadeh’s (2002) ideas meant mainly for increasing deduction capabilities of search engines. We show an implementation for the summarization of Web server logs.
Parallel treebanks provide a systematic way of expressing the structural relationships between source and target texts. In this paper, we present the general design principles behind the Copenhagen Dependency Treebanks, a set of parallel treebanks for Danish, English, German, Italian and Spanish with a unified annotation of morphology, syntax, discourse, and tranlational equivalence. Finally, we suggest some hypotheses about morphology and discourse, and describe how we plan to explore them empirically on the basis of the treebanks.
In this paper we present initial results on parsing Arabic using treebank-based parsers and automatic\nLFG f-structure annotation methodologies. The Arabic Annotation Algorithm (A3) (Tounsi et al., 2009) exploits the rich functional annotations in the Penn Arabic Treebank (ATB) (Bies and Maamouri, 2003; Maamouri and Bies, 2004) to assign LFG f-structure equations to trees. For parsing, we modify Bikel’s (2004) parser to learn ATB functional tags and merge phrasal categories with functional tags in the training data. Functional tags in parser output trees\nare then "unmasked" and available to A3 to assign f-structure equations. We evaluate the resulting\nf-structures against the DCU250 Arabic gold standard dependency bank (Al-Raheb et al., 2006). Currently we achieve a dependency f-score of 77%.
Standard English PoS-taggers generally involve tag-assignment (via dictionary-lookup etc) followed by tag-disambiguation (via a context model, e.g. PoS-ngrams or Brill transformations). We want to PoS-tag our Arabic Corpus, but evaluation of existing PoS-taggers has highlighted shortcomings; in particular, about a quarter of all word tokens are not assigned a fully correct morphological analysis. Tag-assignment is significantly more complex for Arabic. An Arabic lemmatiser program can extract the stem or root, but this is not enough for full PoS-tagging; words should be decomposed into five parts: proclitics, prefixes, stem or root, suffixes and postclitics. The morphological analyser should then add the appropriate linguistic information to each of these parts of the word; in effect, instead of a tag for a word, we need a subtag for each part (and possibly multiple subtags if there are multiple proclitics, prefixes, suffixes and postclitics). Many challenges face the implementation of Arabic morphology, the rich “root-and-pattern” nonconcatenative (or nonlinear) morphology and the highly complex word formation process of root and patterns, especially if one or two long vowels are part of the root letters. Moreover, the orthographic issues of Arabic such as short vowels ( ), Hamzah (ء أ إ ؤ ئ), Taa’ Marboutah ( ة ) and Ha’ ( ه ), Ya’ ( ي ) and Alif Maksorah( ى ), Shaddah ( ) or gemination, and Maddah ( آ ) or extension which is a compound letter of Hamzah and Alif ( أا ). Our morphological analyzer uses linguistic knowledge of the language as well as corpora to verify the linguistic information. To understand the problem, we started by analyzing fifteen established Arabic language dictionaries, to build a broad-coverage lexicon which contains not only roots and single words but also multi-word expressions, idioms, collocations requiring special part-of-speech assignment, and words with special part-of-speech tags. The next stage of research was a detailed analysis and classification of Arabic language roots to address the “tail” of hard cases for existing morphological analyzers, and analysis of the roots, word-root combinations and the coverage of each root category of the Qur’an and the word-root information stored in our lexicon. From authoritative Arabic grammar books, we extracted and generated comprehensive lists of affixes, clitics and patterns. These lists were then cross-checked by analyzing words of three corpora: the Qur’an, the Corpus of Contemporary Arabic and Penn Arabic Treebank (as well as our Lexicon, considered as a fourth cross-check corpus). We also developed a novel algorithm that generates the correct pattern of the words, which deals with the orthographic issues of the Arabic language and other word derivation issues, such as the elimination or substitution of root letters.
The part-of-speech determination is necessary for resolving the part-of-speech ambiguity in English-Korean machine translation. The part-of-speech ambiguity causes high parsing complexity and makes the accurate translation difficult. In order to solve the problem, the resolution of the part-of-speech ambiguity must be performed after the lexical analysis and before the parsing. This paper proposes the CatAmRes model, which resolves the part-of-speech ambiguity, and compares the performance with that of other part-of-speech tagging methods. CatAmRes model determines the part-of-speech using the probability distribution from Bayesian network training and the statistical information, which are based on the Penn Treebank corpus. The proposed CatAmRes model consists of Calculator and POSDeterminer. Calculator calculates the degree of appropriateness of the partof-speech, and POSDeterminer determines the part-of-speech of the word based on the calculated values. In the experiment, we measure the performance using sentences from WSJ, Brown, IBM corpus.
Discovering frequent structures within large natural language corpora is one of the core problems of corpus linguistics, but it is difficult to do for richly structured data. This paper describes a practical algorithm to extract frequent structures from treebanks or annotated corpora that can be represented as a tree structures. It extracts the most frequent structures first, so that not all structures have to be counted in order to find the most frequent ones. This algorithm assumes random constant-time access to all parts of the treebank and has space and time bounds broadly proportionate to the size of the output, which is not readily predictable in most cases. It is efficient enough to be usable with reasonable sized corpora using conventional desktop workstations.