Papers reviewed and determined not to be word norm studies. Use the flag icon to report errors or suggest re-inclusion.
16504 papers
Research suggests that ratings of child psychopathology by parents and teachers are generally not highly correlated. We examined the agreement and discordance between the child behaviour ratings of parents and teachers of a cohort of 6-year-old Pacific children living in New Zealand, based on scores from the Child Behaviour Checklist and the Teacher Report Form. Mother's reports were obtained for 1019 children, of whom, 602 also had father's reports and 559 had teacher's reports. Rater agreement was low between all pairs of informants. Fathers and teachers had higher agreement than mothers and fathers, the latter in turn had higher agreement than mothers and teachers, and agreement was generally higher for Externalizing problems than Internalizing problems. In terms of discordance, mothers reported more aggressive behaviour than fathers, while fathers reported more Internalizing and Total problems than mothers. Mothers and fathers generally reported more behaviour problems than teachers. The higher agreement found between informants from different settings (fathers and teachers) than between informants from similar settings (mothers and fathers) is in contrast with some of the literature. Further research is needed to investigate how child, informant, and setting characteristics affect ratings of children's behaviour.
The conditional collaborative filtering technology does not consider the user's context information which a-ffect user's rating.But recent research show that user's personalized context directly affect rating,so the result of recommendation can be improved if personalize context is incorporated into conditional collaborative filtering technology.Besides,the personalized context and item class are can be combined,Firstly classifying the items,and then making sure user's personalize context under every item class.When predicting the rating of target item,Firstly make sure which item class the target item is belong to,and then identify the user's personalized context used to compute the rating of the target item.The experimental results show that the recommendation accuracy of proposed approach is better than Slope One.
The purpose of this study was to evaluate the diagnostic quality of digital tomosynthesis (DT) images for pediatric imaging of the spine. We performed a phantom image rating study to assess the visibility of anatomical spinal structures in DT images relative to digital radiography (DR) and computed tomography (CT). We collected DT and DR images of the cervical, thoracic and lumbar spine using anthropomorphic phantoms. Four pediatric radiologists and two residents rated the visibility of structures on the DT image sets compared to DR using a four point scale (0 = not visible; 1 = visible; 2 = superior to DR; 3 = excellent, CT unnecessary). In general, the structures in the spine received ratings between 1 and 3 (cervical), or 2 and 3 (thoracic, lumbar), with a few mixed scores for structures that are usually difficult to see on diagnostic images, such as vertebrae near the cervical-thoracic joint and the apophyseal joints of the lumbar spine. The DT image sets allow most critical structures to be visualized as well or better than DR. When DR imaging is inconclusive, DT is a valuable tool to consider before sending a pediatric patient for a higher-dose CT exam.
The phenomenon of morphosyntactic variation in nominal groups was undertaken in the linguistic literature in respect of the theory of variant, the causes of language evolution or linguistic norm in Standard Polish. Regional Polish language in the East Borderland (Kresy) provides an interesting text corpus showing the variation in using the dative case and the dla (Eng. for) prepositional phrase for a much broader range than in today’s Standard Polish, e.g. powiedzieć dla (Prp) księdza (Gen) – SP powiedzieć księdzu (Dat) ‘to tell priest’; spodobali się dla (Prp) was (Gen) – SP spodobali się wam (Dat) ‘you liked them’; dla (Prp) mnie (Gen) pomagała – SP mnie (Dat) pomagała ‘she helped me’. The subject of the article is to analyse the determinants of alternation (variation) between the dative case and dla (Eng. for) prepositional phrase on the basis of homogeneous language material, which is 120 pages of text, recorded in one of Polish dialects in Khmelnytsky district in Ukraine. The aim was to investigate the possible functional and semantic limitations of the variants in the dialect. Statement of the frequency of the competing morphosyntactic markings in the text corpus has shown that the prepositional phrase does not eliminate the dative case in any of analysed functions. The analysis demonstrates that alternation includes all the functions of the dative case, both the dativus commodi and the dativus incommodi. In particular, this second context illustrates the specific usage of dla prepositional phrase, e.g. zniszczyć meble dla (Prp) kogoś (Gen), which are not acceptable as dativus incommodi in Standard Polish (SP zniszczyć meble komuś (Dat) ‘destroy one’s furniture’).
The problem of automatically extracting structured information from texts is an important, unsolved problem within the field of Natural Language Processing. The extraction of such information can facilitate activities such as the building of knowledge bases, automatic \nsummarisation and sentiment analysis. A human reader can easily discern the events described in a text, along with the participants and the relationships between them, \nbut using a computer to automatically discover the same information is much more challenging. Particular focus has been given to extracting relations between the entities in a text, such as those representing geographical locations, personal and social relationships, and employment. In this thesis, we consider two closely related entity relationships, which are interesting, frequent and have not been tackled previously, which we refer to collectively \nas entity instantiations. \nWe define an entity instantiation as an entity relation in which a set of entities is introduced, and either a member or subset of this set is mentioned. In the example below, \nwe see a set membership instantiation, between ‘several EU countries’ and ‘the UK’, along with a subset instantiation, between the same set and ‘the low countries’. Inflation has increased sharply in several EU countries. In the UK, this has accompanied a drop in interest rates, but in the low countries rates have remained steady. This thesis details the creation of the first corpus of entity instantiations. The final corpus consists of 4,521 instantiations, 2,118 of which are intersentential, and 2,403 of which are intrasentential, annotated over 75 Penn Treebank Wall Street Journal newswire texts. The subsequent annotation study shows high levels of inter-annotator agreement and our \ncorpus study analyses the annotated entity instantiations in terms of their internal structure, the distance between arguments and their syntactic relationship, finding a particularly strong link between syntactic parent-child relationships and sentence-internal entity instantiations. \nTo establish that the accurate automatic identification of entity instantiations is possible, we develop the first instantiation identification algorithm, which uses a supervised machine learning approach. The feature set draws on surface, syntactic, contextual, salience and knowledge features to aid classification. We separately apply our classifier to intersentential and intrasentential entity instantiations and experiment with both balanced data, with a 50/50 positive/negative split, and the original unbalanced corpus. The classifier records highly significant performance increases over both unigram-based \nand majority class baselines on the balanced data, and also on the original distribution of intrasentential instantiations. \nIn order to take advantage of the aforementioned link between syntax and intrasentential entity instantiations, tree kernels were employed to learn directly from the syntactic parse trees which contain the two potential participants in an intrasentential instantiation. \nThe tree kernel features perform similarly to the unstructured feature set, with a much shorter development time. Combining tree kernels with unstructured features gives further improvements over both the baselines, and either method in isolation. We also apply our entity instantiations to the difficult problem of implicit discourse relation classification, hypothesising that introducing features identifying the presence of an entity instantiation between the arguments of a discourse relation can improve classification performance. Our experiments show that an entity instantiation is a strong indicator of the presence of an Expansion.Instantiation discourse relation. We create a binary Expansion.Instantiation classifier, based on the feature set detailed in Sporleder \nand Lascarides (2008), but augment it by adding entity instantiation features based on gold standard annotations. The classifier which includes entity instantiation data performs significantly better than the same classifier without entity instantiation data. We also experiment with the incorporation of machine-identified entity instantiations. However, our entity instantiation classifier is not sufficiently accurate to impact on discourse relation classification.
To control negative emotion entails avoiding the harmful influences of bad mood, which may influence attention, memory, subjective and physical well-being, etc. Developing effective methods of negative emotion regulation are critical in improving mental health. The study of cognitive appraisal has been the recent focus of this pursuit. Cognitive appraisal is defined as a type of cognitive regulation that may eliminate negative feelings. While much evidence of cognitive appraisal has been reported, the studies often used inappropriate instructions and hence caused a confounding effect due to uncontrollable cognitive activity. For example, some researchers explicitly asked participants to try to reduce their emotional intensity by using reappraisal. As a result, participants would use unnecessary cognitive activities to decrease emotion, leading to the artificial inflation of appraisal effect. In this study, an improved method was used to solve this problem and probe only the function of appraisal on negative emotion. Three pieces of films, the length of which were all about six minutes, were chosen to elicit emotion. According to the emotional valence ratings, one of them was neutral while the other two were negative. Thirty-seven participants for the main experiment were instructed to watch the films with two physiological indexes being recorded: GSR (Galvanic Skin Reflex) and ECG (Electrocardiography). Before and after each film, the participants were asked to rest for four minutes. A rating for their current mood was also made before and after the clips. Different from previous studies, two distinct appraisals were given to two participant groups before the second negative film started, both asked the participants to watch the films naturally. Nineteen of the participants were told the actors’ own stories and emphasized they just performed. The rest, as a control group, were told the content in the film. At the end of the experiment, all participants were asked if they thought the film was fabled when watching the last clip to assess whether the appraisal background influenced their cognition of the film. The results indicated that only GSR and self emotion rating reflected emotional activity differences between the two groups. Analysis of covariance with the GSR level of the first rest as covariant indicated that the GSR level in the actors-appraisal group was lower than that in the control group when watching the second negative film. However, during the first two films, there were no differences between these two groups. On the other hand, analysis of covariance with the self report before the first rest as covariant indicated that the negative experience of actors–appraisal group was lower than that of the control group when watching the second negative film. During the first two films, there were no such differences. The change of GSR and negative experience, as anticipated, indicated that appraisal decreased physiological reaction to negative emotion. To sum up, people with the knowledge that the emotional stimulus was fabled showed more peaceful physiological activity along with lower negative emotion rating. These results indicate the effect of appraisal on emotion.
Public opinion polls have historically indicated that the US public favours domestic over global priorities. It is not known what influence health knowledge has in shaping public opinion about domestic and global health policy. This study examines how knowledge of HIV/AIDS is related to the rated importance of domestic and global health issues. Participants were recruited to participate in an electronic survey (N = 995) and were predominantly White (86.3%), married (61.9%) and female (71.8%). HIV/AIDS knowledge was significantly associated with both domestic (β = 0.12, p < 0.05) and global health (β = 0.14, p < 0.01) priorities after controlling for sociodemographic variables. In addition, global health was found to act as a mediator between HIV/AIDS knowledge and perceived importance of domestic issues. Study findings suggest that those with greater HIV/AIDS knowledge rate global health issues higher, which in turn affects ratings of more domestic issues. This research has implications for ways to gain support for implementation of public health policy through increasing health knowledge.
A selective tour of annotation in historical corpora begins with extra-linguistic markup: how far can it alert the corpus user to usage which is atypical of the variety being sampled? Several syntactic fossils are discussed, and a playful use of foreign and pseudo-foreign words. Are they a kind of code-switching? In the former case the answer No is given, in the latter a partial Yes. As for grammatical mark-up, with few exceptions a given scheme must privilege one particular analysis for each word, sentence or other unit of analysis. Special tags are available in the CLAWS and Penn Treebank tagsets for cases which remain ambiguous but which are in principle decidable. Grammatical mark-up remains essentially a matter of synchronic analysis, and the guiding principle is to be as specific as possible; tagsets routinely deploy a much finer set of distinctions than traditional word classes. Historical corpora like the Penn family aim also for consistency of analysis. I argue that both principles can be problematic. Consider first the push towards a unique POS tag for every word. I propose that certain kinds of word are vague as to their word class not because of a failure of analysis but because they are genuinely underdetermined. Vagueness is not ambiguity, so ambiguity tags would be inappropriate - at least with their currently intended values. Secondly, the desideratum of consistency does not allow for patterns which arguably have dual analyses synchronically, nor for items which are in transition or which have changed over the time-span of a historical corpus. Among the data discussed are the POS-tagging and parsing of adjectives derived from passive participles (interested, amused), multi-word prepositions (on behalf of), phrasal and prepositional verbs (run over), proper-to-common-noun conversions and noun-to-adjective transitions (BandAid), countable-to-mass conversions (He looked at me across a vast expanse of table) and the converse (two coffees). A brief conclusion argues that while some of the problems considered are statistically unimportant, others demand greater flexibility of mark-up.
This study has a dual focus in that it aims to develop a viable methodology for elicitation experiments in English linguistics, while simultaneously applying the proposed methods to investigate an actual subject, the distribtion of the additive particles 'also' and 'too'. Traditionally, data for linguistic research is gained by sampling natural language corpora. Although this approach is valid and, indeed, has been applied here, elicitation experiments can gain in validity and informative value by additionally introducing questionnaires to accompany corpus research. Online questionnaires particularly are a cost-effective and highly customizable tool to create a linguistic database against which existing data can be tested. For the purpose of this study, I have created six online questionnaires to test three hypotheses about the distribution of 'also' and 'too'. Two interdependent hypotheses assume that the use of the two particles is sensitive to structural properties of the `added constituent' while the third one, the information-structural hypothesis, argues that the use of 'also' and 'too' is controlled by the information structure of the sentence. In addition to the questionnaires, a balanced sample was extracted from the "British National Corpus" and tested against corpus data from previous studies as well as the data elicited online. In the course of this study, the additive particles will firstly be defined in terms of their structural properties, and the hypotheses about their use introduced and explicated. Furthermore, the data elicitation process will be detailed, as well as results from previous studies be taken into account. The hypotheses will subsequently be tested against the data from both corpus research and elicitation per questionnaires, and the outcome discussed. Concluding the study, I will focus on the results of the distribution analysis as well as evaluate the introduction of the online questionnaires and their application in the context of testing the hypotheses against empirical linguistic data.
Treebanks are a necessary prerequisite for many NLP tasks, including, but not limited to, semantic role labeling. For many languages, however, treebanks are either nonexistent or too small to be useful. Time-critical applications may require rapid deployment of natural language software for a new critical language—much faster than the development time of a traditional treebank. This dissertation describes a method for generating a treebank and training syntactic and semantic models using only semantic training information—that is, no human-annotated syntactic training data whatsoever. This will greatly increase the speed of development of natural language tools for new critical languages in exchange for a modest drop in overall accuracy. Using Combinatory Categorial Grammar (CCG) in concert with Propbank semantic role annotations allows us to accurately predict lexical categories in combination with a partially hidden Markov model. By training the Berkeley parser on our generated syntactic data, we can achieve SRL performance of 65.5% without using a treebank, as opposed to 74% using the same feature set with gold-standard data.
This paper deals with a group of entries in the Pralex lexical database, namely with the one letter abbreviations and acronyms, which form a specific and very diverse group. Firstly, the author shortly describes the concept of chosen abbreviations and acronyms; secondly, she focuses on the methodology of processing of this group of entries in the lexical database.
The article deals with the issue of regional lexical units (LU) processed in the Pralex lexical database (LDB), specifically with the processing of regionally marked lexical units in the rising dictionaries while defining regionalisms of different type in the Czech tradition and in the Pralex LDB. It also emphasizes the problems occurring during the processing of this type of LU in the LDB. Furthermore, it examines the shifts in our evaluation of regional LU in comparison with the previous dictionaries (Dictionary of the Standard Czech Language, Dictionary of Standard Czech, Academy Dictionary of Loanwords in Czech).
The first part of this paper deals with the concept of multi-word lexical units and the delimitation of their basic types in the Pralex lexical database, including the analysis of some terminological and conceptual issues. After a short preview of the approaches to the lexicographical treatment of multi-word lexical units in contemporary monolingual dictionaries, the second part of this paper deals with the database processing of multi-word lexical units, both with the general rules of their treatment in the Pralex lexical database and with the specific rules of treatment of one of two basic types of multi-word lexical units, i.e. multiple-word namings.
The article deals with the treatment of hypocorisms in the Pralex lexical database. The core of the article is the description of the work with exemplifications (sentential examples).
This paper analyses a Quebec comic strip, Magasin general, in which language is both an element of group cohesion and of explicit thematisation. Starting from the problematic relationship Quebec has with French linguistic norm, we try to investigate the use of language in the BD itself, the discourse about language by the characters and by a large paratextual apparatus, and also the reception of this comic by the Quebec highbrow press.
The paper deals with the macrostructure of the Pralex lexical database and with the basic principles of building an index, mainly with the choice of sources and with criteria of the choice of lexical units including possibilities of their future completion from own sources of the Institute.
The article deals with the program and content aspects of the treatment of the Czech lexis in the form of a lexical database. It summarises the development of the Praled software, the basic features of the macro- and microstructures of the Pralex LDB and the main principles of the treatment of database items.
In this paper, we give a summary of various dependency chart parsing algorithms in terms of the use of parsing histories for a new dependency arc decision. Some parsing histories are closely related to the target dependency arc, and it is necessary for the parsing algorithm to take them into consideration. Each dependency treebank may have some unique characteristics, and it requires for the parser to model them by certain parsing histories. We show in experiments that proper selection of the parsing algorithm which reflect the dependency annotation of the coordinate structures improves the overall performance. 1
The paper deals with compounds in the Pralex lexical database, especially with their general characteristics and problems with their processing in the database. Apart from the compounds, the author focuses on the particular components and their origin, meanings and mutual relations. Compounds with a quantitative meaning as well as compounds containing such component are not handled in here.
The article describes the development of the typographical practice of placing spaces between words in the early printed Glagolitic books and the swift decline in the 16 th century of the use of so-called word-blocks in favour of full word separation.Little studied but signifi cant for our understanding of a host of writing and reading practices, ranging from the rhetorical and compositional features of the medieval Croatian Church Slavonic texts to the linguistic norms of early modern works, the use of the white space consistently increases over time as modern typographical practices take hold.In the context of widespread changes in mechanical printing practices in the 15 th -17 th centuries, the paper looks at the increasing use of white space between all words, primarily in the CrCS printed liturgical books and on the parallel decrease in the use of word-blocks (zdruenice).Given the larger movement toward regularization of liturgical texts in the 16 th century and the growing awareness of linguistic science our examination of typesetting practices offers some insights into the implementation of regularized linguistic norms for the Glagolitic liturgical books and makes it possible to conclude that the more widespread typesetting practices of the secular presses quickly gained a foothold in the ecclesiastical printeries.There was, moreover, a rapid conformity to the Western typographical practice of separating words as the smallest units of independent meaning; i.e. in accordance with our own contemporary practices.
The article focuses on the contacts between prescriptive and universal grammar in 18th century British linguistics. The author argues that, contrary to the wide-spread belief, there existed two-way relations between them: on the one hand, practical prescriptive grammar used the principles of universal grammar as the foundation for singling out parts of speech and grammar categories and based normative recommendations upon similarities between languages. On the other hand, the authors of universal grammars showed interest to problems of linguistic norm and gave practical recommendations in the vein of prescriptive tradition.
The paper deals with abbreviations, i.e. shortenings, acronyms and abbreviated expressions; it interprets and internally divides this special group of lexical units. The chosen typology was applied for a development of the database Praled. Following the particular types of units, the author monitors possible processing ways of various types of abbreviated forms, while building the Pralex lexical database.
In this paper we aim to automatically identify subjects, which are not expressed but nevertheless understood in Czech sentences. Our system uses the maximum entropy method to identify different types of unstated subjects and the system has been trained and tested on the Prague Dependency Treebank 2.0. The results of our experiments bring out further consideration over the suitability of the chosen corpus for our task.
The reconstruction of standardized texts in the Prague Dependency Treebank of Spoken Czech enables the comparison of authentic spoken utterances and standardized texts The authors concentrate on the questions: What does the syntactic identity of the Czech spoken and written texts consist of? What syntactic constructions are „natural“ in the spoken and in the written text? What is the difference in the density of the cohesive links, in the explicit and implicite relations between units?
Introduction Brain-Computer Interfaces (BCI) can be used for communication and motor restoration (Birbaumer & Cohen, 2007). To our knowledge, no BCI study looked at patients with dementia who have severe communication deficits. BCIs based on operant training could be problematic for patients with cognitive deficits. A paradigm shift from instrumental-operant learning to classical conditioning could possibly overcome this failure (Birbaumer, 2006). Recent findings demonstrated the possibility to classify cognitive and emotional states by the pattern classification of BOLD signals in both offline (Lee et al., 2010a, 2010b) and online situations (Sitaram et al., 2010). The present study aims to investigate the feasibility of an auditory classical conditioning paradigm within a fMRI based BCI setting. The paradigm is designed to condition individuals to associate positive and negative emotional stimuli as unconditioned stimuli (US) with congruent and incongruent word pairs as conditioned stimuli (CS), respectively. Our goal is to ascertain whether the brain signals pertaining to congruent and incongruent word pairs could be classified with more than chance accuracy using our fMRI support vector machine (SVM) with a view to apply for basic online yes/no communication in Alzheimer patients. Methods The paradigm consisted of one single session divided into six blocks, comprising the different phases of conditioning (habituation, acquisition, extinction). The US consisted of auditory emotional stimuli selected from the International Affective Digitized Sounds (IADS, Bradley & Lang, 1999). A segment of baby laughter represented the positive emotional stimulus and a segment of screaming represented the negative emotional stimulus. The CS, presented aurally, were congruent (e.g. ‘animal-elephant’) and incongruent (e.g. ‘animal-Germany’) word-pairs. The unconditioned and conditioned responses (UR and CR) were the changes in the BOLD signal pertaining to the CS and US, respectively. The first block consisted of a randomized presentation of 50 US and 50 CS. In the second and third blocks 25 congruent word pairs, immediately followed by the baby laughter, and 25 incongruent word pairs, immediately followed by the scream, were presented randomly. In the fourth and fifth blocks, respectively 40% and 20% of the CS were paired with the US. In the sixth block, only the CS was presented. Functional imaging was performed continuously during these blocks on 6 healthy subjects (4 females, 2 males, age 21-27) on a 3.0 T scanner (Siemens, Germany). To classify the signals corresponding to various conditions, namely congruent and incongruent word-pairs, a linear SVM (with the regularization parameter, C=1) was implemented. Classification performance from data was evaluated through 2-fold cross validation (CV). Based on the parameters of the trained SVM model, we analyzed the fMRI data with the Effect Mapping method (EM; Lee et al., 2010a, 2010b, Sitaram et al., 2010). To investigate the relative importance of different brain regions in decoding the conditioning brain states, feature vectors from the frontal cortex were used as input to build a separate SVM classifier. Results The Self Assessment Manikin (SAM) test showed that participants reported more negative valence and a higher arousal for the scream compared to the baby laughter. Classification of the BOLD signal as a response to the congruent and incongruent word pairs immediately followed by the emotional US showed above chance level performance (57-64%) on one subject, around chance level (50-56%) performance on three subjects, and below chance level (44-47%) performance on two subjects. Conclusions In this pilot study we have demonstrated an approach for conditioning the BOLD signal by repeated association of the emotional stimuli with semantic stimuli, resulting in a paradigm for basic yes/no communication. Further work includes improving the performance of the classifier by feature selection, an online implementation of the system, and its testing on patients. References: Birbaumer, N. (2006), ‘Brain-computer-interface research: coming of age’, Clinical Neurophysiology, vol. 117, pp. 479-483. Birbaumer, N. & Cohen, L. G. (2007), ‘Brain-computer interfaces: communication and restoration of movement in paralysis’, The Journal of Physiology, vol. 579, no. 3, pp. 621-636. Bradley, M. M. & Lang, P. J. (1999), ‘International Affective Digitized Sounds (IADS): Stimuli, instruction manual and affective ratings’, University of Florida, Gainesville. Sitaram R, Lee S, Ruiz S, Rana M, Veit R, Birbaumer N. Real-time support vector classification and feedback of multiple emotional brain states. Neuroimage, 2010 Aug 6. Lee, S., Halder, S., Kübler, A., Birbaumer, N., Sitaram, R. Effective functional mapping of fMRI data with support-vector machines. Hum Brain Mapp. 2010a, Jan 28. Lee, S., Ruiz, S., Caria, A., Birbaumer, N., Sitaram, R Cerebral reorganization induced by real-time fMRI feedback training of the insular cortex: a multivariate investigation. Neuroreh and Neural Rep (2010b).
Nous montrons comment enrichir une annotation en dependances syntaxiques au format du French Treebank de Paris 7 en utilisant la reecriture de graphes, en vue du calcul de sa representation semantique. Le systeme de reecriture est compose de regles grammaticales et lexicales structurees en modules. Les regles lexicales utilisent une information de controle extraite du lexique des verbes francais Dicovalence.
Based on the assertions of the Theory of Linguistic Variation and Change, this paper proposes a discussion about the possible action of linguistic norms over two variable phenomena in Brazilian Portuguese: the position of clitic pronouns associated with a single verb, and the use of prepositions with verbal complements indicating a?goal/recipient?. By the analysis of data from the newspapers from São Paulo and Rio Claro between (the years of) 1900 and 1915, we intend to describe each phenomenon; compare these descriptions and evaluate the role played by the standard and the common usage which is (already) perceptible in the?paulistas? continuous published pages of that period.
The movement of people and their languages on an unprecedented scale has been exerting increasing pressure on the model of the nation state and the ideal of socially and linguistically homogeneous societies. Sociolinguists analyzing recent changes in language policies and citizenship legislation have focused on the global-local interface, as well as issues of territoriality and group membership, thus connecting with cultural geographers and anthropologists who study the relationship between language and senses of place. Although frequently regarded as the domain of linguistics, language policies and practices are inseparable from issues of power and identity. Scholarly investigations have demonstrated the need to view the dynamics of language contact in relation to “external” linguistic dimensions that are central to the field of sociology, including social class, gender, and ethnicity, together with acts of compliance with or resistance to social and linguistic norms.
The article proposes a reading of two novels by Evgenij Popov (b. 1946), Nakanune nakanune (1993) and Podlinnaja istorija «Zelenych muzykantov» (1999), as a response to “the language question” in post-perestroika Russian culture and society. The heterogeneous linguistic landscape of Nakanune – a remake of Turgenev's classic Nakanune – is compared to the “theoretical” reflections on language and style found in the numerous footnotes contained in Podlinnaja istorija. The comparative view opens up for a discussion of two principal types of metalanguage and the differences and interactions between them: explicit commentary on language and linguistic reflexivity through linguistic practice. For the latter, I propose the term “performative metalanguage”, and point to the numerous ways in which linguistic norms and styles may be negotiated, challenged and commented on through the linguistic practice, and in combination with straightforward commentary.
This paper is to analyze curricular changes of Chongryon Korean schools in Japan. Chongryon Korean schools belong to the category of miscellaneous schools in the classification by the Ministry of Education in Japan. They neither need to follow curricular set by the Japanese government nor receive subsidies from it. They have their own curricula and textbooks. Historically, there have been 6 curricular reforms in Chongryon Korean schools. After characterizing those reforms, this paper compares the ``Korean`` textbooks of the 1993 and 2003 reforms. In general, the political color of advocating socialism and the Juche idea and anti-American and anti-Seoul propoganda gets thinner in new textbooks. The 2003 Korean textbooks emphasize speaking practice and adapt dialogues rich with story-telling. This paper also examines linguistic norms followed in the Korean textbooks. (Osaka University of Economics and Law)
This paper presents our preliminary work on adaptation of parsing technology toward natural language query processing for biomedical domain. We built a small treebank of natural language queries, and tested a state-of-theart parser, the results of which revealed that a parser trained on Wall-Street-Journal articles and Medline abstracts did not work well on query sentences. We then experimented an adaptive learning technique, to seek the chance to improve the parsing performance on query sentences. Despite the small scale of the experiments, the results are encouraging, enlightening the direction for effective improvement. 1
Music processing may be preserved in subjects with Alzheimer disease (AD). It is not known which neural substrates are engaged in music processing, and how music familiarity moderates the engagement of these substrates in AD. We investigated fMRI patterns of brain activation during listening to familiar and non-familiar classical music excerpts in subjects with mild to moderate AD and healthy age-matched controls. We related these patterns to behavioral data on familiarity ratings and musical abilities. Five subjects with AD (age M = 76.2, SD = 6.6, MMSE M = 18.6, SD = 7.7, range 9-26) and five healthy controls (age M72.8, SD = 8.0, MMSE M = 29.6, SD =.6, range 29-30) underwent fMRI with a block design paradigm consisting of 75s of familiar music excerpts followed by 30s white noise, vs. 75s of unfamiliar music excerpts. Participants were instructed to just listen. They received a battery of behavioral tests including repeated familiarity ratings of the presented music excerpts (1=very familiar to 5=very unfamiliar), the Montreal Battery for Evaluation of Amusia (MBEA), and the Seashore test of musical abilities. For fMRI a mixed-model 2x2 ANOVA was used to examine effect of group AD vs. controls) and stimulus (familiar vs. unfamiliar) using a p-value <.01 and a minimum cluster size of 200ÂμL. For behavioural data, independent- and paired-sample t-tests were used with p-value <.05. We found a significant group by stimulus effect in fMRI activation patterns. When familiar to unfamiliar music activations were compared, the following regions showed increases in AD and decreases in controls: left/right lingual gyrus, left inferior parietal, left superior, middle and inferior temporal gyri, left precuneus, left culmen, left/right striatum. Ratings for familiar and unfamiliar excerpts did not differ by group (AD M = 1.2, SD =.2 and M = 1.9, SD =.6; Controls M = 1.2, SD =.3, and M = 1.9, SD =.3). Performance on music ability tests also did not differ by group except for Seashore loudness and rhythm (p <.05). Subjects with AD appear to respond more intensely to non-familiar than familiar music by activating regions associated with recognizing familiar patterns and emotions. They do not differ from controls on behavioral measures. This finding suggests differential neural recruitment with respect to preserved music recognition, and its potential application in diagnosis and treatment.
In this paper, we present a series of enhancements to the English-Japanese version of a multilingual Linguistics Based Machine Translation system, for the translation of complex sentences, modality and complex verbal structures. The system is using a classical transfer-based architecture and dedicated lexical databases. Relying on linguistic data acquired on large corpora or compiled from the web, corrections have been done in constituent reordering, lexical selection and verb conjugation. Even if the system is not, so far, as efficient as state-of-the-art English-Japanese MT systems, the results show a clear progress and underline the interest of using syntactic information in MT.
In the present paper we focus on control as a subtype of anaphora. We work with the theory of control present within the dependency-based framework of Functional Generative Description (FGD), in which control is defined as a relation of a referential dependency between a controller (antecedent � semantic argument of the main clause) and a controllee (anaphor � empty subject of the nonfinite complement (controlled clause)). First this paper presents the rule-based reconstruction of controllees, then, it discusses the perceptron-based determination of the controllees� antecedent. We evaluated our approach on data from the Prague Dependency Treebank 2.0, however, the rules and features of our system are supposed to be language independent and can be tested on other languages in the future.
The core of the processing of individual items in the lexical database Pralex is an example part of an entry (exemplification). While processing the nouns, we try to select a relevant evidence from the corpus SYN, i.e. examples of a concrete use of meanings processed in advance according to SSJČ and, in addition to that, we also try to record evidence of the new meanings. In the exemplification, we focus on a syntactic and semantic linkage. Lemmas are mainly illustrated by means of non-quotational, modified evidences, i.e. example linkage in the form of syntagmas, which are regarded as common collocations of the processed nouns. Sentential records are introduced as well, though in a lower extent.
For bilingual students, attending a university away from home serves as a separation from one's native linguistic norms to an entirely new set of language practices. These U.S. institutions of higher learning almost always demand English in favor of other native or heritage languages. This study examines the ways a small group of second- and third- generation Lithuanian immigrants attempt to maintain their heritage language. It also looks at how language loss can manifest within a college environment. The research suggested the difficulties and challenges college life poses on the retention of a heritage language. Discussions for further research include: how University policy might implicitly prescribe English, use of heritage language among friend as a marker of kinship, the difficulty of resisting English monolingualism within the U.S. public school system, and the loss of ethnic customs and identity in sync with language loss.
Objective To explore the brain mechanism of arousal abnormal in emotional processing in abstinent heroin abstainer.Methods Totally 13 heroin abstainers(experimental group) and 13 healthy subjects(control group) underwent fMRI when viewing pictures with different arousal levels.Then the arousal ratings were evaluated by the subjects.fMRI data were processed with AFNI software.Behavioral data were analyzed with SPSS 13.0 software.Results Looking at emotional pictures,compared with control subjects,brain activation of experimental group decreased in left amygdale and hippocampus,bilateral thalamus,cingulate gyrus,bilateral superior frontal gyrus and middle frontal gyrus,right precentral gyrus,bilateral superior temporal gyrus,right inferior tempotal gyrus,left fusiform gyrus and bilateral caudate nucleus.ROI analysis indicated increased brain activation in experimental group when watching high arousal stimuli than low arousal stimuli in right thalamus,left amygdale and hippocampus,while the controls showed the opposite pattern.Conclusion Heroin abusers have abnormalities in arousal processing.
This paper presents a novel approach for resolving ambiguities in concepts that already reside in semantic databases such as Freebase and DBpedia. Different from standard dictionaries and lexical databases, semantic databases provide a rich hierarchy of semantic relations in ontological structures. Our disambiguation approach decides on the implied sense by computing concept similarity measures as a function of semantic relations defined in ontological graph representation of concepts. Our similarity measures also utilize Wikipedia descriptions of concepts. We performed a preliminary experimental evaluation, measuring disambiguation success rate and its correlation with input text content. The results show that our method outperforms well-known disambiguation methods.
In article the known problem of language ontogenesis is considered in the light of the new knowledge received for last decades in sphere of such sciences, as historical linguistics, glottochronology, archaeological anthropology, anthropological genetics, geology and biology. The contradictions arising by comparison of received knowledge are presented. The hypothesis of mental monogenesis is offered which allows to remove a part of existing contradictions. The new approaches are offered based on use of linguistic databases, created recently, for carrying out of the further researches on a problem of ontogenesis.
This paper introduces, an XML format developed to serialise the object model defined by the ISO Syntactic Annotation Framework SynAF. Based on widespread best practices we adapt a popular XML format for syntactic annotation, TigerXML, with additional features to support a variety of syntactic phenomena including constituent and dependency structures, binding, and different node types such as compounds or empty elements. We also define interfaces to other formats and standards including the Morpho-syntactic Annotation Framework MAF and the ISOCat Data Category Registry. Finally a case study of the German Treebank TueBa-D/Z is presented, showcasing the handling of constituent structures, topological fields and coreference annotation in tandem.
This paper provides a critical appraisal of the methods used in a research project that investigates the linguistic participation of women in the Scottish Parliament, the Northern Ireland Assembly and the National Assembly for Wales. The project adopted a mixed-method approach by firstly undertaking an ethnographic description of the linguistic norms and practices in each institution. This process included observations in the debating chambers and interviews with politicians. Secondly, discourse analytic techniques (drawing upon both CDA and CA) were used to analyse close transcripts of debates. Finally, a quantitative study was undertaken which involved classifying and counting speaking turns across as sample of debates in each institution. Here I reflect on the use of this mixed method approach, both in terms of the practical applicability to the field, and in terms of the effects on the project results. I conclude that mixed-method research can be extremely useful in new research sites as ‘scoping’ studies because different methods can illuminate particular aspects of institutional practices. However, it is often difficult to integrate the different types of data made available by disparate methods into a unified interpretative framework.
The proposed study investigated an attentional bias in experimentally-induced dysphoria using self-relevant pictures in a dot probe study design. Participants generated their own photographs, using digital cameras to capture stimuli that are self-relevant and emotional to them. It was hypothesized that individuals with induced dysphoria would exhibit a greater attentional bias to negative stimuli than participants with induced happiness when self-generated pictures were used in a dot-probe paradigm. In addition, exploratory analyses were conducted to examine possible gender effects. To examine the first hypothesis, a MANOVA was conducted including the priming groups and gender as the predictors and the bias scores as the dependent variables. Results did not support the primary hypothesis. Regarding gender effects, females responded longer on all trials, and the interaction of gender and priming condition neared significance for negative attentional bias scores. It was also hypothesized that the importance and valence ratings of the pictures would significantly predict response latency, but this prediction was not supported by the data. The findings of this study are discussed in terms of cognitive theories of attentional biases in depression as well as methodological issues in this line of work.
This article deals with a processing of pronouns in the Pralex lexical database. Pronouns can be homonymous with other word types, e.g. with nouns or adverbs. Grammatical data for polysemic pronouns are introduced at the level of a full entry or at the level of individual meanings. Within every lexical unit, the author records a pronoun type as well as its primary syntactic function. Sometimes, identification of a word type can be problematic, e.g. the pronouns který (who) and jaký (which), which are used to link a subordinate clause, but have no relative function at all. When defining an explanation of a meaning, we rely on Dictionary of the Standard Czech Language and Dictionary of Standard Czech for Schools and the General Public. An example part of entries is divided in the so-called modified and sentential records.
Communicative structure is central to the linguistic representation at nearly all levels of the Meaning-Text Models (MTMs). Its correlation with lexical and syntactic features makes it also essential for such natural language processing applications as text generation, which is about to undergo a significant shift from the symbolic, rule-based paradigm to the statistical paradigm. In the statistical paradigm, the availability of sufficiently large corpora annotated with linguistic information, and thus also with the communicative structure (CommStr), is critical. However, to the best of our knowledge, so far no corpora annotated with CommStr in the sense of the Meaning-Text Theory are available. We describe two experiments that explore how such corpora can be obtained. In the first experiment, a fragment of a Spanish Treebank is annotated manually. In the second experiment, we exploit the correlation of CommStr with syntactic features to annotate the English PropBank.
En este artículo se presenta la introducción de clases semánticas de WordNet en un analizador de dependencias, obteniendo mejoras en el Penn Treebank completo por primera vez. Probamos diferentes combinaciones de algunas clases semánticas básicas y algoritmos de desambiguación del sentido de las palabras. Nuestros experimentos muestran que seleccionar la combinación adecuada de características semánticas en los datos de desarrollo es clave para el éxito. Dada la naturaleza básica de las clases semánticas y los algoritmos de desambiguación del sentido de las palabras utilizados, creemos que hay un amplio margen para futuras mejoras.
Many approaches to web malware detection tend to consider obfuscated scripts as malicious, although it is has been demonstrated that obfuscator does not indicate malice. In a bid to distinguish obfuscation techniques used in malicious and benign scripts, we propose a subtree matching technique to identify learned structural patterns in analyzed scripts. Our proposal implements two techniques from the realm of natural language processing and string algorithms. We advocate the use of abstract syntax trees to reduce the entropy introduced by string randomization and rather focus on code structure patterns. In a learning phase, we discover frequently occurring trees in abstract syntax treebanks, while in the testing phase, we attempt to identify these trees within a candidate AST by using a pushdown automata that accepts the set of learned trees.
Psycholinguistic studies on whether classifiers facilitate processing object-extracted relative clauses (RC) in Mandarin have often made use of a classifier mismatch-match configuration, wherein a preceding classifier mismatches the following RC-subject but matches the modified head noun. However, an examination of the Chinese Treebank corpus 5.0 shows this configuration rarely occurs. None of the 10 tokens of pre-RC classifiers conforms to the mismatch-match configuration in a real sense. Instead, either a dropped RC-subject or some intervening item successfully avoids anticipated lexical disruption effects induced by a mismatching classifier. The results of analysis suggest that the constructed examples used in previous psycholinguistic studies may not realistically test natural language processing procedures.
Many recent experiments in the automatic classification of discourse relations have limited themselves to a small set of coarse categories. While there are eminent reasons to do so – annotation for a small sets of categories can be created more reliably, and possibly also be defined in a more clear – cut way than finer distinctions – it is an interesting question whether the finer-grained distinctions present in some annotated corpora can be reconstructed reliably. The present paper investigates the feasibility of such fine-grained tagging of discourse relations using data from the Penn Discourse Treebank. 1
The conventional sequence labeling methods for Chinese word segmentation do not fully utilize the linguistic information, which restricts further improvements of the performance. Chinese morphology intensively investigates the constructions and usages of Chinese words, which is helpful to Chinese word segmentation. Furthermore, some word segmentation ambiguities cannot be resolved only by means of the lexical information, and the final disambiguations take place in the parsing process. In this paper, we propose a parsing-based Chinese word segmentation model, which can fully utilize the morphological and syntactic information. Experiments on Penn Chinese Treebank(CTB) 5.0 show that the proposed model obtains competitive performances as the CRFs-based model. To investigate the relationship between our parsing-based model and the CRFs-based model, a maximum entropy model based framework for integrating different knowledge sources is employed. The integrating model obtains an F-measure of 97.9, 25% in segmentation error rate reduction relative to the CRFs-based model, which indicates that the two models are complementary to each other.