Papers reviewed and determined not to be word norm studies. Use the flag icon to report errors or suggest re-inclusion.
18265 papers
Specificity of the number of nouns in Czech and its annotation in Prague Dependency Treebank The paper focuses on the way how the grammatical category of number of nouns will be annotated in the forthcoming version of Prague Dependency Treebank (PDT 3.0), concentrating on the peculiarities beyond the regular opposition of singular and plural. A new semantic feature closely related to the category of number (so-called pair/group meaning) was introduced. Nouns such as ruce ‘hands’ or klíče ‘keys’ refer with their plural forms to a pair or to a typical group even more often than to a larger amount of single entities. Since pairs or groups can be referred to with most Czech concrete nouns, the pair/group meaning is considered as a grammaticalized meaning of nouns in Czech. In the present paper, manual annotation of the pair/group meaning is described, which was carried out on the data of Prague Dependency Treebank. A comparison with a sample annotation of data from Prague Dependency Treebank of Spoken Czech has demonstrated that the pair/group meaning is both more frequent and more easily distinguishable in the spoken than in the written data.
This paper describes the development, composition, and several uses of the Ancient Greek and Latin Dependency Treebanks, large collections of Classical texts in which the syntactic, morphological and lexical information for each word is made explicit. To date, over 200 individuals from around the world have collaborated to annotate over 350,000 words, including the entirety of Homer’s Iliad and Odyssey, Sophocles’ Ajax, all of the extant works of Hesiod and Aeschylus, and selections from Caesar, Cicero, Jerome, Ovid, Petronius, Propertius, Sallust and Vergil. While perhaps the most straightforward value of such an annotated corpus for Classical philology is the morphosyntactic searching it makes possible, it also enables a large number of downstream tasks as well, such as inducing the syntactic behavior of lexemes and automatically identifying similar passages between texts.
The article proposes a reading of two novels by Evgenij Popov (b. 1946), Nakanune nakanune (1993) and Podlinnaja istorija «Zelenych muzykantov» (1999), as a response to “the language question” in post-perestroika Russian culture and society. The heterogeneous linguistic landscape of Nakanune – a remake of Turgenev's classic Nakanune – is compared to the “theoretical” reflections on language and style found in the numerous footnotes contained in Podlinnaja istorija. The comparative view opens up for a discussion of two principal types of metalanguage and the differences and interactions between them: explicit commentary on language and linguistic reflexivity through linguistic practice. For the latter, I propose the term “performative metalanguage”, and point to the numerous ways in which linguistic norms and styles may be negotiated, challenged and commented on through the linguistic practice, and in combination with straightforward commentary.
The movement of people and their languages on an unprecedented scale has been exerting increasing pressure on the model of the nation state and the ideal of socially and linguistically homogeneous societies. Sociolinguists analyzing recent changes in language policies and citizenship legislation have focused on the global-local interface, as well as issues of territoriality and group membership, thus connecting with cultural geographers and anthropologists who study the relationship between language and senses of place. Although frequently regarded as the domain of linguistics, language policies and practices are inseparable from issues of power and identity. Scholarly investigations have demonstrated the need to view the dynamics of language contact in relation to “external” linguistic dimensions that are central to the field of sociology, including social class, gender, and ethnicity, together with acts of compliance with or resistance to social and linguistic norms.
The paper deals with abbreviations, i.e. shortenings, acronyms and abbreviated expressions; it interprets and internally divides this special group of lexical units. The chosen typology was applied for a development of the database Praled. Following the particular types of units, the author monitors possible processing ways of various types of abbreviated forms, while building the Pralex lexical database.
Unsupervised parsing induction has attracted a significant amount of attention over the last few years. However, current systems exhibit a degree of complexity that can shy away newcomers to the field. We challenge the need for such complexity and present a straightforward weak-EM based system. The results we obtained are close to state-of-the-art ones while still making it extremely simple to experiment with different sub-components. We use a k-best parser, an inductor for Probabilistic Bilexical Grammars (PBGs) [1] and a simple treebank builder. Since our algorithm is independent of the PBG inductor, it overlaps with other models from the literature such as Dependency Model with Valence [2]. Our algorithms are fully fleshed and easily reproducible. We experiment in 8 languages that inform intuitions in training- size dependent parameterization.
The paper focuses on proper names and, more specifically, personal names, toponyms and microtoponyms with a high degree of cultural embeddedness. The author draws on the classification of translation rules and methods (techniques) developed by Ermolovich, with a special focus on difficult cases of transfer where transliteration and transcription are employed. Two directions of translation are discussed: Polish into Russian and Russian into Polish. The paper analyses discrepancies between linguistic norms and usage, where problems originate from the existence of recognised equivalents, arbitrariness of transcription and globalisation (widespread use of the Internet). In the face of these challenges, consistency in a translation project is difficult to retain. The author proposes a set of practical exercises that can be used during translation classes (drawing on the analysis of original texts) in order to make beginner translation students aware of the existing problems and potential solutions. Examples of Polish realia in Russian texts and Russian realia in Polish texts are presented.
This paper considers how implemented grammars can enhance descriptive ones. An implemented grammar encodes the analyses in machine readable form, facilitating automatic annotation of morphological, syntactic and semantic structures. A grammar augmented with a collection of such structures would enable, for example, a reader to search for items in which a PP argument fills the third most prominent semantic role.
Therapist self-disclosure has been theorized and found to have both positive and negative effects. These effects depend, in part, on the nature of the disclosure. This study sought to examine the differential effects of therapist disclosures of more and less resolved countertransference issues on perceptions of therapists and therapy sessions. Using an analogue method, undergraduate participants (N = 116) were randomly assigned to watch one of two videos in which a therapist disclosed personal issues that were relatively resolved or relatively unresolved. As hypothesized, therapist disclosure of issues that were more resolved caused the therapist to be rated as more attractive and trustworthy and instilled greater hope than therapist disclosure of less resolved issues. The type of therapist disclosure, however, did not affect ratings of the expertness of the therapist, the depth or smoothness of the session, or the perceived universality between client and therapist. Implications of the results for the judicious use of self-disclosure are discussed.
The Internet has rarely been used in auditory perception studies due to concerns about standardisation and calibration across different systems and settings. However, not all auditory research is based on the investigation of fine-grained differences in auditory thresholds. Where meaningful ‘real-world’ listening, for instance the perception of speech, is concerned, the Internet may be a more appropriate and ecologically valid setting to collect data. This study compared affective ratings of low-pass-filtered infant-, foreigner- and British adult-directed speech obtained with traditional methods in the laboratory, with those obtained from an Internet sample. Dropout rates and demographic distribution of participants in the Internet condition were also assessed. The results show that affective ratings were similar for both the Internet and laboratory samples. These findings indicate the viability of Internet-based research into affective speech perception and suggest that precise acoustic environmental control may not always be necessary.
Manually performed treebanking is an expensive effort compared with automatic annotation.In return, manual treebanking is generally believed to provide higherquality/value syntactic annotation than automatic methods.Unfortunately, there is little or no empirical evidence for or against this belief, though arguments have been voiced for the high degree of subjectivity in other levels of linguistic analysis (e.g.morphological annotation).We report a double-blind annotation experiment at the level of dependency syntax, using a small Finnish corpus as the analysis data.The results suggest that an interannotator agreement can be reached as a result of reviews and negotiations that is much higher than the corresponding labelled attachment scores (LAS) reported for stateof-the-art dependency parsers.
”[Der findes ikke abstract til denne artikel]”
From two corpus studies into varieties of clausal coordination in English (Meyer, 1995 and Greenbaum & Nelson, 1999), it is known that the incidence of clausal coordinate ellipsis (CCE) is about two times higher in written than in spoken language. We present a treebank study into CCE in written and spoken Dutch and German which confirms this tendency. Moreover, we observe considerable differences between written and spoken language with respect to the incidence of four main types of clausal coordinate ellipsis—Gapping, Forward Conjunction Reduction (FCR), Backward Conjunction Reduction (BCR), and Subject Gap with Finite/Fronted Verb (SGF). We argue that the detailed data pattern cannot be accounted for in terms of audience design, and propose an explanation based on the assumption that during spontaneous speaking—but not during writing—, the scope of online grammatical planning is basically restricted to one (finite) clause.
Identifying factors that improve the assessment of athletes' psychological functioning is imperative to make proper return-to-play decisions following concussion. Prior research indicates that an individual's affect is related to symptom reporting. The present study examines two novel methods of affect assessment in college athletes at baseline participating in a sports-concussion management program. A total of 256 athletes completed a neuropsychological baseline battery with measurements of psychological symptoms (BDI-Fast Screen, Post-Concussion Symptom Scale, and ImPact Total Symptom Score) and a measure of affective memory bias (the Affective Verbal Learning Test; AVLT). Examiners completed an observation-based rating of affect. Multivariate analysis of variance and χ2 analyses were conducted to examine the effect of affect on symptom reports. Examiners' Affect Ratings were predictive of broad symptom reporting, while the performance based index of affect (Affective Verbal Learning Test, AVLT) was more predictive of depressive symptoms. These findings suggest that performance on the AVLT may be a useful indicator of self-reported depression in a collegiate athlete sample. Additionally, these results demonstrate that examiners' behavioral assessments of affect are important in the assessment of psychological functioning in athletes. Continued work should focus on developing objective measures that are sensitive and valid for the evaluation of outcomes from concussion.
Data integration systems attempt to provide users with seamless and flexible access to information from multiple autonomous, distributed and heterogeneous data sources through a unified query interface. Besides data are continuously growing, maintained by different organizations and managed autonomously, querying data from heterogeneous data sources faces new challenges. As data integration has been automated, the ambiguity in concept interpretation also known as semantic heterogeneity has become one of the main obstacles to this process. Introduction of the Semantic Web Vision Ontologies WordNet ontology [3] is a large lexical database that is used in many schema matching algorithms to match schemas based on the semantics of attributes. In this paper ontology based semantic query reformulation technique is followed to improve the recall of the query. The reformulated query is optimized by removing disjunctive clauses in the query to reduce the computational cost of the semantic query execution. Experimental results show that the proposed optimization technique improves recall with minimal execution time.
FinnWordNet is a wordnet for Finnish that complies with the format of the Princeton WordNet (PWN) (Fellbaum, 1998).It was built by translating the Princeton WordNet 3.0 synsets into Finnish by human translators.It is open source and contains 117000 synsets.The Finnish translations were inserted into the PWN structure resulting in a bilingual lexical database.In natural language processing (NLP), wordnets have been used for infusing computers with semantic knowledge assuming that humans already have a sufficient amount of this knowledge.In this paper we present a case study of using wordnets as an electronic dictionary.We tested whether native Finnish speakers benefit from using a wordnet while completing English sentence completion tasks.We found that using either an English wordnet or a bilingual English-Finnish wordnet significantly improves performance in the task.This should be taken into account when setting standards and comparing human and computer performance on these tasks.
Among the most salient and extensively researched phonological processes of Caribbean Spanish after the categorical weakening of /-s / is the prolific behavior of the implosive or post-nuclear liquids // and /l/. It has been repeatedly claimed yet remarkably unsubstantiated that the gemination of word-medial, post-nuclear liquids to a following consonantal segment is a pervasive characteristic of Cuban Spanish, particularly of the western dialect region. To this end, the fundamental objective of the present study was to acoustically investigate said-phenomenon as it is purported to occur in the province of Havana, whose capital city models the linguistic norm for the rest of the country. Speech samples were elicited from twenty-four native speakers and spectrographic analysis was performed on these collected tokens in order to precisely identify the characteristics of both liquids in the above-mentioned segmental environment. Although we did not find any evidence of liquid gemination in the 120 words under analysis, we did observe two systematic and conditioned processes that may potentially help account for the impressionistic identification of gemination in word-internal position: 1) the insertion of an excrescent vowel between all [.C] sequences; and 2) an increase in the duration of the closure of the stop positionally subsequent to /-L / → [Ø] by an average 30.9%. In light of the evident lack of empirical studies dedicated to the allophony of final liquids, we believe the implications of the findings in this investigation to be important for Caribbean Spanish in general and Cuban Spanish in particular and hope that this study may provide a foundation for future work on the phenomenon of liquid gemination.
This paper gives a description of an annotation scheme for annotating a corpus of computer-mediated communication in Hindi (CO3H) with certain semantic, pragmatic and situational features. The annotation scheme is based on the theory of register analysis, where it is assumed that a registeral difference entails difference in certain linguistic features. It adapts and integrates the annotation schemes of sense annotation in the Penn Discourse Treebank and dialogue act annotation of DIT++ within this larger registeral framework. The situational and linguistic features that will be used to annotate the corpus for PoRT is described in the paper, along with some proposed labels for these features.
Corpus studies by Schuler, AbdelRahman, Miller, and Schwartz (2010), appear to support a model of comprehension taking place in a general-purpose working memory store, by providing an existence proof that a simple probabilistic sequence model over stores of up to four syntacticallycontiguous memory elements has the capacity to reconstruct phrase structure trees for over 99.9% of the sentences in the Penn Treebank Wall Street Journal corpus (Marcus, Santorini, & Marcinkiewicz, 1993), in line with capacity estimates for general-purpose working memory, e.g. by Cowan (2001).But capacity predictions of this simple structure-based model ignore non-structural dependencies, such as long-distance fillergap dependencies, that may place additional demands on working memory.Distinguishing unattached gap fillers from open attachment sites in syntactically-contiguous memory elements requires this contiguity constraint to be strengthened to a constraint that working memory elements be semantically contiguous.This paper presents corpus results showing that this stricter semantic contiguity constraint still predicts working memory requirements in line with capacity estimates such as that of Cowan (2001).
The paper demonstrates how the generic parser of a minimally supervised information extraction framework can be adapted to a given task and domain for relation extraction (RE). For the experiments a generic deep-linguistic parser was employed that works with a largely hand-crafted head-driven phrase structure grammar (HPSG) for English. The output of this parser is a list of n best parses selected and ranked by a MaxEnt parse-ranking component, which had been trained on a more or less generic HPSG treebank. It will be shown how the estimated confidence of RE rules learned from the n best parses can be exploited for parse reranking. The acquired reranking model improves the performance of RE in both training and test phases with the new first parses. The obtained significant boost of recall does not come from an overall gain in parsing performance but from an application-driven selection of parses that are best suited for the RE task. Since the readings best suited for successful rule extraction and instance extraction are often not the readings favored by a regular parser evaluation, generic parsing accuracy actually decreases. The novel method for task-specific parse reranking does not require any annotated data beyond the semantic seed, which is needed anyway for the RE task.
The purpose of this research was to compare the lexemes used in books for children aged 7 to 12 and the basic vocabulary for the same age groups. To this end, a comparable monolingual corpus of Persian translational and non-translational children’s literature was used from Persian Linguistic Database (PLDB). This corpus was compared with the list of Iranian Primary School Students Core Vocabulary. Results indicated that there was no acceptable conformity between the lexemes used in translational and non-translational texts prepared for children aged 7 to 12 and the basic vocabulary for the same ages. From the findings of this study, it can be assumed that in writing and translating books for children, the list of Iranian Primary School Students’ Core Vocabulary may be a reliable source; though other supplementary elements are also required.
Research suggests that ratings of child psychopathology by parents and teachers are generally not highly correlated. We examined the agreement and discordance between the child behaviour ratings of parents and teachers of a cohort of 6-year-old Pacific children living in New Zealand, based on scores from the Child Behaviour Checklist and the Teacher Report Form. Mother's reports were obtained for 1019 children, of whom, 602 also had father's reports and 559 had teacher's reports. Rater agreement was low between all pairs of informants. Fathers and teachers had higher agreement than mothers and fathers, the latter in turn had higher agreement than mothers and teachers, and agreement was generally higher for Externalizing problems than Internalizing problems. In terms of discordance, mothers reported more aggressive behaviour than fathers, while fathers reported more Internalizing and Total problems than mothers. Mothers and fathers generally reported more behaviour problems than teachers. The higher agreement found between informants from different settings (fathers and teachers) than between informants from similar settings (mothers and fathers) is in contrast with some of the literature. Further research is needed to investigate how child, informant, and setting characteristics affect ratings of children's behaviour.
The conditional collaborative filtering technology does not consider the user's context information which a-ffect user's rating.But recent research show that user's personalized context directly affect rating,so the result of recommendation can be improved if personalize context is incorporated into conditional collaborative filtering technology.Besides,the personalized context and item class are can be combined,Firstly classifying the items,and then making sure user's personalize context under every item class.When predicting the rating of target item,Firstly make sure which item class the target item is belong to,and then identify the user's personalized context used to compute the rating of the target item.The experimental results show that the recommendation accuracy of proposed approach is better than Slope One.
The phenomenon of morphosyntactic variation in nominal groups was undertaken in the linguistic literature in respect of the theory of variant, the causes of language evolution or linguistic norm in Standard Polish. Regional Polish language in the East Borderland (Kresy) provides an interesting text corpus showing the variation in using the dative case and the dla (Eng. for) prepositional phrase for a much broader range than in today’s Standard Polish, e.g. powiedzieć dla (Prp) księdza (Gen) – SP powiedzieć księdzu (Dat) ‘to tell priest’; spodobali się dla (Prp) was (Gen) – SP spodobali się wam (Dat) ‘you liked them’; dla (Prp) mnie (Gen) pomagała – SP mnie (Dat) pomagała ‘she helped me’. The subject of the article is to analyse the determinants of alternation (variation) between the dative case and dla (Eng. for) prepositional phrase on the basis of homogeneous language material, which is 120 pages of text, recorded in one of Polish dialects in Khmelnytsky district in Ukraine. The aim was to investigate the possible functional and semantic limitations of the variants in the dialect. Statement of the frequency of the competing morphosyntactic markings in the text corpus has shown that the prepositional phrase does not eliminate the dative case in any of analysed functions. The analysis demonstrates that alternation includes all the functions of the dative case, both the dativus commodi and the dativus incommodi. In particular, this second context illustrates the specific usage of dla prepositional phrase, e.g. zniszczyć meble dla (Prp) kogoś (Gen), which are not acceptable as dativus incommodi in Standard Polish (SP zniszczyć meble komuś (Dat) ‘destroy one’s furniture’).
Public opinion polls have historically indicated that the US public favours domestic over global priorities. It is not known what influence health knowledge has in shaping public opinion about domestic and global health policy. This study examines how knowledge of HIV/AIDS is related to the rated importance of domestic and global health issues. Participants were recruited to participate in an electronic survey (N = 995) and were predominantly White (86.3%), married (61.9%) and female (71.8%). HIV/AIDS knowledge was significantly associated with both domestic (β = 0.12, p < 0.05) and global health (β = 0.14, p < 0.01) priorities after controlling for sociodemographic variables. In addition, global health was found to act as a mediator between HIV/AIDS knowledge and perceived importance of domestic issues. Study findings suggest that those with greater HIV/AIDS knowledge rate global health issues higher, which in turn affects ratings of more domestic issues. This research has implications for ways to gain support for implementation of public health policy through increasing health knowledge.
The problem of automatically extracting structured information from texts is an important, unsolved problem within the field of Natural Language Processing. The extraction of such information can facilitate activities such as the building of knowledge bases, automatic \nsummarisation and sentiment analysis. A human reader can easily discern the events described in a text, along with the participants and the relationships between them, \nbut using a computer to automatically discover the same information is much more challenging. Particular focus has been given to extracting relations between the entities in a text, such as those representing geographical locations, personal and social relationships, and employment. In this thesis, we consider two closely related entity relationships, which are interesting, frequent and have not been tackled previously, which we refer to collectively \nas entity instantiations. \nWe define an entity instantiation as an entity relation in which a set of entities is introduced, and either a member or subset of this set is mentioned. In the example below, \nwe see a set membership instantiation, between ‘several EU countries’ and ‘the UK’, along with a subset instantiation, between the same set and ‘the low countries’. Inflation has increased sharply in several EU countries. In the UK, this has accompanied a drop in interest rates, but in the low countries rates have remained steady. This thesis details the creation of the first corpus of entity instantiations. The final corpus consists of 4,521 instantiations, 2,118 of which are intersentential, and 2,403 of which are intrasentential, annotated over 75 Penn Treebank Wall Street Journal newswire texts. The subsequent annotation study shows high levels of inter-annotator agreement and our \ncorpus study analyses the annotated entity instantiations in terms of their internal structure, the distance between arguments and their syntactic relationship, finding a particularly strong link between syntactic parent-child relationships and sentence-internal entity instantiations. \nTo establish that the accurate automatic identification of entity instantiations is possible, we develop the first instantiation identification algorithm, which uses a supervised machine learning approach. The feature set draws on surface, syntactic, contextual, salience and knowledge features to aid classification. We separately apply our classifier to intersentential and intrasentential entity instantiations and experiment with both balanced data, with a 50/50 positive/negative split, and the original unbalanced corpus. The classifier records highly significant performance increases over both unigram-based \nand majority class baselines on the balanced data, and also on the original distribution of intrasentential instantiations. \nIn order to take advantage of the aforementioned link between syntax and intrasentential entity instantiations, tree kernels were employed to learn directly from the syntactic parse trees which contain the two potential participants in an intrasentential instantiation. \nThe tree kernel features perform similarly to the unstructured feature set, with a much shorter development time. Combining tree kernels with unstructured features gives further improvements over both the baselines, and either method in isolation. We also apply our entity instantiations to the difficult problem of implicit discourse relation classification, hypothesising that introducing features identifying the presence of an entity instantiation between the arguments of a discourse relation can improve classification performance. Our experiments show that an entity instantiation is a strong indicator of the presence of an Expansion.Instantiation discourse relation. We create a binary Expansion.Instantiation classifier, based on the feature set detailed in Sporleder \nand Lascarides (2008), but augment it by adding entity instantiation features based on gold standard annotations. The classifier which includes entity instantiation data performs significantly better than the same classifier without entity instantiation data. We also experiment with the incorporation of machine-identified entity instantiations. However, our entity instantiation classifier is not sufficiently accurate to impact on discourse relation classification.
The first part of this paper deals with the concept of multi-word lexical units and the delimitation of their basic types in the Pralex lexical database, including the analysis of some terminological and conceptual issues. After a short preview of the approaches to the lexicographical treatment of multi-word lexical units in contemporary monolingual dictionaries, the second part of this paper deals with the database processing of multi-word lexical units, both with the general rules of their treatment in the Pralex lexical database and with the specific rules of treatment of one of two basic types of multi-word lexical units, i.e. multiple-word namings.
This paper analyses a Quebec comic strip, Magasin general, in which language is both an element of group cohesion and of explicit thematisation. Starting from the problematic relationship Quebec has with French linguistic norm, we try to investigate the use of language in the BD itself, the discourse about language by the characters and by a large paratextual apparatus, and also the reception of this comic by the Quebec highbrow press.
The article deals with the program and content aspects of the treatment of the Czech lexis in the form of a lexical database. It summarises the development of the Praled software, the basic features of the macro- and microstructures of the Pralex LDB and the main principles of the treatment of database items.
In this paper, we give a summary of various dependency chart parsing algorithms in terms of the use of parsing histories for a new dependency arc decision. Some parsing histories are closely related to the target dependency arc, and it is necessary for the parsing algorithm to take them into consideration. Each dependency treebank may have some unique characteristics, and it requires for the parser to model them by certain parsing histories. We show in experiments that proper selection of the parsing algorithm which reflect the dependency annotation of the coordinate structures improves the overall performance. 1
The paper deals with compounds in the Pralex lexical database, especially with their general characteristics and problems with their processing in the database. Apart from the compounds, the author focuses on the particular components and their origin, meanings and mutual relations. Compounds with a quantitative meaning as well as compounds containing such component are not handled in here.
The article describes the development of the typographical practice of placing spaces between words in the early printed Glagolitic books and the swift decline in the 16 th century of the use of so-called word-blocks in favour of full word separation.Little studied but signifi cant for our understanding of a host of writing and reading practices, ranging from the rhetorical and compositional features of the medieval Croatian Church Slavonic texts to the linguistic norms of early modern works, the use of the white space consistently increases over time as modern typographical practices take hold.In the context of widespread changes in mechanical printing practices in the 15 th -17 th centuries, the paper looks at the increasing use of white space between all words, primarily in the CrCS printed liturgical books and on the parallel decrease in the use of word-blocks (zdruenice).Given the larger movement toward regularization of liturgical texts in the 16 th century and the growing awareness of linguistic science our examination of typesetting practices offers some insights into the implementation of regularized linguistic norms for the Glagolitic liturgical books and makes it possible to conclude that the more widespread typesetting practices of the secular presses quickly gained a foothold in the ecclesiastical printeries.There was, moreover, a rapid conformity to the Western typographical practice of separating words as the smallest units of independent meaning; i.e. in accordance with our own contemporary practices.
The article focuses on the contacts between prescriptive and universal grammar in 18th century British linguistics. The author argues that, contrary to the wide-spread belief, there existed two-way relations between them: on the one hand, practical prescriptive grammar used the principles of universal grammar as the foundation for singling out parts of speech and grammar categories and based normative recommendations upon similarities between languages. On the other hand, the authors of universal grammars showed interest to problems of linguistic norm and gave practical recommendations in the vein of prescriptive tradition.
The reconstruction of standardized texts in the Prague Dependency Treebank of Spoken Czech enables the comparison of authentic spoken utterances and standardized texts The authors concentrate on the questions: What does the syntactic identity of the Czech spoken and written texts consist of? What syntactic constructions are „natural“ in the spoken and in the written text? What is the difference in the density of the cohesive links, in the explicit and implicite relations between units?
Introduction Brain-Computer Interfaces (BCI) can be used for communication and motor restoration (Birbaumer & Cohen, 2007). To our knowledge, no BCI study looked at patients with dementia who have severe communication deficits. BCIs based on operant training could be problematic for patients with cognitive deficits. A paradigm shift from instrumental-operant learning to classical conditioning could possibly overcome this failure (Birbaumer, 2006). Recent findings demonstrated the possibility to classify cognitive and emotional states by the pattern classification of BOLD signals in both offline (Lee et al., 2010a, 2010b) and online situations (Sitaram et al., 2010). The present study aims to investigate the feasibility of an auditory classical conditioning paradigm within a fMRI based BCI setting. The paradigm is designed to condition individuals to associate positive and negative emotional stimuli as unconditioned stimuli (US) with congruent and incongruent word pairs as conditioned stimuli (CS), respectively. Our goal is to ascertain whether the brain signals pertaining to congruent and incongruent word pairs could be classified with more than chance accuracy using our fMRI support vector machine (SVM) with a view to apply for basic online yes/no communication in Alzheimer patients. Methods The paradigm consisted of one single session divided into six blocks, comprising the different phases of conditioning (habituation, acquisition, extinction). The US consisted of auditory emotional stimuli selected from the International Affective Digitized Sounds (IADS, Bradley & Lang, 1999). A segment of baby laughter represented the positive emotional stimulus and a segment of screaming represented the negative emotional stimulus. The CS, presented aurally, were congruent (e.g. ‘animal-elephant’) and incongruent (e.g. ‘animal-Germany’) word-pairs. The unconditioned and conditioned responses (UR and CR) were the changes in the BOLD signal pertaining to the CS and US, respectively. The first block consisted of a randomized presentation of 50 US and 50 CS. In the second and third blocks 25 congruent word pairs, immediately followed by the baby laughter, and 25 incongruent word pairs, immediately followed by the scream, were presented randomly. In the fourth and fifth blocks, respectively 40% and 20% of the CS were paired with the US. In the sixth block, only the CS was presented. Functional imaging was performed continuously during these blocks on 6 healthy subjects (4 females, 2 males, age 21-27) on a 3.0 T scanner (Siemens, Germany). To classify the signals corresponding to various conditions, namely congruent and incongruent word-pairs, a linear SVM (with the regularization parameter, C=1) was implemented. Classification performance from data was evaluated through 2-fold cross validation (CV). Based on the parameters of the trained SVM model, we analyzed the fMRI data with the Effect Mapping method (EM; Lee et al., 2010a, 2010b, Sitaram et al., 2010). To investigate the relative importance of different brain regions in decoding the conditioning brain states, feature vectors from the frontal cortex were used as input to build a separate SVM classifier. Results The Self Assessment Manikin (SAM) test showed that participants reported more negative valence and a higher arousal for the scream compared to the baby laughter. Classification of the BOLD signal as a response to the congruent and incongruent word pairs immediately followed by the emotional US showed above chance level performance (57-64%) on one subject, around chance level (50-56%) performance on three subjects, and below chance level (44-47%) performance on two subjects. Conclusions In this pilot study we have demonstrated an approach for conditioning the BOLD signal by repeated association of the emotional stimuli with semantic stimuli, resulting in a paradigm for basic yes/no communication. Further work includes improving the performance of the classifier by feature selection, an online implementation of the system, and its testing on patients. References: Birbaumer, N. (2006), ‘Brain-computer-interface research: coming of age’, Clinical Neurophysiology, vol. 117, pp. 479-483. Birbaumer, N. & Cohen, L. G. (2007), ‘Brain-computer interfaces: communication and restoration of movement in paralysis’, The Journal of Physiology, vol. 579, no. 3, pp. 621-636. Bradley, M. M. & Lang, P. J. (1999), ‘International Affective Digitized Sounds (IADS): Stimuli, instruction manual and affective ratings’, University of Florida, Gainesville. Sitaram R, Lee S, Ruiz S, Rana M, Veit R, Birbaumer N. Real-time support vector classification and feedback of multiple emotional brain states. Neuroimage, 2010 Aug 6. Lee, S., Halder, S., Kübler, A., Birbaumer, N., Sitaram, R. Effective functional mapping of fMRI data with support-vector machines. Hum Brain Mapp. 2010a, Jan 28. Lee, S., Ruiz, S., Caria, A., Birbaumer, N., Sitaram, R Cerebral reorganization induced by real-time fMRI feedback training of the insular cortex: a multivariate investigation. Neuroreh and Neural Rep (2010b).
The paper deals with the macrostructure of the Pralex lexical database and with the basic principles of building an index, mainly with the choice of sources and with criteria of the choice of lexical units including possibilities of their future completion from own sources of the Institute.
Nous montrons comment enrichir une annotation en dependances syntaxiques au format du French Treebank de Paris 7 en utilisant la reecriture de graphes, en vue du calcul de sa representation semantique. Le systeme de reecriture est compose de regles grammaticales et lexicales structurees en modules. Les regles lexicales utilisent une information de controle extraite du lexique des verbes francais Dicovalence.
Based on the assertions of the Theory of Linguistic Variation and Change, this paper proposes a discussion about the possible action of linguistic norms over two variable phenomena in Brazilian Portuguese: the position of clitic pronouns associated with a single verb, and the use of prepositions with verbal complements indicating a?goal/recipient?. By the analysis of data from the newspapers from São Paulo and Rio Claro between (the years of) 1900 and 1915, we intend to describe each phenomenon; compare these descriptions and evaluate the role played by the standard and the common usage which is (already) perceptible in the?paulistas? continuous published pages of that period.
In this paper we aim to automatically identify subjects, which are not expressed but nevertheless understood in Czech sentences. Our system uses the maximum entropy method to identify different types of unstated subjects and the system has been trained and tested on the Prague Dependency Treebank 2.0. The results of our experiments bring out further consideration over the suitability of the chosen corpus for our task.
This paper presents our preliminary work on adaptation of parsing technology toward natural language query processing for biomedical domain. We built a small treebank of natural language queries, and tested a state-of-theart parser, the results of which revealed that a parser trained on Wall-Street-Journal articles and Medline abstracts did not work well on query sentences. We then experimented an adaptive learning technique, to seek the chance to improve the parsing performance on query sentences. Despite the small scale of the experiments, the results are encouraging, enlightening the direction for effective improvement. 1
In this paper, we present a series of enhancements to the English-Japanese version of a multilingual Linguistics Based Machine Translation system, for the translation of complex sentences, modality and complex verbal structures. The system is using a classical transfer-based architecture and dedicated lexical databases. Relying on linguistic data acquired on large corpora or compiled from the web, corrections have been done in constituent reordering, lexical selection and verb conjugation. Even if the system is not, so far, as efficient as state-of-the-art English-Japanese MT systems, the results show a clear progress and underline the interest of using syntactic information in MT.
In the present paper we focus on control as a subtype of anaphora. We work with the theory of control present within the dependency-based framework of Functional Generative Description (FGD), in which control is defined as a relation of a referential dependency between a controller (antecedent � semantic argument of the main clause) and a controllee (anaphor � empty subject of the nonfinite complement (controlled clause)). First this paper presents the rule-based reconstruction of controllees, then, it discusses the perceptron-based determination of the controllees� antecedent. We evaluated our approach on data from the Prague Dependency Treebank 2.0, however, the rules and features of our system are supposed to be language independent and can be tested on other languages in the future.
The core of the processing of individual items in the lexical database Pralex is an example part of an entry (exemplification). While processing the nouns, we try to select a relevant evidence from the corpus SYN, i.e. examples of a concrete use of meanings processed in advance according to SSJČ and, in addition to that, we also try to record evidence of the new meanings. In the exemplification, we focus on a syntactic and semantic linkage. Lemmas are mainly illustrated by means of non-quotational, modified evidences, i.e. example linkage in the form of syntagmas, which are regarded as common collocations of the processed nouns. Sentential records are introduced as well, though in a lower extent.
For bilingual students, attending a university away from home serves as a separation from one's native linguistic norms to an entirely new set of language practices. These U.S. institutions of higher learning almost always demand English in favor of other native or heritage languages. This study examines the ways a small group of second- and third- generation Lithuanian immigrants attempt to maintain their heritage language. It also looks at how language loss can manifest within a college environment. The research suggested the difficulties and challenges college life poses on the retention of a heritage language. Discussions for further research include: how University policy might implicitly prescribe English, use of heritage language among friend as a marker of kinship, the difficulty of resisting English monolingualism within the U.S. public school system, and the loss of ethnic customs and identity in sync with language loss.
Various linguistic notions subsumed by or closely related to point of view have been discussed under different terms in different frameworks (e.g. deictic center, empathy, subjectivity). These show that grammatical forms in the sentence depend in some way on the speaker’s point of view. This paper aims to propose that Korean discourse markers ‘-nun’ and ‘-ga’ can represent the speaker’s point of view toward the sentence as a system, i.e. a deictic center, subjectivity, human animacy or activity marker ‘-nun’ and a speaker-distal, objectivity, thing animacy or passivity marker ‘-ga’. This also implies that various properties related to the speaker’s point of view can be subsumed under just two categories. In this paper, quantitative patterns of these markers based on verbs sensitive to point of view, as found in the 21 Century Sejong Treebank, are estimated by statistical method Correspondence Analysis.