Papers reviewed and determined not to be word norm studies. Use the flag icon to report errors or suggest re-inclusion.
16504 papers
Automatic text analysis systems can lexically recognize a word only if it already exists in the electronic dictionary. The same thing is true for the NOOJ system analysis programs. One understands here by electronic dictionaries the lexical databases where all information is explicit because they are intended for computer programs use. These bases aim at the modelling of the language, which distinguishes them from the electronic lexicons created for particular applications needs. For technical languages or of speciality, work remains to be made to build dictionaries. Our work applies to the NOOJ system, with for immediate objective the installation of an electronic dictionary of French computer science terms (compound words), with an aim of analysing, automatically indexing texts. The linguistic aspects of the terminology are retained.
Recently the rational principle and the customary principle have been regarded as the main judging basis for linguistic norm.But the practice proves that the rational principle is hard to conduct and the customary principle less popularized.The establishment of the judging basis for linguistic norm should take the consideration of those target's characteristics.Langue and parole must be divided and ruled by their dissimilarity.Stability,sociality and pragmatic function should be the basis for langue,while effect principle and ethic principle should be the basis for parole.When dealing with the transitional area between langue and parole,the normal and the abnormal,we should not go to extremes.
AIM: (i) To determine the ability of general practitioners (GPs) and paediatricians to correctly identify children as overweight or obese by visual cues alone; (ii) to describe the current management practices of overweight and obese children by these practitioners; and (iii) to compare these with National Health and Medical Research Council (NHMRC) Clinical Practice Guidelines. METHODS: Forty-four GPs and 29 paediatricians participated in the study. Respondents completed a questionnaire based on a series of body images, rating the size of the child as acceptable weight, overweight or obese and indicating the likelihood of carrying out a series of management options. RESULTS: There was considerable variation in ability to rate images correctly with the total number of correct responses being 72% and 68%, respectively, for GPs and paediatricians. There were statistically significant differences in management between GPs and paediatricians in terms of conducting appropriate anthropometry and screening for co-morbidities, with paediatricians performing closer to the NHMRC Clinical Practice Guidelines. CONCLUSION: GPs and paediatricians have the opportunity to screen children for overweight and obesity during their everyday practice. Accurate determination of weight status cannot be performed by visualisation alone and all children should have height and weight measured and correctly interpreted. Some areas of current GP and paediatrician management of overweight and obese children fall short of the NHMRC clinical guidelines and areas for improvement are highlighted in this paper.
When compared with its two large neighbours, Russian and Polish, the Ukrainian language presents a picture of striking internal variation. Not only are Ukrainian dialects more mutually divergent than those of Polish or of territorially more widespread Russian,2 but on the literary level the language has long been characterized by the existence of two variants of the standard which have never been perfectly harmonized, in spite of the efforts of nationalist writers for a century and a half. While Ukraine’s modern standard language is based on the eastern dialect of the Kyiv-Poltava-Kharkiv triangle, the literary Ukrainian cultivated by most of the diaspora communities continues to follow to a greater or lesser degree the norms of the Lviv koiné in 1 The authors would like to thank Dr Lance Eccles of Macquarie University for technical assistance in producing this paper. 2 De Bray (1969: 30-35) identifies three main groups of Russian dialects, but the differences are the result of internal evolutionary divergence rather than of external influences. The popular perception is that Russian has minimal dialectal variation
In this paper, we present a system which can extract syntactic feature structures from a Korean Treebank (Sejong Treebank) to develop a Feature-based Lexicalized
We examined the effects of mood and the content (a priori valence and involvement) and formal (presentation modality: text vs. video) characteristics of messages presented on a small screen on emotional responses and involvement among 47 young adults. Mood was induced by autobiographical memories varying in affective valence and arousal. Facial electromyography (EMG) and cardiac interbeat intervals were used as physiological indexes of valence and arousal. Both mood and the emotional tone of a message exerted an independent influence on the emotional response to the message. A strong valence-related mood-congruency effect emerged in predicting involvement. The text modality elicited higher involvement, arousal ratings, and orbicularis oculi EMG activity compared to the video modality when in a depressed mood, whereas the reverse was true when in a joyful, relaxed, or fearful mood. The results point to the possibility of mood-adapted media services.
The correct attachment of prepositional phrases (PPs) is a central disambiguation problem in parsing natural languages. This paper compares the baseline situation in English, German and Swedish based on manual PP attachments in various treebanks for these languages. We argue that cross-language comparisons of the disambiguation results in previous research is impossible because of the different selection procedures when building the training and test sets. We perform uniform treebank queries and show that English has the highest noun attachment rate followed by Swedish and German. We also show that the high rate in English is dominated by the preposition of. From our study we derive a list of criteria for profiling data sets for PP attachment experiments.
We exploit the resources in the Arabic Treebank (ATB) for the novel task of automatically creating lexical semantic verb classes for Modern Standard Arabic (MSA). Verbs are clustered into groups that share semantic elements of meaning as they exhibit similar syntactic behavior. The results of the clustering experiments are compared with a gold standard set of classes, which is approximated by using the noisy English translations provided in the ATB to create Levin-like classes for MSA. The quality of the clusters is found to be sensitive to the inclusion of information about lexical heads of the constituents in the syntactic frames, as well as parameters of the clustering algorithm. The best set of parameters yields an Fβ=1 score of 0.501, compared to a random baseline with an Fβ=1 score of 0.37.
A focus on student diversity presumes that educational decisions, from statewide policies to individual classroom practices, may affect different student populations differently. Therefore, while the various aspects of student diversity are reflected in differing science outcomes, the ways in which policies and schools define, delimit, and manage student diversity may affect outcomes at least as much as does “diversity” itself. Regardless of the origin or nature of students' marginalization, academic success depends to a significant degree on assimilation to mainstream cultural and linguistic norms, for example, particular ways of structuring narratives, displaying competence, or interacting with adults, not to mention the phonological and grammatical conventions of standard English (Delpit, 1995; Heath, 1983). Traditional science instruction generally assumes that students have access to certain educational resources at home (such as computers, or adults with the time and academic skills to help with homework), and it requires students living in poverty to adopt learning habits that necessitate a certain level of socioeconomic stability (such as a quiet place to study, and freedom from child care or work-related responsibilities). While some students may overcome these barriers to academic success through exceptional talent, effort, or family support, the existence of such individuals does not negate the inequity of their educational circumstances or the need for social solutions to what are social, not individual, problems. Such issues must be taken into account in interpreting gaps in science outcomes among diverse student groups and in devising instructional programs to close the gaps.
We present an automatic approach to tree annotation in which basic nonterminal symbols are alternately split and merged to maximize the likelihood of a training treebank. Starting with a simple X-bar grammar, we learn a new grammar whose nonterminals are subsymbols of the original nonterminals. In contrast with previous work, we are able to split various terminals to different degrees, as appropriate to the actual complexity in the data. Our grammars automatically learn the kinds of linguistic distinctions exhibited in previous work on manual tree annotation. On the other hand, our grammars are much more compact and substantially more accurate than previous work on automatic annotation. Despite its simplicity, our best grammar achieves an F1 of 90.2% on the Penn Treebank, higher than fully lexicalized systems.
In this paper, we describe an annotation scheme for the attribution of abstract objects (propositions, facts, and eventualities) associated with discourse relations and their arguments annotated in the Penn Discourse TreeBank. The scheme aims to capture both the source and degrees of factuality of the abstract objects through the annotation of text spans signalling the attribution, and of features recording the source, type, scopal polarity, and determinacy of attribution. RESUME. Dans cet article, nous decrivons un schema d’annotation pour l’encodage des objets abstraits (propositions, faits et possibilites) associes aux relations de discours et a leurs arguments tels qu’annotes dans le Penn Discourse TreeBank. Ce schema a pour objet la capture de la source et du degre de factualite des objets abstraits. Les aspects cles de ce schema comprennent l’annotation des intervalles textuels signalant l’attribution, ainsi que l’annotation des proprietes caracterisant la source, le type, la polarite de la portee, et le degre de determination de l’attribution.
This paper investigates the complexity of dependencies at the discourse level, in particular the dependencies between discourse connectives and their argu-ments. Our study is based on data from the Penn Discourse Treebank (PDTB) and is therefore an exploration into the ways treebanks can inform linguistic issues. We observe that, unlike in syntax, there is more uncertainty and flex-ibility with regards to the location and extent of discourse arguments. This leads to a variety of possible patterns of dependencies between pairs of dis-course relations, including nested, crossed and a range of other non-tree-like configurations. Nevertheless, our main conclusion is that the types of dis-course dependencies are highly restricted since the more complex cases can be factored out by appealing to discourse notions like anaphora and attribu-tion. We conjecture that the complexity of dependencies is far more restricted at the discourse level as compared to the syntactic level. 1
This paper describes an ongoing effort to parse the Hebrew Bible. The parser consults the bracketing information extracted from the cantillation marks of the Masoetic text. We first constructed a cantillation treebank which encodes the prosodic structures of the text. It was found that many of the prosodic boundaries in the cantillation trees correspond, directly or indirectly, to the phrase boundaries of the syntactic trees we are trying to build. All the useful boundary information was then extracted to help the parser make syntactic decisions, either serving as hard constraints in rule application or used probabilistically in tree ranking. This has greatly improved the accuracy and efficiency of the parser and reduced the amount of manual work in building a Hebrew treebank.
Based on the analysis of the usage and the syntactic function of Chinese punctuations,this paper proposes a new hierarchical approach to parse the long Chinese sentences.In traditional parsing approaches,the parsing procedure is performed in an one-level way and the punctuation marks are not specially treated.Correspondingly,in our approach,the complex long Chinese sentences are broken into sub-sentences or units(say 'units' hereafter) by using punctuation marks with special functions,so that the original whole sentence is parsed unit by unit.This idea of 'divide-and-conquer' greatly reduces the difficulty in the traditional parsing approaches to recognize the syntactic relationship between the sub-sentences and phrases or inside the sub-sentences or phrases.And also,in our approach,the grammatical rules with punctuation marks and their probabilities are extracted from the large scale treebank,which are very beneficial to the syntactic disambiguation.Our experimental results have shown that comparing with the traditional Chart parsing algorithm,our approach can significantly reduce the time consumption and the numbers of ambiguous edges,and get about 7% of the correct rate and the recall rate increasing while parsing long Chinese sentences.
Statistical parsers trained and tested on the Penn Wall Street Journal (WSJ) treebank have shown vast improvements over the last 10 years. Much of this improvement, however, is based upon an ever-increasing number of features to be trained on (typically) the WSJ treebank data. This has led to concern that such parsers may be too finely tuned to this corpus at the expense of portability to other genres. Such worries have merit. The standard "Charniak parser" checks in at a labeled precision-recall f-measure of 89.7% on the Penn WSJ test set, but only 82.9% on the test set from the Brown treebank corpus.This paper should allay these fears. In particular, we show that the reranking parser described in Charniak and Johnson (2005) improves performance of the parser on Brown to 85.2%. Furthermore, use of the self-training techniques described in (McClosky et al., 2006) raise this to 87.8% (an error reduction of 28%) again without any use of labeled Brown data. This is remarkable since training the parser and reranker on labeled Brown data achieves only 88.4%.
The thesis presents tools for analysis at analytical and tectogrammatical layers that the Prague Dependency Treebank is based on. The tools for analytical annotation consist of two parsers and a tool for assigning syntactic tags. Although the performance of the parsers is far below that of the state-of-the-art parsers, they both can be considered a certain contribution to parsing, since the methods they are based on are novel. The tool for assigning syntactic tags makes 15% less errors than a tool used for this purpose previously. The tool developed for tectogrammatical annotation is the only one that can currently perform this task in such a breadth. Although other, specialized tools may have a better performance of some of its particular subtasks, my tool makes 29% and 47% less errors for the Czech language than the combination of existing tools for annotating the tectogrammatical structure and deep functors, respectively, which are the core of the tectogrammatical layer. The proposed tools are designed the way they can be used for other languages as well.
Abstract Facial masculinity may be used as a cue in female mate choice, as it reflects the success of the male genotype in its developmental environment. Women may maximize reproductive success by using a conditional strategy favoring highly masculine facial features for short‐term relationships and feminized facial features in men for long‐term relationships. Three studies examine reactions to masculinized and feminized male facial composites. Properties of the original composite image affect ratings of critical attributes and the magnitude of the differences in ratings between versions undergoing identical processes of geometric manipulation (Study 1). Both men and women attribute personality, behavior, and mating strategies consistent with predictions derived from the good genes and mating trade‐off hypotheses (Study 2). Participants accurately grouped behavioral tendencies related to high mating effort/risky strategies and high parenting effort/risk adverse strategies and associated mating effort more so with masculinized faces and parenting effort more so with feminized faces (Study 3). These results indicate that male facial masculinity serves as a visual cue for inferring personality and reproductive strategy.
This paper presents a comparative study of probabilistic treebank parsing of German, using the Negra and TüBa-D/Z tree-banks. Experiments with the Stanford parser, which uses a factored PCFG and dependency model, show that, contrary to previous claims for other parsers, lexicalization of PCFG models boosts parsing performance for both treebanks. The experiments also show that there is a big difference in parsing performance, when trained on the Negra and on the TüBa-D/Z treebanks. Parser performance for the models trained on TüBa-D/Z are comparable to parsing results for English with the Stanford parser, when trained on the Penn treebank. This comparison at least suggests that German is not harder to parse than its West-Germanic neighbor language English.
Abstract We present a new method for learning to parse a bilingual sentence using Inversion Transduction Grammar trained on a parallel corpus and a monolingual treebank. The method produces a parse tree for a bilingual sentence, showing the shared syntactic structures of individual sentence and the differences of word order within a syntactic structure. The method involves estimating lexical translation probability based on a word-aligning strategy and inferring probabilities for CFG rules. At runtime, a bottom-up CYK-styled parser is employed to construct the most probable bilingual parse tree for any given sentence pair. We also describe an implementation of the proposed method. The experimental results indicate the proposed model produces word alignments better than those produced by Giza++, a state-of-the-art word alignment system, in terms of alignment error rate and F-measure. The bilingual parse trees produced for the parallel corpus can be exploited to extract bilingual phrases and train a decoder for statistical machine translation.
Previous articleNext article FreeCurrent ApplicationsLinguistic AnthropologyV.ChandV.Chand Search for more articles by this author PDFPDF PLUSFull Text Add to favoritesDownload CitationTrack CitationsPermissionsReprints Share onFacebookTwitterLinked InRedditEmailQR Code SectionsMoreImmigration Practices in Belgium: African AsylumSeeker DiscourseJan Blommaert, a Belgian anthropologist currently at the Institute of Education, University of London, is known in Belgium as a public campaigner on immigration issues. He became involved in African asylumseekers rights in 1998, when the death of Semira Adamu during her forced repatriation provoked a public outcry over the implementation of Belgian immigration policies, in which more than 95% of applicants for asylum are rejected. Adamu had fled Nigeria in her teens to avoid entering into a polygamous marriage with a 65yearold man. Shackled at the ankles and vigorously resisting, she was suffocated by Belgian police attempting to restrain her.Poster for a 2003 commemoration of Adamu's death.View Large ImageDownload PowerPointAdamus death was a catalyst for the formation of new action groups, and existing organizations also became involved: churches opened their doors to asylum seekers, and NGOs such as OXFAM and the League of Human Rights campaigned for asylum seekers rights. Early on, Blommaert contributed to this organic collaborative effort: I gave tons of public lectures for any audience willing to listen, wrote opeds in major newspapers, campaigned with MPs close to the government, participated in public debates on these matters, and wrote expert articles for a wider audience. Additionally, he saw that his understanding of the legal and sociolinguistic issues was directly applicable to the problem. He was motivated, he explains, by awareness of the real stakes and real people involved.African asylum seekers face circumstances not shared with those from Europe and the Gulf because they come from wartorn areas with unclear state boundaries, may be illiterate and lack documentation, have long migration paths to Belgium, and do not share a language with immigration authorities. In particular, Blommaert points out, their choices of language, background texts, and genres for storytelling affect their chances of acquiring refugee status. African asylumseeker language is stereotypically filled with language mixing, language impurities, varying degrees of literacy, and different styles of storytelling and discourse. When asylum seekers present their stories, they are unaware of how the officials are judging them on their linguistic habits. The officials note these details and find them inconsistent with Belgian expectations; they judge the asylum seekers in terms of these nave choices and reject them. In many such cases, rejection is a matter of life or death for the refugees, given the risks associated with deportation and with repatriation into the home country that originally motivated them to seek asylum.Blommaert has mobilized an informal network of Belgian academics and institutions that has produced many academically informed public statements. Additionally, by documenting and analyzing oral immigration interviews with an eye to understanding the disconnect between asylum seekers presentations and immigration officials expectations, he has improved practice in the Immigration Department. He has used his analysis to train members of the department, helping to adjust views of what can be gathered in an immigration interview, to improve interview techniques, and to promote an awareness of variability in sociocultural linguistic norms and presentation styles.He continues to work as an official expert for Belgian legal, government, and security authorities, translating and analyzing documentation and corroboratory written texts provided for immigration procedures and advising on specific issues. He has been able to influence outcomes for some individual asylum seekers. His public campaigning has raised awareness of the issues surrounding African asylum seekers, and his publications have provoked work within academia that may lead to further interventions by academics with regard to the interview process.Blommaert hopes that public campaigning by NGOs and academic scholarship will promote continued dialogue with immigration officials, eventually producing policies and procedures that take into account the politics of migration and displacement and the way in which asylum seekers frame their life stories. Previous articleNext article DetailsFiguresReferencesCited by Current Anthropology Volume 47, Number 3June 2006 Sponsored by the Wenner-Gren Foundation for Anthropological Research Article DOIhttps://doi.org/10.1086/504161 Views: 287Total views on this site Citations: 1Citations are reported from Crossref PDF download Crossref reports the following articles citing this article:Kevin D. O'Gorman, Cailein Gillespie The mythological power of hospitality leaders?, International Journal of Contemporary Hospitality Management 22, no.55 (Jul 2010): 659–680.https://doi.org/10.1108/09596111011053792
A 66-year-old man was admitted to our hospital on suspicion of lung cancer with bone metastasis. He suffered multiple joint and muscle pain. 18F-Fluorodeoxy glucose positron emission tomography (FDG-PET) showed multiple accumulations in the lung, bones including the vertebrae, and mediastinal lymph nodes. Anti-human immunodeficiency virus (HIV) antibody was negative. Because Mycobacterium avium complex (MAC) was isolated from bronchial lavage fluid, bronchial wall, peripheral blood, and muscle abscess, he was diagnosed as having disseminated MAC infection. Although multidrug chemotherapy was initiated, his condition rapidly deteriorated at first. After surgical curettage of the musculoskeletal abscess, his condition gradually improved. As for etiology, we suspected that neutralizing factors against interferon-gamma (IFN-γ) might be present in his serum because a whole blood IFN-γ release assay detected low IFN-γ level even with mitogen stimulation. By further investigation, autoantibodies to IFN-γ were detected, suggesting the cause of severe MAC infection. We should consider the presence of autoantibodies to IFN-γ when a patient with disseminated NTM infection does not indicate the presence of HIV infection or other immunosuppressive condition.
This paper investigates how to extend coverage of a domain independent lexicon tailored for natural language understanding. We introduce two algorithms for adding lexical entries from VerbNet to the lexicon of the Trips spoken dialogue system. We report results on the efficiency of the method, discussing in particular precision versus coverage issues and implications for mapping to other lexical databases.
In this paper we describe the structure and development of the Brandeis Semantic Ontology (BSO), a large generative lexicon ontology and lexical database. The BSO has been designed to allow for more widespread access to Generative Lexicon-based lexical resources and help researchers in a variety of computational tasks. The specification of the type system used in the BSO largely follows that proposed by the SIMPLE specification (Busa et al., 2001), which was adopted by the EU-sponsored SIMPLE project (Lenci et al., 2000). 1.
Abstract Attempting to automatically learn to identify verb complements from natural language corpora without the help of sophisticated linguistic resources like grammars, parsers or treebanks leads to a significant amount of noise in the data. In machine learning terms, where learning from examples is performed using class-labelled feature-value vectors, noise leads to an imbalanced set of vectors: assuming that the class label takes two values (in this work complement/non-complement), one class (complements) is heavily underrepresented in the data in comparison to the other. To overcome the drop in accuracy when predicting instances of the rare class due to this disproportion, we balance the learning data by applying one-sided sampling to the training corpus and thus by reducing the number of non-complement instances. This approach has been used in the past in several domains (image processing, medicine, etc) but not in natural language processing. For identifying the examples that are safe to remove, we use the value difference metric, which proves to be more suitable for nominal attributes like the ones this work deals with, unlike the Euclidean distance, which has been used traditionally in one-sided sampling. We experiment with different learning algorithms which have been widely used and their performance is well known to the machine learning community: Bayesian learners, instance-based learners and decision trees. Additionally we present and test a variation of Bayesian belief networks, the COr-BBN (Class-oriented Bayesian belief network). The performance improves up to 22% after balancing the dataset, reaching 73.7% f-measure for the complement class, having made use only a phrase chunker and basic morphological information for preprocessing.
This paper discusses a novel probabilistic synchronous TAG formalism, synchronous Tree Substitution Grammar with sister adjunction (TSG+SA). We use it to parse a language for which there is no training data, by leveraging off a second, related language for which there is abundant training data. The grammar for the resource-rich side is automatically extracted from a treebank; the grammar on the resource-poor side and the synchronization are created by handwritten rules. Our approach thus represents a combination of grammar-based and empirical natural language processing. We discuss the approach using the example of Levantine Arabic and Standard Arabic.
This presentation reports the methodology followed and the results attained on an on-going project aiming at building a large lexical database of corpus-extracted multiword (MW) expressions for the Portuguese language. MW expressions were automatically extracted from a balanced 50 million word corpus compiled for this project, furthermore statistically interpreted using lexical association measures and are undergoing a manual validation process. The lexical database covers different types of MW expressions, from named entities to lexical associations with different degrees of cohesion, ranging from totally frozen idioms to favoured co-occurring forms, like collocations. We aim to achieve two main objectives with this resource: to build on the large set of data of different types of MW expressions to revise existing typologies of collocations and to integrate them in a larger theory of MW units; to use the extensive hand-checked data as training data to evaluate existing statistical lexical association measures.
Objective:This study aimed to examine the effects of haloperidol and amphetamine on human startle response modulated by emotionally-toned film clips. Method: Sixty participants, in two groups (one receiving haloperidol and the other receiving amphetamine) were tested using electromyography (EMG) to measure eye-blink muscle (orbicular oculi) while different emotions were induced by six 2-minute film clips. Results: An affective rating shows the negative and positive effects of the two drugs on emotional reactivity, neither amphetamine nor haloperidol had any impact on the modulation of the startle response. Conclusion: The methodological and theoretical aspects of the study and findings will be discussed.
The databases record instances of deponency, which is the term we have adopted to describe mismatches between morphology and morphosyntax. The prototypical example are the deponent verbs of Latin, which involve a mismatch between passive form and active meaning. That is, a normal Latin verb had active forms such as amō 'I love' and amāvī 'I have loved', which contrasted with the passive forms amor 'I am loved' and amātus sum 'I have been loved' (in this case, with a masculine subject). A deponent verb, on the other hand, looks like the passive but functions like the active, as in mīror 'I admire', mīrātus sum 'I have admired'. In the the databases we construe deponency in an extended fashion, covering any mismatch between the apparent morphosyntactic value of a morphological form and its actual value in a given syntactic context. Two databases are housed on this site, accessible through the links above. The cross-linguistic database looks at the presence of morphological mismatches in a controlled sample of genetically and geographically diverse languages (based on the 100-language sample from the World Atlas of Language Structures). The typological database records the logical space of deponency: what features may be affected, and what are the characteristics of the resulting paradigm? Every logical combination of parameters is represented by one exemplar (or where none has been found, this is noted too). The typological database is supplemented by a set of formal analyses of examples which hold particular interest for morphological theory.
Automatic analysis of syntax is one of the core problems in natural language processing. Despite significant advances in syntactic parsing of written text, the application of these techniques to spontaneous spoken language has received more limited attention. The recent explosive growth of online, accessible corpora of spoken language interactions opens up new opportunities for the development of high accuracy parsing approaches to the analysis of spoken language. The availability of high accuracy parsers will in turn provide a platform for development of a wide range of new applications, as well as for advanced research on the nature of conversational interactions. One concrete field of investigation that is ripe for the application of such parsing tools is the study of child language acquisition. In this thesis, we describe an approach for analyzing the syntactic structure of spontaneous conversational language in parent-child interactions. Specific emphasis is placed on the challenge of accurately annotating the English corpora in the CHILDES database with grammatical relations (such as subject, objects and adjuncts) that are of particular interest and utility, to researchers in child language acquisition. This work involves rule-based and corpus-based natural language processing techniques, as well as methodology for combining results from different parsing approaches. We present novel strategies for integrating the results of different parsers into a system with improved accuracy. One practical application of this research is the automation of language competence measures used by clinicians and researchers of child language development. We present an implementation of an automatic version of one such measurement scheme. This provides not only a useful tool for the child language research community, but also a task-based evaluation framework for grammatical relation identification. Through experiments using data from the Penn Treebank, we show that several of the techniques and ideas presented in this thesis are applicable not just to analysis of parent-child dialogs, but to parsing in general.
The paper focuses on the description of the process of “mining” lexical semantic information from published dictionaries of Brazilian Portuguese language. Specifically, it is described the manual approach of compiling and filtering hyperonymy/hyponymy and holonymy/meronym logic-conceptual relations of concrete nouns. These relations, filtering from the dictionaries, will be use to organize part of the nouns of Wordnet.Br lexical database.
The goal of the present study was to examine the relationships and differences between emotion perceived ( i.e., the emotional quality expressed by music) and emotion felt ( i.e., the individual's emotional response to music). Thirty-two participants listened to 12 music pieces differing in terms of a priori basic emotional quality, and rated the music from two points of view ( i.e., emotion felt and emotion perceived) using 16 adjectives from dimensional models of emotion. As expected, in general, music seemed to arouse emotions similar to the emotional quality perceived in music. However, the affect ratings were significantly moderated by the point of view from which the emotions were assessed. Felt emotions were stronger than perceived emotions in connection with pleasure, but weaker in connection with arousal, positive activation, and negative activation. As also expected, negative perceived quality in music elicited less or an opposite felt emotion. That is, fearful music was perceived as negative but felt as positive.
Since ancient linguistics, the studies of Indo-European word order work with the conception of universal natural word order (ordo naturalis) – an order of the verb-dependent constituents in the linear organization of a clause. The description of the natural word order is usually based on occasional (and in some degree random) observations of clauses in a certain language. In Czech linguistics, the idea of the natural word order was formulated in a more precise way as the hypothesis of the systemic ordering (Sgall, Hajicova and Buraňova, 1980). According to the authors, the contextually non-bound participants and adverbials are ordered as follows (o. c., page 77):
This paper evaluates four of the most commonly used, freely available, state-of-the-art parsers on a standard benchmark as well as with respect to a set of data relevant for measuring text cohesion, as one example of a learning technology application that requires fast and accurate syntactic parsing. We outline advantages and disadvantages of existing technologies and make recommendations. Our performance report uses traditional measures based on a gold standard as well as novel dimensions for parsing evaluation. To our knowledge, this is the first attempt to evaluate parsers across genres and grade levels for the implementation in learning technology using both gold standard and directed evaluation methods.
A grammatical method of combining two kinds of speech repair cues is presented. One cue, prosodic disjuncture, is detected by a decision tree-based ensemble classifier that uses acoustic cues to identify where normal prosody seems to be interrupted (Lickley, 1996). The other cue, syntactic parallelism, codifies the expectation that repairs continue a syntactic category that was left unfinished in the reparandum (Levelt, 1983). The two cues are combined in a Treebank PCFG whose states are split using a few simple tree transformations. Parsing performance on the Switchboard and Fisher corpora suggests that these two cues help to locate speech repairs in a synergistic way.