Papers reviewed and determined not to be word norm studies. Use the flag icon to report errors or suggest re-inclusion.
16504 papers
The paper demonstrates how the generic parser of a minimally supervised information extraction framework can be adapted to a given task and domain for relation extraction (RE). For the experiments a generic deep-linguistic parser was employed that works with a largely hand-crafted head-driven phrase structure grammar (HPSG) for English. The output of this parser is a list of n best parses selected and ranked by a MaxEnt parse-ranking component, which had been trained on a more or less generic HPSG treebank. It will be shown how the estimated confidence of RE rules learned from the n best parses can be exploited for parse reranking. The acquired reranking model improves the performance of RE in both training and test phases with the new first parses. The obtained significant boost of recall does not come from an overall gain in parsing performance but from an application-driven selection of parses that are best suited for the RE task. Since the readings best suited for successful rule extraction and instance extraction are often not the readings favored by a regular parser evaluation, generic parsing accuracy actually decreases. The novel method for task-specific parse reranking does not require any annotated data beyond the semantic seed, which is needed anyway for the RE task.
In the Chinese teaching and research, it often needs to draw syntax tree for the analyzing of relationship among compositions of sentence. Drawn syntax tree manually has many defects, such as having huge workload, needing immense storage capacity and so on. So it is of importance to do research of automatically generating syntax tree and its visualization. In this paper, it proposes a method to generate syntax tree automatically and display the syntax tree in web page by using VML technology. By comparing the syntax trees generated in this paper with the syntax trees got from Treebank of PKU, the result shows that the method of coordinate computing proposed in this paper is more precision, which also demonstrates the efficiency of the visualization method.
Corpus studies by Schuler, AbdelRahman, Miller, and Schwartz (2010), appear to support a model of comprehension taking place in a general-purpose working memory store, by providing an existence proof that a simple probabilistic sequence model over stores of up to four syntacticallycontiguous memory elements has the capacity to reconstruct phrase structure trees for over 99.9% of the sentences in the Penn Treebank Wall Street Journal corpus (Marcus, Santorini, & Marcinkiewicz, 1993), in line with capacity estimates for general-purpose working memory, e.g. by Cowan (2001).But capacity predictions of this simple structure-based model ignore non-structural dependencies, such as long-distance fillergap dependencies, that may place additional demands on working memory.Distinguishing unattached gap fillers from open attachment sites in syntactically-contiguous memory elements requires this contiguity constraint to be strengthened to a constraint that working memory elements be semantically contiguous.This paper presents corpus results showing that this stricter semantic contiguity constraint still predicts working memory requirements in line with capacity estimates such as that of Cowan (2001).
This paper gives a description of an annotation scheme for annotating a corpus of computer-mediated communication in Hindi (CO3H) with certain semantic, pragmatic and situational features. The annotation scheme is based on the theory of register analysis, where it is assumed that a registeral difference entails difference in certain linguistic features. It adapts and integrates the annotation schemes of sense annotation in the Penn Discourse Treebank and dialogue act annotation of DIT++ within this larger registeral framework. The situational and linguistic features that will be used to annotate the corpus for PoRT is described in the paper, along with some proposed labels for these features.
Among the most salient and extensively researched phonological processes of Caribbean Spanish after the categorical weakening of /-s / is the prolific behavior of the implosive or post-nuclear liquids // and /l/. It has been repeatedly claimed yet remarkably unsubstantiated that the gemination of word-medial, post-nuclear liquids to a following consonantal segment is a pervasive characteristic of Cuban Spanish, particularly of the western dialect region. To this end, the fundamental objective of the present study was to acoustically investigate said-phenomenon as it is purported to occur in the province of Havana, whose capital city models the linguistic norm for the rest of the country. Speech samples were elicited from twenty-four native speakers and spectrographic analysis was performed on these collected tokens in order to precisely identify the characteristics of both liquids in the above-mentioned segmental environment. Although we did not find any evidence of liquid gemination in the 120 words under analysis, we did observe two systematic and conditioned processes that may potentially help account for the impressionistic identification of gemination in word-internal position: 1) the insertion of an excrescent vowel between all [.C] sequences; and 2) an increase in the duration of the closure of the stop positionally subsequent to /-L / → [Ø] by an average 30.9%. In light of the evident lack of empirical studies dedicated to the allophony of final liquids, we believe the implications of the findings in this investigation to be important for Caribbean Spanish in general and Cuban Spanish in particular and hope that this study may provide a foundation for future work on the phenomenon of liquid gemination.
FinnWordNet is a wordnet for Finnish that complies with the format of the Princeton WordNet (PWN) (Fellbaum, 1998).It was built by translating the Princeton WordNet 3.0 synsets into Finnish by human translators.It is open source and contains 117000 synsets.The Finnish translations were inserted into the PWN structure resulting in a bilingual lexical database.In natural language processing (NLP), wordnets have been used for infusing computers with semantic knowledge assuming that humans already have a sufficient amount of this knowledge.In this paper we present a case study of using wordnets as an electronic dictionary.We tested whether native Finnish speakers benefit from using a wordnet while completing English sentence completion tasks.We found that using either an English wordnet or a bilingual English-Finnish wordnet significantly improves performance in the task.This should be taken into account when setting standards and comparing human and computer performance on these tasks.
While some visual objects prompt strong affective responses (e.g., guns and ice cream), most objects are thought to be affectively neutral. Last year we reported evidence for the existence of “micro-valences” (Lebrecht & Tarr, VSS, 2010): that nominally neutral objects actually possess subtle valences that we hypothesize form an integral part of object perception. In the current experiment we used fMRI to investigate: a) the extent to which micro-valences are coded within the extended visual object recognition network (Bar, 2007); b) how micro-valences are neurally instantiated with respect to valence strength and direction. Using slow event-related fMRI, participants viewed an object picture for 500ms and evaluated the object's “pleasantness” on each trial. Participants were shown 120 everyday, nominally neutral objects (e.g., teapots and clocks) and 120 strongly valenced objects (e.g., gold and a skull). Objects were assigned to these conditions based on mean valence ratings acquired in a prior experiment with a different population of participants. Individualized ratings for all objects were also acquired for our fMRI participants during a post-scan session. Regions of interest for further analysis were identified using two independent localizers: a) objects versus scrambled objects; b) strongly valenced objects versus minimally valenced objects (e.g., paperclips). Two results stand out. First, somewhat consistent with previous findings, lateral regions of PFC and regions of medial OFC are selective to a positive versus negative comparison for strongly valenced objects. Second, and intriguingly, almost all participants show selectivity for micro-valence objects, comparing positive to negative, in a region adjacent to the region for strongly valenced objects. We posit that intrinsic to visual object perception, object valence – for all objects – is evaluated in PFC. This valence metric forms one of many associated object properties that can influence subsequent perceptual and non-perceptual object-related processing.
Translating is often thought of as transfer of meanings from one linguistic-cultural sphere to another. Often, however, in order for the transfer to succeed the translator needs to revise the target system itself. This is particularly true in software localization. When the concepts in the text are new even for the source culture, translators become the very tools with which the target culture expands and changes. The route to this are the ever-changing norms of language. This article presents some preliminary findings from a study of the Finnish localization team of the KDE SC desktop environment. I will demonstrate the kinds of rhetorical ‘moves’ localizers use when tackling with the linguistic norms and also briefly consider the methodological problems one encounters in investigating this ‘normative work’.
Support vector machines (SVMs) have played a significant role in the field of pattern recognition. This study utilizes the SVM as a classifier for the analysis of Malay cheque word recognition using Malay lexical database (Ahmad et al., 2007). The SVM system was used for individual character recognition and then lexical verification was applied for word level. Several pre-processing steps were taken such as noise removal, image normalization, and skeletonization prior to feature extraction to improve the dataset perspective and hence the recognition accuracy. Statistical and geometrical extraction techniques have been applied in the approach. The results show that the statistical feature is reliable, accessible and provides more accurate results. The results also show that the new approach passed 97.15% character recognition, and combined with word lexical verification, the recognition rate surpassed 98.2% recognition rate.
OBJECTIVE: This prospective study compares MRI of atherosclerotic plaque in the abdominal aorta at 3 T with that at 1.5 T in patients suffering from hereditary hyperlipidaemia, a major risk factor for atherosclerosis. METHODS: MRI of the abdominal aorta at 1.5 and 3 T was performed in 21 patients (mean age 58 years). The study protocol consisted of proton density (PD), T(1), T(2) and fat-saturated T(2) weighted black blood images of the abdominal aorta in corresponding orientation. Two independent radiologists performed image rating. First, image quality was rated on a five-point scale. Second, atherosclerotic plaques were scored according to the modified American Heart Association (AHA) classification and analysed for field strength-related differences. Weighted κ statistics were calculated to assess interobserver agreement. RESULTS: Interobserver agreement was substantial for nearly all categories. MRI at 3 T offered superior image quality in all contrast weightings, most significantly in T(1) and T(2) weighted techniques. Plaque burden in the study collective was unexpectedly moderate. The majority of plaques were classified as AHA III lesions; no lesions were classified above AHA V. There was no significant influence of the field strength regarding the AHA classification. CONCLUSION: Abdominal aortal plaque screening is basically feasible at both field strengths, whereas the image quality is rated superior at 3 T. However, the role of the method in clinical practice remains uncertain, since substantial findings in the high-risk collective were scarce.
ABSTRACT \nUCHI-SOTO (INSIDE-OUTSIDE): LANGUAGE AND CULTURE IN \nCONTEXT FOR THE JAPANESE AS A FOREIGN \nLANGUAGE (JFL) LEARNER \nby \n?? Jamie Louise Goekler 2010 \nMaster of Arts in Teaching International Languages \nCalifornia State University, Chico \nFall 2010 \nJapanese as a Foreign Language (JFL) learners have frequently been exposed \nto learning materials which are neither contextualized culturally nor linguistically \nfrom the target language perspective. This research review addresses the cultural and \ncommunicative gaps which exist in many JFL textbooks and enhances JFL students??? \nawareness of similarities and differences between the Japanese culture and their own. \nTo teach JFL from an emic perspective, teachers must first provide students with cultural \nand communicative content that matches target culture and linguistic norms; students \nmust come to recognize the meaning of uchi (insider) and soto (outsider) if they \nare to communicate from an insider perspective. This body of research provides information \non cultural, linguistic, and paralinguistic factors essential to communicative \ncompetence \nx \nin Japanese. This information will help JFL students develop communicative competence \nby becoming linguistic and cultural insiders, viewing Japanese from an emic perspective. \nStudents will also learn about the implications of insider relationships and \nhow they influence language and social relations. This research details uchi-soto relationships, \nhierarchy, honorific language use, communication styles and strategies, gendered \nlanguage, and aidzuchi (Japanese discourse markers and techniques).
In this paper, we argue that there are two seemingly incompatible perceptions of discourse structure: a semantics-centered view and a syntax-centered view. In the semantics-based view, discourse structure is viewed as a structure that identifies the most important portions of the text and describes how they combine semantically. In the syntax-based view, discourse structure is viewed as an extension of syntax to the discourse level, which essentially links the syntactic trees for the individual sentences into one big tree structure. We will argue that these differences in perception may explain some of the central disagreements in the literature about the nature of discourse structure, in particular whether discourse structure is best viewed as a tree or a general graph. However, the two views are not as incompatible as they may seem at first sight, since the semantic discourse structure can be reinterpreted as a functor-argument structure that is derived from the syntactic tree structure. We describe the ramifications of the two views for the analysis of discourse markers, which are the focus of the discourse annotation in the Penn Discourse Treebank, and show how the syntax-based view can maintain a tree structure even for examples that seem to exhibit non-tree like properties in a semantics-based view.
The three syllable-final nasals /m/, /n/, and /ŋ/ in old Chinese have merged into two (/n, ŋ/) or only one (/ŋ/) nasal in modern Chinese languages. The perceptual confusability of place of articulation was investigated with speakers of Southern Min, a Chinese language which preserves all three final nasals. Three experiments of forced-choice nasal-identification were conducted: (1) complete CVN syllables embedded in noise, (2) CV-truncations of the CVN syllables, and (3) the excised nasal murmur, −N. The first experiment revealed that /m/ was the most and /n/ was the least confusable. Responses were highly accurate in the second experiment (above 85%) and around chance in the third experiment (below 40%), which indicated that listeners relied on the information in the vowel rather than the nasal murmur to identify final nasals. The vowel /i/ resulted in the most and /a/ the least misidentification of final nasals among /i, ə, a/. Low-level and falling tones resulted in more misidentification of final nasals than mid-level, high-level, and rising tones. Segment duration, F0, and formant transitions are analyzed across the vowel and tone types for insight into these findings. Lexical familiarity ratings did not show significant correlation with the perceptual results.
Supervised learning algorithms often require large amounts of labeled data.Creating this data can be time consuming and expensive.Recent work has used untrained annotators on Mechanical Turk to quickly and cheaply create data for NLP tasks, such as word sense disambiguation, word similarity, machine translation, and PP attachment.In this experiment, we test whether untrained annotators can accurately perform the task of POS tagging.We design a Java Applet, called the Interactive Tagging Guide (ITG) to assist untrained annotators in accurately and quickly POS tagging words using the Penn Treebank tagset.We test this Applet on a small corpus using Mechanical Turk, an online marketplace where users earn small payments for the completion of short tasks.Our results demonstrate that, given the proper assistance, untrained annotators are able to tag parts of speech with approximately 90% accuracy.Furthermore, we analyze the performance of expert annotators using the ITG and discover nearly identical levels of performance as compared to the untrained annotators.
We employ syntactic parsing to describe and to discover lexico-grammatical features of English regional varieties. In the absence of suitable Treebanks, automatically parsed corpora (tree jungles) can be used. As an example we focus on Indian English, using the International Corpus of English (ICE), and the British National Corpus (BNC). We use a largely corpus-driven method. There are few differences in frequencies of syntactic relations between the corpora, but considerable differences when taking the intricate relations between grammar and lexis into account. We describe differences in the use of zero articles, verb-preposition constructions, and ditransitive verbs. We show that relatively small corpora can be used to discover subtle lexico-grammatical differences.
We present an approach for Seman-tic Role Labeling (SRL) using Condi-tional Random Fields in a joint identifi-cation/classification step. The approach is based on shallow syntactic information (chunks) and a number of lexicalized fea-tures such as selectional preferences and automatically inferred similar words, ex-tracted using lexical databases and distri-butional similarity metrics. We use se-mantic annotations from the Proposition Bank for training and evaluate the system using CoNLL-2005 test sets. The addi-tional lexical information led to improve-ments of 15 % (in-domain evaluation) and 12 % (out-of-domain evaluation) on over-all semantic role classification in terms of F-measure. The gains come mostly from a better recall, which suggests that the addi-tion of richer lexical information can im-prove the coverage of existing SRL mod-els even when very little syntactic knowl-edge is available. 1
Animacy is known to play a role in postverbal argument ordering for various languages (W02, B08a, B08b). Together with definiteness, pronominality, and shortness, animacy favors positioning an argument first. Nevertheless, it is sometimes difficult to distinguish linear ordering from grammatical function or semantic role assignment, since those are also subject to various factors including animacy (K77). In order to investigate the role of animacy in French argument ordering, we limit ourselves to complements of ditransitive verbs with object and indirect object nominal complements. Using statistics on treebanks for language production and a questionnaire study for language perception, we show that animacy seems to play no role in the relative ordering.
This article gives a survey of the main issues confronting the compilers of monolingual dictionaries in the age of the Internet. Among others, it discusses the relationship between a lexical database and a monolingual dictionary, the role of corpus evidence, historical principles in lexicography vs. synchronic principles, the instability of word meaning, the need for full vocabulary coverage, principles of definition writing, the role of dictionaries in society, and the need for dictionaries to give guidance on matters of disputed word usage. It concludes with some questions about the future of dictionary publishing.OPSOMMING: Die samestelling van 'n eentalige woordeboek vir moedertaalsprekers. Hierdie artikel gee 'n oorsig van die hoofkwessies waarmee die samestellers van eentalige woordeboeke in die eeu van die Internet te kampe het. Dit bespreek onder andere die verhouding tussen 'n leksikale databasis en 'n eentalige woordeboek, die rol van korpusgetuienis, historiese beginsels vs sinchroniese beginsels in die leksikografie, die onstabiliteit van woordbetekenis, die noodsaak van 'n volledige woordeskatdekking, beginsels van die skryf van definisies, die rol van woordeboeke in die maatskappy, en die noodsaak vir woordeboeke om leiding te gee oor sake van betwiste woordgebruik. Dit sluit af met 'n aantal vrae oor die toekoms van die publikasie van woordeboeke.Sleutelwoorde: EENTALIGE WOORDEBOEKE, LEKSIKALE DATABASIS, WOORDEBOEKSTRUKTUUR, WOORDBETEKENIS, BETEKENISVERANDERING, GEBRUIK, GEBRUIKSAANTEKENINGE, HISTORIESE BEGINSELS VAN DIE LEKSIKOGRAFIE, SINCHRONIESE BEGINSELS VAN DIE LEKSIKOGRAFIE, REGISTER, SLANG, STANDAARDENGELS, WOORDESKATDEKKING, KONSEKWENSIE VAN VERSAMELINGS, FRASEOLOGIE, SINTAGMATIESE PATRONE, PROBLEME VAN KOMPOSISIONALITEIT, LINGUISTIESE PRESKRIPTIVISME, LEKSIKALE GETUIENIS
Unsupervised parsing induction has attracted a significant amount of attention over the last few years. However, current systems exhibit a degree of complexity that can shy away newcomers to the field. We challenge the need for such complexity and present a straightforward weak-EM based system. The results we obtained are close to state-of-the-art ones while still making it extremely simple to experiment with different sub-components. We use a k-best parser, an inductor for Probabilistic Bilexical Grammars (PBGs) [1] and a simple treebank builder. Since our algorithm is independent of the PBG inductor, it overlaps with other models from the literature such as Dependency Model with Valence [2]. Our algorithms are fully fleshed and easily reproducible. We experiment in 8 languages that inform intuitions in training- size dependent parameterization.
The paper focuses on proper names and, more specifically, personal names, toponyms and microtoponyms with a high degree of cultural embeddedness. The author draws on the classification of translation rules and methods (techniques) developed by Ermolovich, with a special focus on difficult cases of transfer where transliteration and transcription are employed. Two directions of translation are discussed: Polish into Russian and Russian into Polish. The paper analyses discrepancies between linguistic norms and usage, where problems originate from the existence of recognised equivalents, arbitrariness of transcription and globalisation (widespread use of the Internet). In the face of these challenges, consistency in a translation project is difficult to retain. The author proposes a set of practical exercises that can be used during translation classes (drawing on the analysis of original texts) in order to make beginner translation students aware of the existing problems and potential solutions. Examples of Polish realia in Russian texts and Russian realia in Polish texts are presented.
In this paper, we present results of an ongoing investigation of a manually aligned parallel treebank and an automatic tree aligner.We establish the features that show a significant correlation with alignment performance.We present those features with the biggest correlation scores and discuss their significance, with mention of future applications of these findings.
While the 2007-2010 financial crisis has hit a variety of countries asymmetrically, the case of Spain is particularly illustrative: this country experienced a pronounced housing bubble partly funded via spectacular developments in its securitization markets leading to looser credit standards and subsequent financial stability problems. We analyze the sequential deterioration of credit in this country considering rating changes in individual securitized deals and on balance sheet bank conditions. Using a sample of 20, 286 observations on securities and rating changes from 2000Q1 to 2010Q1 we build a model in which loan growth, on balancesheet credit quality and rating changes are estimated simultaneously. Our results suggest that loan growth significantly affects on balance-sheet loan performance with a lag of at least two years. Additionally, loan performance is found to lead rating changes with a lag of four quarters. Importantly, bank characteristics (in particular, observed solvency, cash flow generation and cost efficiency) also affect ratings considerably. Additionally, these other bank characteristics seem to have a higher weight in the rating changes of securities issued by savings banks as compared to those issued by commercial banks. JEL Classification: G21, G12
Data integration systems attempt to provide users with seamless and flexible access to information from multiple autonomous, distributed and heterogeneous data sources through a unified query interface. Besides data are continuously growing, maintained by different organizations and managed autonomously, querying data from heterogeneous data sources faces new challenges. As data integration has been automated, the ambiguity in concept interpretation also known as semantic heterogeneity has become one of the main obstacles to this process. Introduction of the Semantic Web Vision Ontologies WordNet ontology [3] is a large lexical database that is used in many schema matching algorithms to match schemas based on the semantics of attributes. In this paper ontology based semantic query reformulation technique is followed to improve the recall of the query. The reformulated query is optimized by removing disjunctive clauses in the query to reduce the computational cost of the semantic query execution. Experimental results show that the proposed optimization technique improves recall with minimal execution time.
Identifying factors that improve the assessment of athletes' psychological functioning is imperative to make proper return-to-play decisions following concussion. Prior research indicates that an individual's affect is related to symptom reporting. The present study examines two novel methods of affect assessment in college athletes at baseline participating in a sports-concussion management program. A total of 256 athletes completed a neuropsychological baseline battery with measurements of psychological symptoms (BDI-Fast Screen, Post-Concussion Symptom Scale, and ImPact Total Symptom Score) and a measure of affective memory bias (the Affective Verbal Learning Test; AVLT). Examiners completed an observation-based rating of affect. Multivariate analysis of variance and χ2 analyses were conducted to examine the effect of affect on symptom reports. Examiners' Affect Ratings were predictive of broad symptom reporting, while the performance based index of affect (Affective Verbal Learning Test, AVLT) was more predictive of depressive symptoms. These findings suggest that performance on the AVLT may be a useful indicator of self-reported depression in a collegiate athlete sample. Additionally, these results demonstrate that examiners' behavioral assessments of affect are important in the assessment of psychological functioning in athletes. Continued work should focus on developing objective measures that are sensitive and valid for the evaluation of outcomes from concussion.
Recent applications including the Semantic Web, Web ontology and XML have sparked a renewed interest on graph-structured databases. Among others, twig queries have been a popular tool for retrieving subgraphs from graph-structured databases. To optimize twig queries, selectivity estimation has been a crucial and classical step. However, the majority of existing works on selectivity estimation focuses on relational and tree data. In this paper, we investigate selectivity estimation of twig queries on possibly cyclic graph data. To facilitate selectivity estimation on cyclic graphs, we propose a matrix representation of graphs derived from prime labeling - a scheme for reachability queries on directed acyclic graphs. With this representation, we exploit the consecutive ones property (C1P) of matrices. As a consequence, a node is mapped to a point in a two-dimensional space whereas a query is mapped to multiple points. We adopt histograms for scalable selectivity estimation. We perform an extensive experimental evaluation on the proposed technique and show that our technique controls the estimation error under 1.3% on XMARK and DBLP, which is more accurate than previous techniques. On TREEBANK, we produce RMSE and NRMSE 6.8 times smaller than previous techniques.
UFAL). Abstract. Annotated corpora such as treebanks are important for the development of parsers, language applications as well as understanding of the language itself. Only very few languages possess these scarce resources. In this paper, we describe our eort in syntactically annotating a small corpora (600 sentences) of Tamil language. Our annotation is similar to Prague Dependency Treebank (PDT 2.0) and consists of 2 levels or layers: (i) morphological layer (m-layer) and (ii) analytical layer (a-layer). For both the layers, we introduce annotation schemes i.e. positional tagging for m-layer and dependency relations (and how dependency structures should be drawn) for a-layers. Finally, we evaluate our corpora in the tagging and parsing task using well known taggers and parsers and discuss some general issues in annotation for Tamil language.
We want to demonstrate on some selected linguistic issues that classical structural and functional linguistics even with its seemingly traditional approaches has something to offer to a formal description of language and its applications in natural language processing and to illustrate by a brief reference to Functional Generative Grammar (on the theoretical side of CL) and Prague Dependency Treebank (on the applicational side) a possible interaction between linguistics and CL.
Prepositional phrase (PP) consists of two parts which are a preposition as the leading part and a word or phrase as the tail part. In accordance with this fact, this paper proposes a new approach for identifying PP. In this method, PP identification is transformed into the collocation identification of preposition itself and the right boundary word. The Cascaded Conditional Random Fields (CCRFs) is used in this approach. With the Penn Chinese Treebank 5.1 as our experiment corpus, the F <sub xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">1</sub> rises to 94.63%. This approach obtains breakthrough in this specific field as the current F <sub xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">1</sub> is about 8.6% higher than any publicly published paper.
This paper is based on the assertions of Theory of Linguistic Variation and Change and proposes a discussion about possible actions of linguistic norms over two variable phenomena in Brazilian Portuguese: the position of clitic pronouns associated with a single verb, and the use of prepositions with verbal complements indicating a “goal/recipient”. We intend to describe each phenomenon by analysing data selected from newspapers from Sao Paulo and Rio Claro published between 1900 and 1915. We compare the descriptions and evaluate the role played by the standard and the common usage which is (already) perceptible in the ‘paulistas’ published pages of that period.
In this paper, we present a simple and effective fine-grained feature generation scheme for dependency parsing. We focus on the problem of grammar representation, introducing fine-grained features by splitting various POS tags to different degrees using HowNet hierarchical semantic knowledge. To prevent the oversplitting, we adopt a threshold-constrained bottomup strategy to merge the derived subcategories. We conduct the experiments on the Penn Chinese Treebank. The results show that, with the fine-grained features, we can improve the dependency parsing accuracies by 0.52 % (absolute) for the unlabeled first-order parser, and in the case of second-order parser, we can improve the dependency parsing accuracies by 0.61% (absolute). 1
The paper presents a preliminary research on possible relations between the syntactic structure and the polarity of a Czech sentence by means of the so-called sentiment analysis of a computer corpus. The main goal of sentiment analysis is the detection of a positive or negative polarity, or neutrality of a sentence (or, more broadly, a text). Most often this process takes place by looking for the polarity items, i.e. words or phrases inherently bearing positive or negative values. These words (phrases) are collected in the subjectivity lexicons and implemented into a computer corpus. However, when using sentences as the basic units to which sentiment analysis is applied, it is always important to look at their semantic and morphological analysis, since polarity items may be influenced by their morphological context. It is expected that some syntactic (and hypersyntactic) relations are useful for the identification of sentence polarity, such as negation, discourse relations or the level of embeddedness of the polarity item in the structure. Thus, we will propose such an analysis for a convenient source of data, the richly annotated Prague Dependency Treebank.
The short-term memory for the flavour of a wine in frequent but nonexpert drinkers was studied in an experiment comparing memory for wine under conditions where participants imagined and remembered a target wine with memory under conditions where participants carried out subsequent imaging and image rating tasks on the wine and on competing mental images. The results show that the cognitive operation of mental image rating produced enhanced memory when the operations were performed on the image of the target wine and somewhat impaired memory when operations were performed on an image of a competing flavour or operations performed on mental images in other domains. It is concluded that sensory traces of flavour may be effectively maintained in short-term memory when supported by appropriate encoding operations.
In this paper we try to present how information technologies as tools for the creation of digital bilingual dictionaries can help the preservation of natural languages. Natural languages are an outstanding part of human cultural values and for that reason they should be preserved as part of the world cultural heritage. We describe our work on the bilingual lexical database supporting the Bulgarian-Polish Online dictionary. The main software tools for the webpresentation of the dictionary are shortly described. We focus our special attention on the presentation of verbs, the richest from a specific characteristics viewpoint linguistic category in Bulgarian.
The Plain Meaning Rule is often assailed on the grounds that it is unprincipled—that it substitutes for careful analysis an interpreter's ad hoc and impressionistic intuition about the meaning of legal texts. But what if judges and lawyers had the means to test their intuitions about plain meaning systematically? Then initial linguistic impressions about the meaning of a legal text might be viewed as hypotheses to be tested, rather than determinative criteria upon which to base important decisions. There exists very little legal scholarship on corpus linguistics—the study of language function and use through large, electronic linguistic databases called corpora—and the role that corpus methods might play in legal interpretation. This omission becomes more and more striking as scholars and jurists (and even the United States Supreme Court) have found themselves persuaded by corpus-based arguments. This Article argues that the plain or ordinary meaning of a given term in a given context is an empirical matter that may be quantified through corpus-based methods. These methods, when applied to questions of legal ambiguity, present significant advantages over existing empirical approaches to plain meaning and over the prevailing intuition-based interpretive approach of many courts. Because large, sophisticated linguistic corpora are widely available and easy to use, and because corpus methods offer a more principled and systematic alternative to the impressionistic interpretation of legal texts, corpus linguistics may one day revolutionize the process of legal interpretation.
Using semi-supervised EM, we learn finegrained but sparse lexical parameters of a generative parsing model (a PCFG) initially estimated over the Penn Treebank. Our lexical parameters employ supertags, which encode complex structural information at the pre-terminal level, and are particularly sparse in labeled data – our goal is to learn these for words that are unseen or rare in the labeled data. In order to guide estimation from unlabeled data, we incorporate both structural and lexical priors from the labeled data. We get a large error reduction in parsing ambiguous structures associated with unseen verbs, the most important case of learning lexico-structural dependencies. We also obtain a statistically significant improvement in labeled bracketing score of the treebank PCFG, the first successful improvement via semi-supervised EM of a generative structured model already trained over large labeled data. 1
This preliminary study explored whether neurophysiological responses to visual stimuli, including attachment-related pictures, differed based on attachment status. Along with self-reported valence ratings and reaction times, recorded electroencephalographic (EEG) responses to a total of 100 images, 25 each of Positive, Negative, Neutral, and Personal (each participant's parents and child), were analyzed within and among three mothers with three attachment statuses (Dismissing, Preoccupied, and Secure), as judged by the Adult Attachment Interview (AAI). All three mothers gave their highest pleasantness ratings for Personal photographs. However, differences emerged when cross-region Alpha2 activation patterns in response to each picture type were compared amongst attachment categories. Alpha2 activation recorded during viewing of the participants' children's photographs was similar to viewing Negative pictures for mothers with insecure (Dismissing and Preoccupied) status; whereas the Alpha2 activation of the mother with Secure status towards photographs of her child was similar to Positive pictures. Different patterns of hemispheric asymmetry in Beta1 frequency when processing different picture types were also found. The mother with Dismissing status showed significantly stronger left-hemisphere Beta1 activation across all image types. The Preoccupied mother showed significantly stronger right-hemisphere Beta1 activation for all but the Neutral images, during which activation did not differ between the two hemispheres. The mother with Secure status showed significantly stronger Beta1 activation in the left hemisphere for all but parental Personal photos, during which activation did not differ between the two hemispheres. Implications from the current findings and future research possibilities are discussed.
A Morphological Analyzer and Generator are two crucial tools involving any Natural Language Processing of Dravidian Languages. The present paper discusses the improvization of the existing Morphological Analyzer and Generator for Tamil by defining and describing the relevant linguistic database required for the purpose of developing them. The implementation of an open source platform called Apertium to handle inflection as well as derivation for word level analysis and generation of Tamil is also discussed. The paper also presents the efficacy, coverage and speed of the module against the large corpora. The paper also draws inferences of the morphological categories in their inflection and problems in analysing them. I. INTRODUCTION A language like Tamil is regarded as morphologically rich wherein the words are formed of one or more stems/roots plus one or more suffixes. So the complexity of morphology requires a more sophisticated morphological analyzer and generator. A morphological analyzer is a computational tool to analyze word forms into their roots along with their constituent functional elements. The morphological generator is the reverse process of an analyzer i.e. from a given root and functional elements, it generates the well-formed word forms. The present attempt involves a practical adoption of lttoolbox for the Modern Standard Written Tamil in order to develop an improvised open source morphological analyzer and generator. The tool uses the computational algorithm called Finite State Transducers for one-pass analysis and generation, and the database is based on the morphological model called Word and Paradigm.
Nous presentons une architecture pour l’analyse syntaxique en deux etapes. Dans un premier temps un analyseur syntagmatique construit, pour chaque phrase, une liste d’analyses qui sont converties en arbres de dependances. Ces arbres sont ensuite reevalues par un reordonnanceur discriminant. Cette methode permet de prendre en compte des informations auxquelles l’analyseur n’a pas acces, en particulier des annotations fonctionnelles. Nous validons notre approche par une evaluation sur le corpus arbore de Paris 7. La seconde etape permet d’ameliorer significativement la qualite des analyses retournees, quelle que soit la metrique utilisee.
Syntactic structures have been good features for opinion analysis, but it is not easy to use them. To find these features by supervised learning methods, correct syntactic labels are indispensible. Two possible sources to acquire syntactic structures are parsing trees and dependency trees. For the annotation processing, parsing trees are more readable for annotators, while dependency trees are easier to use by programs. To use syntactic structures as features, this paper tried to annotate on human friendly materials and transform these annotations to the corresponding machine friendly materials. We annotated the gold answers of opinion syntactic structures on the parsing tree from Chinese Treebank, and then proposed methods to find their corresponding dependency relations on the dependency trees generated from the same sentence. With these relations, we could train a model to annotate opinion dependency relations automatically to provide an opinion dependency parser, which is language independent if language resources are incorporated. Experiment results show that the annotated syntactic structures and their corresponding dependency relations improve at least 8% of the performance of opinion analysis.
This paper examines how W.E.B. DuBois' concept of double consciousness influenced the interactions of 13 Black youth inside an after school Community Literacy Intervention Program (CLIP). Du Bois, a pre-eminent 20th century Black sociologist, used double consciousness as a lens to help explain social and psychological tensions that African Americans encounter while negotiating their identities in a societal context structured mainly upon dominant white cultural and linguistic norms and values. The authors provide a conceptual framework for understanding the interpretive processes that signify double consciousness which includes: surveying the context; assessing risks and identity consequences; articulating mainstream or race conscious reads, and bridging/or disengaging. Implications for pre-service teachers and particularly urban educators are discussed.
This paper presents a method for word sense disambiguation based on Lesk algorithm which uses lexical database WordNet as knowledge base. The word sense disambiguation is the process of automatically clarifying a meaning of a word in its context. In general, ontology means the meaning or analogous term. It can be interpreted by relating the word with other words in the sentence. This tool accepts English statement as input and gives best possible meaning of given word. Method is experimented with senseval-2 test data for lexical sample task. The results show the betterment over the original Lesk algorithm.
We consider a very simple, yet effective, approach to cross language adaptation of dependency parsers. We first remove lexical items from the treebanks and map part-of-speech tags into a common tagset. We then train a language model on tag sequences in otherwise unlabeled target data and rank labeled source data by perplexity per word of tag sequences from less similar to most similar to the target. We then train our target language parser on the most similar data points in the source labeled data. The strategy achieves much better results than a non-adapted baseline and stateof-the-art unsupervised dependency parsing, and results are comparable to more complex projection-based cross language adaptation algorithms. 1
The poor grammatical output of Machine Translation (MT) systems appeals syntax-based approaches within language modeling. However, previous studies showed that syntax-based language modeling using (Context-Free) Treebank Grammars was not very helpful in improving BLEU scores for Chinese-English machine translation. In this article we further study this issue in the context of Chinese-English syntax-based Statistical Machine Translation (SMT) where Synchronous Tree Substitution Grammars (STSGs) are utilized to model the translation process. In particular, we develop a Tree Substitution Grammar-based language model for syntax-based MT, and present three methods to efficiently integrate the proposed language model into MT decoding. In addition, we design a simple and effective method to adapt syntax-based language models for MT tasks. We demonstrate that the proposed methods are able to benefit a state-of-the-art syntax-based MT system. On the NIST Chinese-English MT evaluation corpora, we finally achieve an improvement of 0.6 BLEU points over the baseline.
AIM: To evaluate a standardised MRI acquisition protocol and a new image rating scale for disease severity in patients with progressive supranuclear palsy (PSP) and multiple systems atrophy (MSA) in a large multicentre study. METHODS: The MRI protocol consisted of two-dimensional sagittal and axial T1, axial PD, and axial and coronal T2 weighted acquisitions. The 32 item ordinal scale evaluated abnormalities within the basal ganglia and posterior fossa, blind to diagnosis. Among 760 patients in the study population (PSP = 362, MSA = 398), 627 had per protocol images (PSP = 297, MSA = 330). Intra-rater (n = 60) and inter-rater (n = 555) reliability were assessed through Cohen's statistic, and scale structure through principal component analysis (PCA) (n = 441). Internal consistency and reliability were checked. Discriminant and predictive validity of extracted factors and total scores were tested for disease severity as per clinical diagnosis. RESULTS: Intra-rater and inter-rater reliability were acceptable for 25 (78%) of the items scored (≥ 0.41). PCA revealed four meaningful clusters of covarying parameters (factor (F) F1: brainstem and cerebellum; F2: midbrain; F3: putamen; F4: other basal ganglia) with good to excellent internal consistency (Cronbach α 0.75-0.93) and moderate to excellent reliability (intraclass coefficient: F1: 0.92; F2: 0.79; F3: 0.71; F4: 0.49). The total score significantly discriminated for disease severity or diagnosis; factorial scores differentially discriminated for disease severity according to diagnosis (PSP: F1-F2; MSA: F2-F3). The total score was significantly related to survival in PSP (p<0.0007) or MSA (p<0.0005), indicating good predictive validity. CONCLUSIONS: The scale is suitable for use in the context of multicentre studies and can reliably and consistently measure MRI abnormalities in PSP and MSA. Clinical Trial Registration Number The study protocol was filed in the open clinical trial registry (http://www.clinicaltrials.gov) with ID No NCT00211224.
Flat noun phrase structure was, up until recently, the standard in annotation for the Penn Treebanks. With the recent addition of internal noun phrase annotation, dependency parsing and applications down the NLP pipeline are likely affected. Some machine translation systems, such as TectoMT, use deep syntax as a language transfer layer. It is proposed that changes to the noun phrase dependency parse will have a cascading effect down the NLP pipeline and in the end, improve machine translation output, even with a reduction in parser accuracy that the noun phrase structure might cause. This paper examines this noun phrase structure’s effect on dependency parsing, in English, with a maximum spanning tree parser and shows a 2.43%, 0.23 Bleu score, improvement for English to Czech machine translation. 1