Papers reviewed and determined not to be word norm studies. Use the flag icon to report errors or suggest re-inclusion.
18265 papers
The reconstruction of standardized texts in the Prague Dependency Treebank of Spoken Czech enables the comparison of authentic spoken utterances and standardized texts The authors concentrate on the questions: What does the syntactic identity of the Czech spoken and written texts consist of? What syntactic constructions are „natural“ in the spoken and in the written text? What is the difference in the density of the cohesive links, in the explicit and implicite relations between units?
Treebanks are a necessary prerequisite for many NLP tasks, including, but not limited to, semantic role labeling. For many languages, however, treebanks are either nonexistent or too small to be useful. Time-critical applications may require rapid deployment of natural language software for a new critical language—much faster than the development time of a traditional treebank. This dissertation describes a method for generating a treebank and training syntactic and semantic models using only semantic training information—that is, no human-annotated syntactic training data whatsoever. This will greatly increase the speed of development of natural language tools for new critical languages in exchange for a modest drop in overall accuracy. Using Combinatory Categorial Grammar (CCG) in concert with Propbank semantic role annotations allows us to accurately predict lexical categories in combination with a partially hidden Markov model. By training the Berkeley parser on our generated syntactic data, we can achieve SRL performance of 65.5% without using a treebank, as opposed to 74% using the same feature set with gold-standard data.
This study has a dual focus in that it aims to develop a viable methodology for elicitation experiments in English linguistics, while simultaneously applying the proposed methods to investigate an actual subject, the distribtion of the additive particles 'also' and 'too'. Traditionally, data for linguistic research is gained by sampling natural language corpora. Although this approach is valid and, indeed, has been applied here, elicitation experiments can gain in validity and informative value by additionally introducing questionnaires to accompany corpus research. Online questionnaires particularly are a cost-effective and highly customizable tool to create a linguistic database against which existing data can be tested. For the purpose of this study, I have created six online questionnaires to test three hypotheses about the distribution of 'also' and 'too'. Two interdependent hypotheses assume that the use of the two particles is sensitive to structural properties of the `added constituent' while the third one, the information-structural hypothesis, argues that the use of 'also' and 'too' is controlled by the information structure of the sentence. In addition to the questionnaires, a balanced sample was extracted from the "British National Corpus" and tested against corpus data from previous studies as well as the data elicited online. In the course of this study, the additive particles will firstly be defined in terms of their structural properties, and the hypotheses about their use introduced and explicated. Furthermore, the data elicitation process will be detailed, as well as results from previous studies be taken into account. The hypotheses will subsequently be tested against the data from both corpus research and elicitation per questionnaires, and the outcome discussed. Concluding the study, I will focus on the results of the distribution analysis as well as evaluate the introduction of the online questionnaires and their application in the context of testing the hypotheses against empirical linguistic data.
In this paper, we expand Morzycki (2009)’s claims that degree readings of size adjectives are attributed to syntax. We introduce a corpus-based analysis in Dutch to verify and extend his claim into the semantic domain. Using the LASSY Treebank, we extract syntactic and semantic properties of noun phrases consisting of the adjectives “gigantisch”, “kolossaal”, and “reusachtig ” and manually annotate each adjective-noun pair with a gradable or nongradable label. Using these features, we construct a statistical model based on logistic regression and find that the grammatical role, definiteness, and particular semantic noun groups derived from Cornetto (a Dutch WordNet with referential relations) have a significant effect on the likelihood that an adjective-noun pair is interpreted by the reader to have a degree reading. 1.
This paper proposes a method for shallow parsing on the basis of CRF and transformation-based error-driven learning.The method is applied to Penn Chinese Treebank and gets a good performance of chunking identification.First,CRF model is used to identify chunks to acquire candidate transformation rules by error-driven learning.Then,an evaluation function is used to filter candidate transformation rules.And last,transformation rules are used to revise the chunking results of CRF.The experimental results show that this approach is effective,and outperforms the single CRF-based approach in shallow parsing.Precision,recall and F-values are improved respectively.
While some visual objects prompt strong affective responses (e.g., guns and ice cream), most objects are thought to be affectively neutral. Last year we reported evidence for the existence of “micro-valences” (Lebrecht & Tarr, VSS, 2010): that nominally neutral objects actually possess subtle valences that we hypothesize form an integral part of object perception. In the current experiment we used fMRI to investigate: a) the extent to which micro-valences are coded within the extended visual object recognition network (Bar, 2007); b) how micro-valences are neurally instantiated with respect to valence strength and direction. Using slow event-related fMRI, participants viewed an object picture for 500ms and evaluated the object's “pleasantness” on each trial. Participants were shown 120 everyday, nominally neutral objects (e.g., teapots and clocks) and 120 strongly valenced objects (e.g., gold and a skull). Objects were assigned to these conditions based on mean valence ratings acquired in a prior experiment with a different population of participants. Individualized ratings for all objects were also acquired for our fMRI participants during a post-scan session. Regions of interest for further analysis were identified using two independent localizers: a) objects versus scrambled objects; b) strongly valenced objects versus minimally valenced objects (e.g., paperclips). Two results stand out. First, somewhat consistent with previous findings, lateral regions of PFC and regions of medial OFC are selective to a positive versus negative comparison for strongly valenced objects. Second, and intriguingly, almost all participants show selectivity for micro-valence objects, comparing positive to negative, in a region adjacent to the region for strongly valenced objects. We posit that intrinsic to visual object perception, object valence – for all objects – is evaluated in PFC. This valence metric forms one of many associated object properties that can influence subsequent perceptual and non-perceptual object-related processing.
Emotion recognition algorithms for spoken dialogue applications typically employ lexical models that are trained on labeled in-domain data. In this paper, we propose a domainindependent approach to affective text modeling that is based on the creation of an affective lexicon. Starting from a small set of manually annotated seed words, continuous valence ratings for new words are estimated using semantic similarity scores and a kernel model. The parameters of the model are trained using least mean squares estimation. Word level scores are combined to produce sentence-level scores via simple linear and non-linear fusion. The proposed method is evaluated on the SemEval news headline polarity task and on the ChIMP politeness and frustration detection dialogue task, achieving state-of-theart results on both. For politeness detection, best results are obtained when the affective model is adapted using in domain data. For frustration detection, the domain-independent model and non-linear fusion achieve the best performance. Index Terms: language understanding, emotion, affect, affective lexicon
OBJECTIVES: This research compared sensory processing and personality traits involved in deciding to try a novel fruit (guava) in adults and children. DESIGN: The research employed an age, sex, and food neophobia matched between-participant design to examine sensory decision making in choosing to eat a novel fruit. METHODS: Forty-four adults (Study 1) and 68 children (Study 2) took part. In each study, participants were separated into two groups to investigate whether prior assessment of a familiar and liked fruit (apple) that shares similar visual characteristics to the target novel fruit (guava) increased the likelihood that an individual would decide to try it. All participants completed appetitive and familiarity ratings by sensory stages: vision, smell, and touch, prior to trying (tasting) the fruit. Participants (or their parents) also completed the general and food neophobia scales and adults also completed the sensation-seeking scale. RESULTS: Twenty-eight adults (64%) tried the guava and 16 did not (36%). In the second study, 22 children decided not to try the novel fruit (32%). Significant predictors of whether the adult tried the target fruit were Thrill and Adventure Seeking, Experience Seeking, General Neophobia, and 'appealing to touch'. In children, Food Neophobia, concurrent presentation of a familiar fruit alongside the target and visual assessment of the target predicted decision to try the novel fruit. CONCLUSIONS: This study suggests that touch is pertinent to adults' decision to try a novel fruit, whereas visual cues appear to be more important for children.
This study investigated cognitive and emotional effects of syncopation, a feature of musical rhythm that produces expectancy violations in the listener by emphasising weak temporal locations and de-emphasising strong locations in metric structure. Stimuli consisting of pairs of unsyncopated and syncopated musical phrases were rated by 35 musicians for perceived complexity, enjoyment, happiness, arousal, and tension. Overall, syncopated patterns were more enjoyed, and rated as happier, than unsyncopated patterns, while differences in perceived tension were unreliable. Complexity and arousal ratings were asymmetric by serial order, increasing when patterns moved from unsyncopated to syncopated, but not significantly changing when order was reversed. These results suggest that syncopation influences emotional valence (positively), and that while syncopated rhythms are objectively more complex than unsyncopated rhythms, this difference is more salient when complexity increases than when it decreases. It is proposed that composers and improvisers may exploit this asymmetry in perceived complexity by favoring formal structures that progress from rhythmically simple to complex, as can be observed in the initial sections of musical forms such as theme and variations.
Part-of-speech (POS) is an indispensable feature in dependency parsing. Current research usually models POS tagging and dependency parsing independently. This may suffer from error propagation problem. Our experiments show that parsing accuracy drops by about 6 % when using automatic POS tags instead of gold ones. To solve this issue, this paper proposes a solution by jointly optimizing POS tagging and dependency parsing in a unique model. We design several joint models and their corresponding decoding algorithms to incorporate different feature sets. We further present an effective pruning strategy to reduce the search space of candidate POS tags, leading to significant improvement of parsing speed. Experimental results on Chinese Penn Treebank 5 show that our joint models significantly improve the state-of-the-art parsing accuracy by about 1.5%. Detailed analysis shows that the joint method is able to choose such POS tags that are more helpful and discriminative from parsing viewpoint. This is the fundamental reason of parsing accuracy improvement. 1
Previous research suggests that neural and behavioral responses to surprised faces are modulated by explicit contexts (e.g., "He just found $500"). Here, we examined the effect of implicit contexts (i.e., valence of other frequently presented faces) on both valence ratings and ability to detect surprised faces (i.e., the infrequent target). In Experiment 1, we demonstrate that participants interpret surprised faces more positively when they are presented within a context of happy faces, as compared to a context of angry faces. In Experiments 2 and 3, we used the oddball paradigm to evaluate the effects of clearly valenced facial expressions (i.e., happy and angry) on default valence interpretations of surprised faces. We offer evidence that the default interpretation of surprise is negative, as participants were faster to detect surprised faces when presented within a happy context (Exp. 2). Finally, we kept the valence of the contexts constant (i.e., surprised faces) and showed that participants were faster to detect happy than angry faces (Exp. 3). Together, these experiments demonstrate the utility of the oddball paradigm to explore the default valence interpretation of presented facial expressions, particularly the ambiguously valenced facial expression of surprise.
There has been a rapid increase in the volume of research on data-driven dependency parsers in the past five years. This increase has been driven by the availability of treebanks in a wide variety of languages—due in large part to the CoNLL shared tasks—as well as the straightforward mechanisms by which dependency theories of syntax can encode complex phenomena in free word order languages. In this article, our aim is to take a step back and analyze the progress that has been made through an analysis of the two predominant paradigms for data-driven dependency parsing, which are often called graph-based and transition-based dependency parsing. Our analysis covers both theoretical and empirical aspects and sheds light on the kinds of errors each type of parser makes and how they relate to theoretical expectations. Using these observations, we present an integrated system based on a stacking learning framework and show that such a system can learn to overcome the shortcomings of each non-integrated system.
This paper proposes a method to improve the accuracy of bilingual texts (bitexts) dependency parsing by using an auto-generated bilingual treebank created with the help of statistical machine translation (SMT) systems. Previous bitext parsing methods use human-annotated bilingual treebanks that are costly and troublesome to obtain. In the proposed method, we use an auto-generated bilingual treebank to train the parsing models. First, an SMT system is used to translate a monolingual treebank into the target language; then, a monolingual parser for the target language is used to parse the translated sentences. Since the auto-translated sentences and auto-parsed trees in the auto-generated bilingual treebank are far from perfect, the bilingual constraints are not sufficiently reliable. To overcome this problem, we propose a method to verify the reliability of the constraints using a large amount of target monolingual and bilingual unannotated data. Finally, we design a set of effective bilingual features for parsing models on the basis of the verified constraints. We conduct the experiments using a standard test data. The experimental results show that our bitext parser significantly outperforms monolingual parsers. Moreover, our method is still able to provide improvement when we use a larger monolingual treebank containing over 50 000 sentences. We also test the proposed method with different SMT systems and the results show that our method is very robust to the noise. In particular, the proposed method can be used in a purely monolingual setting with the help of SMT. That is, it does not need the human translation of the test set as previous methods do.
Transition-based dependency parsers generally use heuristic decoding algorithms but can accommodate arbitrarily rich feature representations. In this paper, we show that we can improve the accuracy of such parsers by considering even richer feature sets than those employed in previous systems. In the standard Penn Treebank setup, our novel features improve attachment score form 91.4 % to 92.9%, giving the best results so far for transitionbased parsing and rivaling the best results overall. For the Chinese Treebank, they give a signficant improvement of the state of the art. An open source release of our parser is freely available.
Text clustering is of substantial importance to information retrieval.The method of applying the information of syntactic distribution to text clustering is presented,in order to avoid the complex clustering algorithm whileenabling the linguistic interpretation of clustering features and the results of clustering.According to the dependency Treebank,ten dependency relations are suggested with distinctive distribution between oral and written Chinese By using five of them as clustering feature,the similarity of spoken and written classes achieves 71.98% and 83.13%,respectively.The experiment result shows that the proposed method of applying dependency relations to text clustering is feasible and effective.
For the task of automatic treebank conversion, this paper presents a feature-based approach which encodes bracketing structures in a treebank into features to guide the conversion of this treebank to a different standard. Experiments on two Chinese treebanks show that our approach improves conversion accuracy by 1.31 % over a strong baseline. 1
Summary.-Images of pleasant scenes usually produce increased activity over the zygomaticus major muscie, as measured by electromyography (EMG), while less activity is elicited by unpleasant images. However, increases in zygomaticus major EMG activity while viewing unpleasant images have occasionally been reported in the literature on affective facial expression (i.e., "grimacing"). To examine the possibility that individual differences in emotion regulation might be responsible for this inconsistently observed phenomenon, the habitual emotion regulation tendencies of 63 participants (32 women) were assessed and categorized according to their regulatory tendencies. Participants viewed emotionally salient images while zygomaticus major EMG activity was recorded. Participants also provided self-report ratings of their experienced emotional valence and arousal while viewing the pictures. Despite demonstrating intact affective ratings, the "grimacing" pattern of zygomaticus major activity was observed in those who were less likely to use the cognitive reappraisal strategy to regulate their emotions.
We discuss cognitive topics related to an Intelligence Analyst's Geospatial and Ontological Assistant (IAGOA) under development that associates an analyst's understanding of an agent's activities with the geospatial features of the area where they take place. IAGOA uses linguistic resources from the lexical database FrameNet, based on Fillmore's frame semantics (not related to Minsky's frames), and it provides clear explications of the central notions of situation awareness that are particularly appropriate when the focus is on the activities of a single agent.
In the Pralex lexical database, antonyms are preserved as an explanatory part in the same way as in the Dictionary of the Standard Czech Language, except for the entries, which require a modification of the meaning explanation (e.g. some word-formative variants, homonyms etc.). Members of the antonym pairs are hypertextually linked while using The List of Antonyms, which is available for every entry at the level of particular meanings.
The idea of the Czech Academic Corpus (CAC) came to life in 1971 thanks to the Department of Mathematical Linguistics within the Czech Language Institute. By the mid 1980s, a total of 540,000 words were morphologically and syntactically annotated manually. After the Prague Dependency Treebank (PDT) – the largest annotated treebank of Czech written texts – was built, the conversion from CAC to PDT format began. The main goal was to make the CAC and the PDT compatible, and thus to enable the integration of the CAC into the PDT. The second version of the CAC is thus a complete conversion of the internal format and annotation schemes. The conversion of syntactic annotation began three years after the syntactic annotation of PDT was finished. Such a situation is exceptional because, to our knowledge, there is no other language for which such a significant amount of data is being annotated in two subsequent projects. This article summarizes the experience acquired during the conversion of the CAC syntactic annotation.
This paper briefly depicts major achievements in the Chinese language information processing and roughly reviews the computational linguistic research for the recent 20 years in China.The author questions the current methodologies such as POS tagging and treebank for Chinese.The paper presents some new ideas about the construction of Chinese data resources.The authors suggest that for Chinese we should address deep and semantic annotation instead of current shallow and syntactic one.The future annotation will include targeted tackling,diversified content,stepwise procedure,and non-professional annotators.The paper predicts some eye-catching features: merge of technologies and human-centered computing.
State of the art Word Sense Disambiguation (WSD) systems require large sense-tagged corpora along with lexical databases to reach satisfactory results. The number of English language resources for developed WSD increased in the past years, while most other languages are still under-resourced. The situation is no different for Dutch. In order to overcome this data bottleneck, the DutchSemCor project will deliver a Dutch corpus that is sense-tagged with senses from the Cornetto lexical database. Part of this \ncorpus (circa 300K examples) is manually tagged. The remainder is automatically tagged using different WSD systems and validated by human annotators. The project uses existing corpora compiled in other projects; these are extended with Internet \nexamples for word senses that are less frequent and do not (sufficiently) appear in the corpora. We report on the status of the project and the evaluations of the WSD systems with the current training data.
In this paper we describe a mechanism for parallel treebank generation between an intense studied language (i.e. English) and a less studied language, like Romanian. The Romanian constituents of the treebank are induced from the corresponding constituents of the English part taking into account the words alignments of the corpus. The proposed mechanism reuses and adjusts existing tools and algorithms for automatic Part-Of-Speech annotation and syntactic trees alignment.
This paper describes the development, composition, and several uses of the Ancient Greek and Latin Dependency Treebanks, large collections of Classical texts in which the syntactic, morphological and lexical information for each word is made explicit. To date, over 200 individuals from around the world have collaborated to annotate over 350,000 words, including the entirety of Homer’s Iliad and Odyssey, Sophocles’ Ajax, all of the extant works of Hesiod and Aeschylus, and selections from Caesar, Cicero, Jerome, Ovid, Petronius, Propertius, Sallust and Vergil. While perhaps the most straightforward value of such an annotated corpus for Classical philology is the morphosyntactic searching it makes possible, it also enables a large number of downstream tasks as well, such as inducing the syntactic behavior of lexemes and automatically identifying similar passages between texts.
La ressource présentée dans cet article combine un corpus de noms déverbaux annotés sémantiquement et syntaxiquement, fondé sur le French Treebank, et un lexique électronique fournissant des informations d'ordre morphologique, syntaxique et sémantique sur les noms présents dans le corpus ainsi que sur les verbes dont ils sont dérivés. ABSTRACT. The resource presented in this paper combines a semantically and syntactically anno-tated corpus of deverbal nouns based on the French Treebank, and an electronic lexicon, pro-viding descriptions of morphological, syntactic and semantic properties of the deverbal nouns found in our corpus and of their verbal sources. MOTS-CLÉS: corpus annoté, lexique fondé sur corpus, nominalisation, aspect lexical, structure argumentale.
Eighteenth-century language usage is markedly under-represented in the first two editions of the OED, whose quotations for this period were gathered almost entirely during the late nineteenth and early twentieth centuries. This article reviews some of the possible causes, characteristics and consequences of OED’s gap in eighteenth-century documentation and shows that female authors were particularly scanted. The role of quotations in the OED, as the evidential basis for the dictionary, is briefly considered, along with eighteenth-century (and Victorian/Edwardian) views on women and language, and the availability of female-authored texts for quotation by the lexicographers. The article reports sample reading in eighteenth-century female writers (especially Jean Adam, Penelope Aubin and Anna Seward), which shows that OED could easily have supplied its eighteenth-century deficiency from such authors, and that it often favoured distinctive usages in female-authored texts—innovative, eccentric or domestic vocabulary—rather than usage which exemplified linguistic norms (especially in poetry, where Seward’s case is examined). It also discusses revisions to the OED so far conducted in the third (ongoing) edition, and their implications for readers and editors of eighteenth-century texts.
Abstract The aim of computational semantics is to capture the meaning of natural language expressions in representations suitable for performing inferences, in the service of understanding human language in written or spoken form. First‐order logic is a good starting point, both from the representation and inference point of view. But even if one makes the choice of first‐order logic as representation language, this is not enough: the computational semanticist needs to make further decisions on how to model events, tense, modal contexts, anaphora and plural entities. Semantic representations are usually built on top of a syntactic analysis, using unification, techniques from the lambda‐calculus or linear logic, to do the book‐keeping of variable naming. Inference has many potential applications in computational semantics. One way to implement inference is using algorithms from automated deduction dedicated to first‐order logic, such as theorem proving and model building. Theorem proving can help in finding contradictions or checking for new information. Finite model building can be seen as a complementary inference task to theorem proving, and it often makes sense to use both procedures in parallel. The models produced by model generators for texts not only show that the text is contradiction‐free; they also can be used for disambiguation tasks and linking interpretation with the real world. To make interesting inferences, often additional background knowledge is required (not expressed in the analysed text or speech parts). This can be derived (and turned into first‐order logic) from raw text, semi‐structured databases or large‐scale lexical databases such as WordNet. Promising future research directions of computational semantics are investigating alternative representation and inference methods (using weaker variants of first‐order logic, reasoning with defaults), and developing evaluation methods measuring the semantic adequacy of systems and formalisms.
The dictionaries are among the wellknown tools for applications in everyday life, education, sciences, humanities, and human communication. The recent developments of information technologies contribute to the design and creation of new software tools with a wide range of applications, especially for natural language processing. The paper presents an online bilingual dictionary as a technological tool for applications in digital humanities, and describes the structure and content of the bilingual Lexical Database supporting a Bulgarian-Polish online dictionary. It focuses especially on the presentation of verbs, which form the richest from a specific characteristics viewpoint linguistic category in Bulgarian. The main software modules for webpresentation of this digital dictionary are also shortly described. 1
This paper is about the development of Pashto Treebank in the form of Extensible Markup Language (XML) code. A Chart Parser has been developed that uses Chart Parsing Algorithm for building parse trees for Pashto sentences. The output of the parser is the parsed text which can be obtained in one of its three forms such as reduced graph, parse tree and XML code. For parsing, the parser needs Context Free Grammar (CFG) of Pashto language and Tagged Input Text as input. The system has been tested on real world text taken from Pashto novels and web sites and tagged manually. Eighty seven (87) sentences were parsed by the parser in which fifty four (54) were correctly parsed with a single parse tree and the rest 33 were parsed with multiple trees and thus the accuracy obtained is 62.06%.
We introduce synchronous tree adjoining grammars (TAG) into tree-to-string transla-tion, which converts a source tree to a target string. Without reconstructing TAG deriva-tions explicitly, our rule extraction algo-rithm directly learns tree-to-string rules from aligned Treebank-style trees. As tree-to-string translation casts decoding as a tree parsing problem rather than parsing, the decoder still runs fast when adjoining is included. Less than 2 times slower, the adjoining tree-to-string system improves translation quality by +0.7 BLEU over the baseline system only al-lowing for tree substitution on NIST Chinese-English test sets. 1
Good Dictionary Examples or GDEX is a tool in the Sketch Engine designed to help lexicographers with identifying dictionary examples by ranking sentences according to how likely they are to be good candidates. The ranking is done automatically using various syntactic and lexical features. So far, only GDEX for English has been available. This paper presents the design and evaluation of Slovene GDEX, which was used for finding good examples for the new lexical database of Slovene, one of the activities in the Communication in Slovene project. Several different GDEX configurations were designed, evaluated and compared. The evaluation involved examining sentences of lemmas belonging to different word classes. Good sentences were logged for subsequent analysis with external data-mining software, WEKA. The observed behaviour was then used to adjust the parameters of the GDEX classifiers. We believe that the procedure of identifying features of good examples and their values, described in this paper, can be used for the development of GDEX for any language.
Identifying factors that improve the assessment of athletes' psychological functioning is imperative to make proper return-to-play decisions following concussion. Prior research indicates that an individual's affect is related to symptom reporting. The present study examines two novel methods of affect assessment in college athletes at baseline participating in a sports-concussion management program. A total of 256 athletes completed a neuropsychological baseline battery with measurements of psychological symptoms (BDI-Fast Screen, Post-Concussion Symptom Scale, and ImPact Total Symptom Score) and a measure of affective memory bias (the Affective Verbal Learning Test; AVLT). Examiners completed an observation-based rating of affect. Multivariate analysis of variance and χ2 analyses were conducted to examine the effect of affect on symptom reports. Examiners' Affect Ratings were predictive of broad symptom reporting, while the performance based index of affect (Affective Verbal Learning Test, AVLT) was more predictive of depressive symptoms. These findings suggest that performance on the AVLT may be a useful indicator of self-reported depression in a collegiate athlete sample. Additionally, these results demonstrate that examiners' behavioral assessments of affect are important in the assessment of psychological functioning in athletes. Continued work should focus on developing objective measures that are sensitive and valid for the evaluation of outcomes from concussion.
While most dialectological research so far focuses on phonetic and lexical phenomena, we use recent fieldwork in the domain of dialect syntax to guide the development of multidialectal natural language processing tools. In particular, we develop a set of rules that transform Standard German sentence structures into syntactically valid Swiss German sentence structures. These rules are sensitive to the dialect area, so that the dialects of more than 300 towns are covered. We evaluate the transformation rules on a Standard German treebank and obtain accuracy figures of 85% and above for most rules. We analyze the most frequent errors and discuss the benefit of these transformations for various natural language processing tasks. 1
Based on two syntactic dependency treebanks built with two different styles of Chinese, a statistical study is conducted regarding word-frequency and distributions. We extracted three grammatical words as the research objects and analyzed their network features, including all degree, out-degree, in-degree, all closeness, in-closeness, out-closeness and betweenness. Then these three nodes were removed from the networks. We recorded and compared the network features of the two original networks and the three networks from which one node is respectively removed, including the number of vertices, average degree, average path length, diameter, the number of isolated vertices, domain and density. The results show that all three function words are central nodes of the Chinese syntactic networks but have different status. Their influence to the overall structure is also quite different. The research not only provides a new method for the study about Chinese grammatical words but also provides a new way of thinking the node characteristics in the complex network.
"External sandhi is a linguistic phenomenon which refers to a set of sound changes that occur at word boundaries. These changes are similar to phonological processes such as assimilation and fusion when they apply at the level of prosody, such as in connected speech. External sandhi formation can be orthographically reflected in some languages. External sandhi formation in such languages, causes the occurrence of forms which are morphologically unanalyzable, thus posing a problem for all kind of NLP applications. In this paper, we discuss the implications that this phenomenon has for the syntactic annotation of sentences in Telugu, an Indian language with agglutinative morphology. We describe in detail, how external sandhi formation in Telugu, if not handled prior to dependency annotation, leads either to loss or misrepresentation of syntactic information in the treebank. This phenomenon, we argue, necessitates the introduction of a sandhi splitting stage in the generic annotation pipeline currently being followed for the treebanking of Indian languages. We identify one type of external sandhi widely occurring in the previous version of the Telugu treebank (version 0:2) and manually split all its instances leading to the development of a new version 0:5. We also conduct an experiment with a statistical parser to empirically verify the usefulness of the changes made to the treebank. Comparing the parsing accuracies obtained on versions 0:2 and 0:5 of the treebank, we observe that splitting even just one type of external sandhi leads to an increase in the overall parsing accuracies."
The possibility to analyse vast amounts of linguistic data has brought about changes both in methodology as well as in the ways we perceive certain language phenomena. A key insight gained by computational methods in language analysis is undoubtedly the importance of lexical co-occurrence and usage patterns for the description of lexical meaning. Corpus analysis and new methods in the analysis of pragmatic components of meaning have also yielded significant results in areas such as the treatment of semantic prosody. The present paper does not focus on what is traditionally subsumed under connotation or the speaker’s attitude (e.g., swear words, pejorative and offensive language, praise, excuses, requests, demands, etc.), but on ways in which the pragmatic (functional) meaning that arises from various contextual features can become an integral part of lexicographic descriptions. This is important for the treatment of all of those lexical items whose meanings reside in their function rather than in their bare lexical-semantic meaning, as this is particularly the case with phraseology and idiomatics. From another perspective, pragmatics turns out to be an effective means of sense discrimination in works of lexical and lexicographic relevance, as will be shown in the continuation.
In this paper we present a user-centered approach for defining the dependency syntactic specification for a treebank. We show that by collecting information on syntactic interpretations from the future users of the treebank, we can model so far dependency-syntactically undefined syntactic structures in a way that corresponds to the users ’ intuition. By consulting the users at the grammar definition phase we aim at better usability of the treebank in the future. We focus on two complex syntactic phenomena: elliptical comparative clauses and participial NPs or NPs with a verbderived noun as their head. We show how the phenomena can be interpreted in several ways and ask for the users ’ intuitive way of modeling them. The results aid in constructing the syntactic specification for the treebank. 1
Data integration systems attempt to provide users with seamless and flexible access to information from multiple autonomous, distributed and heterogeneous data sources through a unified query interface. Besides data are continuously growing, maintained by different organizations and managed autonomously, querying data from heterogeneous data sources faces new challenges. As data integration has been automated, the ambiguity in concept interpretation also known as semantic heterogeneity has become one of the main obstacles to this process. Introduction of the Semantic Web Vision Ontologies WordNet ontology [3] is a large lexical database that is used in many schema matching algorithms to match schemas based on the semantics of attributes. In this paper ontology based semantic query reformulation technique is followed to improve the recall of the query. The reformulated query is optimized by removing disjunctive clauses in the query to reduce the computational cost of the semantic query execution. Experimental results show that the proposed optimization technique improves recall with minimal execution time.
In this paper, we argue that there are two seemingly incompatible perceptions of discourse structure: a semantics-centered view and a syntax-centered view. In the semantics-based view, discourse structure is viewed as a structure that identifies the most important portions of the text and describes how they combine semantically. In the syntax-based view, discourse structure is viewed as an extension of syntax to the discourse level, which essentially links the syntactic trees for the individual sentences into one big tree structure. We will argue that these differences in perception may explain some of the central disagreements in the literature about the nature of discourse structure, in particular whether discourse structure is best viewed as a tree or a general graph. However, the two views are not as incompatible as they may seem at first sight, since the semantic discourse structure can be reinterpreted as a functor-argument structure that is derived from the syntactic tree structure. We describe the ramifications of the two views for the analysis of discourse markers, which are the focus of the discourse annotation in the Penn Discourse Treebank, and show how the syntax-based view can maintain a tree structure even for examples that seem to exhibit non-tree like properties in a semantics-based view.
While the 2007-2010 financial crisis has hit a variety of countries asymmetrically, the case of Spain is particularly illustrative: this country experienced a pronounced housing bubble partly funded via spectacular developments in its securitization markets leading to looser credit standards and subsequent financial stability problems. We analyze the sequential deterioration of credit in this country considering rating changes in individual securitized deals and on balance sheet bank conditions. Using a sample of 20, 286 observations on securities and rating changes from 2000Q1 to 2010Q1 we build a model in which loan growth, on balancesheet credit quality and rating changes are estimated simultaneously. Our results suggest that loan growth significantly affects on balance-sheet loan performance with a lag of at least two years. Additionally, loan performance is found to lead rating changes with a lag of four quarters. Importantly, bank characteristics (in particular, observed solvency, cash flow generation and cost efficiency) also affect ratings considerably. Additionally, these other bank characteristics seem to have a higher weight in the rating changes of securities issued by savings banks as compared to those issued by commercial banks. JEL Classification: G21, G12
It is hypothesized that ratings of emotional stimuli are affected by a constant threat of traumatic events. Ratings of valence and arousal on the International Affective Picture System from young adults in the United States were compared to those of young Israeli adults. Israelis rated the pictures as less negative and less positive than did participants from the United States. Israeli women gave higher arousal ratings compared to the American women. These differences may be due to compulsory military service in Israel, during which exposure to traumatic events is more likely to occur, and to the timing of the study which followed a year of frequent suicide bomb attacks. The authors suggest that these findings may reflect mild symptoms of stress disorders.
In the era of globalisation and simultaneous grouping, multilingualism has become a norm. In their job or studies, most of educated Estonians have to mediate information from one or more foreign languages. At the same time, several difficulties arise: (1) these people are not familiar with theoretical issues of translation; (2) they may be faced with a lack of suitable special terms in the target language; (3) for marking the same concept, several parallel scientific paradigms and groups characteristic of the era may use different 114 signifiers, and the other way around: the same lexical units may mark different notions. The article mediates empirical findings of editing (and retranslating) CEFR (2001) and PISA 2009 (2008) terminology; some more general conclusions may address practitioners of every-day translating, some others point to problems to be solved on the higher level of society. Keywords multilingualism, LSP, English, Estonian, translation
Abstract Prior research on relative clauses (RCs) in Mandarin Chinese has led to conflicting results regarding ease of processing subject-extracted RCs (SRCs) versus object-extracted RCs (ORCs) and has often used animacy configurations that are rare in corpora. Building on animacy patterns observed in a corpus, we used self-paced reading to explore how animacy influences real-time processing of Chinese RCs. Experiment 1 tested SRCs, and found marginal facilitation effects with animate heads (subjects) and inanimate objects. Experiment 2 tested ORCs and found significant facilitation effects with inanimate head (objects). Experiment 3 showed that when the subject is animate and the object inanimate, ORCs are as easy to process as SRCs, but when the subject is inanimate and the object is animate, SRCs are processed faster. Thus, the animacy of the head and the embedded noun must be taken into account when evaluating processing ease. Keywords: AnimacyRelative clause (RC)Mandarin ChineseProcessing Acknowledgments We would like to thank audiences at the 14th Annual Conference on Architectures and Mechanisms for Language Processing (AMLaP), the 2008 Western Conference on Linguistics (WECOL) and the 83rd Annual Meeting of the Linguistics Society of America (LSA), where earlier versions of some of this research were presented. Preliminary analyses of some of the data reported here appeared in Wu, Kaiser, and Andersen (2010). The stimuli used in this research are a revised version of the stimuli used in Wu (2009). We thank Yanan Sheng for assistance in the stimulus revision and running of participants, Xiaomei Qiao, and Tangfeng Yang for assistance in carrying out norming studies, Mei Li for providing facilities in running Experiment 3 at Tongji University, and Rudolf Troike for help with finalising the translations of our Chinese stimuli. This research was partially supported by a project sponsored by the Scientific Research Foundation for Returned Overseas Chinese Scholars, State Education Ministry, and by a grant from the Shanghai Municipal Philosophy and Social Sciences Foundation (2010BYY003) to the first author. Notes 1As a reviewer pointed out, the relation between head animacy and RC-type is clear with object RCs (which tend to occur with inanimate heads), but less so with subject RCs. Indeed, in Mak et al.'s (2002, pp. 54–55) German corpus, the 144 subject-extracted RCs have inanimate heads almost as frequently as animate heads: 57% animate heads and 43% inanimate heads. Also, in Roland et al.'s (2007, p. 357) analysis of the English-language Brown corpus, 47% of 100 randomly-selected subject-extracted RCs have inanimate heads. However, existing corpus data from Chinese suggest that subject RCs' head animacy patterns (at least in Chinese) may vary depending on the grammatical role of the RC's head noun. For Chinese, Pu (2007, p. 45) and Wu (2009) found that (1) when SRCs modify sentential subjects, animate heads significantly outnumber inanimate heads, but (2) when SRCs modify sentential objects, there is no particular bias toward animate or inanimate heads. 2The term "experiencer" refers to a change of psychological state on a human participant caused by someone or something in the context of certain intransitive verbs (e.g., win, die); experience-theme verbs (e.g., love, discover, like); or causer-experiencer verbs (e.g., please, amuse, amaze, and annoy). 3Lin and Garnsey (Citation2010) manipulated animacy in their stimuli, but they also topicalised their RCs to a sentence-initial position and used null head nouns. Headless RCs and topicalisation in Mandarin normally occur only when supportive discourse contexts are given, but their stimuli were presented in isolation. Thus their stimuli had a marked structure, which may have complicated their results. 4One reviewer pointed out that the percentage of RCs where both nouns have the same animacy is 28% in Mak et al.'s (2002) Dutch corpus and 40% in their German corpus. However, viewed from another perspective, this means that the percentage of RCs with contrastive animacy configuration is 72% in Dutch and 60% in German, a pattern similar to Wu's (2009) corpus analysis. Furthermore, at least in Wu's (2009) analyses of Chinese Treebank Corpus, RCs with matched animacy (double-animates or double-inanimates) occurred significantly less frequently than RCs with nonmatched animacy (p'<.05). 5In Experiment 1, the log frequencies for the different verbs and for the embedded nouns were matched. The mean log frequencies for the verbs from the SUBTLEX-CH are as follows: 3.14 for Oi-Sa and Oa-Sa, 3.32 for Oa-Si and Oi-Si. The frequencies do not differ significantly, F(3, 76) = 0.1069, p=.9558. The mean log frequencies for the verbs from the 2008 frequency dictionary are as follows: 9.46 for Oi-Sa and Oa-Sa, 9.04 for Oa-Si and Oi-Si. These frequencies also do not differ significantly, F(3, 78) = 0.6015, p=.616. The mean log frequencies for the embedded nouns from the SUBTLEX-CH are as follows: 3.07 for Oi-Sa and Oi-Si, 3.28 for Oa-Sa and Oa-Si. The frequencies do not differ significantly, F(3, 78) = 0.1726, p=.9146. The log frequencies for the embedded nouns from the 2008 frequency dictionary are as follows: 9.38 for Oi-Sa and Oi-Si, 9.36 for Oa-Sa and Oa-Si. The frequencies also do not differ significantly, F(3, 82) = 0.0087, p=.9989. Log frequencies for the head nouns were matched for the 2008 dictionary (means: 8.5 for Oi-Sa and Oa-Sa, 9.05 for Oa-Si and Oi-Si). According to this corpus, the frequencies of the different head nouns do not differ significantly, F(3, 72) = 1.7263, p=.1692. However, according to the SUBTLEX-CH corpus, the log frequencies for the head nouns are not matched [means: 4.32 for Oi-Sa and Oa-Sa, 2.88 for Oa-Si and Oi-Si; F(3, 74) = 4.8757, p=.0038]. As said, the frequency check reported above are based on an incomplete list of words that have their frequencies listed in either resource. 6At the sentence-initial RC-verb position (pos 1, e.g., raokai "bypass"), there was a marginal main effect of Head Animacy (t=1.84, p=.0663), and a marginal interaction between Head Animacy and Embedded-noun Animacy (t=−1.8, p=.073). However, given that this is the first word region, these weak effects are probably due to lexical differences. 7In Experiment 3, the log frequencies were matched for the verbs, but not for the embedded nouns and for the head nouns. The mean log frequencies of the verbs from SUBTLEX-CH are as follows: 3.83 for SRCs with animate heads and for ORCs with inanimate heads, 3.41 for SRCs with inanimate heads and for ORCs with animate heads. The frequencies do not differ significantly, F(3, 82) = 0.2254, p=.8785. The mean log frequencies of the verb from the 2008 frequency dictionary are as follows: 9.26 for SRCs with inanimate heads and for ORCs with animate heads, 9.19 for SRCs with inanimate heads and for ORCs with animate heads. These frequencies also do not differ significantly, F(3, 76) = 0.1274, p=.9436. The log frequencies for the embedded nouns from the SUBTLEX-CH are as follows: 2.86 for SRCs with animate heads and for ORCs with animate heads, 4.03 for SRCs with inanimate heads and for ORCs with inanimate heads. The differences in frequencies are marginally significant, F(3, 74) = 2.581, p=.062. The log mean frequencies of the embedded nouns from the 2008 frequency dictionary are as follows: 9.52 for SRCs with animate heads and for ORCs with animate heads, 9.11 for SRCs with inanimate heads and for ORCs with inanimate heads). These frequencies differ significantly, F(3, 76) = 2.833, p=.044. Reversely for the head nouns, their log frequencies for SRCs with inanimate heads and for ORCs with inanimate heads are more frequent than the log frequencies for SRCs with animate heads and for ORCs with animate heads. However, because we used a Latin-square design such that the different nouns rotated through the different conditions, we do not think this affects our results.