Papers reviewed and determined not to be word norm studies. Use the flag icon to report errors or suggest re-inclusion.
16504 papers
The Noun Phrase (NP) is the dominant construct in natural language text. While base NPs (BNP) and maximal length NPs (MNP) are relatively easy to identified and extracted, the internal structure of NPs is rather a challenge in natural language processing. Penn Treebank leaves the BNPs flat as implicit right branching. Vadas and Curran added BNP internal structure to the Penn Treebank. But the results of the BNP structure are very often incorrect when it is considered within a longer complex NP (CNP). Structural ambiguity prevails in most CNPs and multilingual comparison may help improve disambiguation. We introduce a new NP annotation scheme, which is applicable to multilingual parallel corpora and discriminate genuine flat branching and right branching. Flat branching is preferred instead of binary branching wherever appropriate so as to achieve inter-lingual consistency. As a pilot task to build a gold standard corpus for structural and semantic analysis of CNPs, 381 document titles are extracted from the UN resolutions as typical examples of CNPs. Document titles in Chinese, English and Russian are manually annotated in XML format with the hope to help acquire rules for parsers or machine translators targeted at CNPs. The problems encountered are reported.
The objective of textology (text linguistics) is to analyse and describe interpersonal communication in all its aspects, as it happens now and did in the past, in various discourse communities. The basic categories of so defined textology are: text, utterance and discourse, regarded as different perspectives on the phenomenon of interpersonal communication. This phenomenon manifests itself in established works, with their specific semantic and syntactic structure, but also in genre affinity, in interactive events fixed in a widely understood situational context, and in social, cultural and linguistic norms which regulate communication activities in individual human communities and make interpersonal communication possible. These three categories can be said to reify in definite communication phenomena: in written and recorded text, and in spoken utterances, i.e. in actualized discourses.
Learning vocabulary and understanding texts present difficulty for language learners due to, among other things, the high degree of lexical ambiguity. By developing an intelligent tutoring system, this dissertation examines whether automatically providing enriched sense-specific information is effective for vocabulary learning and reading comprehension of second language learners. The system developed in this study contributes to an extended understanding of how NLP techniques can be applied more effectively in an educational environment. The system allows learners to upload texts and click on any content word in order to obtain sense-appropriate lexical information for unfamiliar or unknown words during reading. The system consists of three components: (1) the system manager controls the interaction among each learner, the NLP server, and the lexical database; (2) the NLP server converts a raw input text to a linguistically-analyzed text; (3) the lexical database is used to provide a sense-appropriate definition and example sentences of a word to the learner. To obtain the sense-appropriate information, the system first performs word sense disambiguation (WSD) on the input text. Pointing to appropriate examples tuned for language learners, however, is complicated by the fact that the database of examples is from one repository (COBUILD), while automatic WSD systems generally rely on senses from another (WordNet). The lexical database, then, is indexed by WordNet senses, each of which points to an appropriate corresponding COBUILD sense. The fact that every sense inventory has its own standards of sense distinction poses a serious problem in integrating these inventories into one. To redirect an input WordNet sense to a corresponding COBUILD sense, thus, a word sense alignment algorithm was developed, following a heuristic of favoring flatter alignment structures. With this system, an empirical study was conducted with 60 intermediate learners of English as a second language to examine whether this system can lead learners to improve their vocabulary acquisition and reading comprehension. The findings show that learners demonstrated higher performance when receiving sense-specific information. Furthermore, the qualitative examination of the effect of automatic system errors show that, although learners showed learning regardless of the appropriateness of lexical information, they still showed relatively greater learning when given appropriate lexical information.
According to the questionnaire held in England, in 2000s, 333 from 547 of the informers admit the usage of split infinitive constructions in casual speech only, while 108 — in standard literary language and 106 consider these constructions to be non-standard. Such discrepancy in appraisal from the author’s point of view evidently proves the fact, that linguistic norm of this phenomenon is in the process of taking up its position. The debate history on the subject has deep roots and goes back to the Old English. Today this process is still not completed. In the author’s opinion the question is revealed and actual in modern linguistics. The work is devoted to a certain research of the problem, where the author applies to the sources of appearance and development of disputes, analyzies deferent positions and puts forward her own view point.
This paper discusses hybridization in contemporary Russian language. In particular, the work focuses on Dina Rubina�s novel Here comes the Messiah! (?????????????!.....).Her prose is characterized by a wide range of linguistic and communicative means; a key role is played by linguistic hybridization which alienates the text from any contexts, notions of time, as well as the ordinary and accepted linguistic norms. In her work, Rubina - a migrant writer living in Israel - describes a carnival atmosphere (based on Bachtin�s idea of carnival) as one of the main leitmotivs of the Russian community in Jerusalem. Words - which are are the primary means of expression in a world considered as a stage - often undergo hybridization processes. Therefore, new meanings, neologisms, original lexemes decorate Rubina�s multicultural text.
In the article, the currently existing codified orthographic norms are presented with reference to four selected problem topics and paralleled with certain other norms significantly influencing the synchronous usage of written language. The expression “norm” or “norms” is to be understood in its broadest sense: at one end of the continuum as a formalized set of rules and regulations which, if ignored, may result in sanctions, and, at the other end, as a set of principles and guidelines organized according to unified rules and affecting, or regulating, any one sphere of human activity and behaviour. The discussion includes other possible reasons for the discrepancies between orthographically prescribed use and use deviating from it. Despite the weight of these reasons, the non-linguistic norms presented in this article appear to be an influential factor that can by no means be overlooked by contemporary normative linguistics.
In this paper we introduce a new approach to transition-based depen- dency parsing. We propose that the parser construct an undirected graph during the parsing process, instead of a standard directed dependency structure. A pos- teriori, the output undirected structure is converted into a dependency tree. This alleviates error propagation, a characteristic problem of these systems. We apply this approach to obtain undirected variants of the Planar and 2-Planar parsers and of Covington's non-projective parser. We perform experiments on several treebanks from the CoNLL-X shared task, showing that these variants outperform the original directed algorithms in most of the cases.
The goal of this paper is to expose the character of ‘Grammar’ in curriculum standards of Korean Language Education revision in 2011. To achieve this goal, we examined the contents of grammar education in this curriculum, comparing with previous curricula and viewing different points of view on grammar education in the Korean Language Education. The grammar curriculum revision in 2011 tried to raise the value of grammar education through connecting the goal of grammar education to language skills. For example, the contents items chosen from the field of ‘Linguistic Norms’ and ‘Vocabulary’ was increased comparing with previous curricula. But in terms of contents, this new curriculum is unsatisfactory, because it selected its contents from linguistic content system as ever and it was unsuccessful in pursuing some important value of grammar education. Therefore, it is necessary to try constructing contents system of grammar education with the items based on the intrinsic value of language.
This article presents an interpretive study of âSiapa Menyuruh?â, a poem by Indonesiaâs contemporary poet Mustofa Bisri. The study is carried out within the framework of Sperber and Wilsonâs relevance theory (RT), which is based on the principles that human cognition tends to the maximization of relevance (I), and that every act of communication presumes its own optimal relevance (II). This suggests that in the relevance-theoretic perspective, Bisri wrote the poem âSiapa Menyuruhâ not because he wanted to violate certain linguistic norms or any communicative maxims, but because this was the most relevant utterance he could produce. He intentionally raised the effort to process his poem because he promised greater cognitive effects to the reader. The reader who is willing to process his utterances in the poem further is granted not with one, single strong implicature, but with a number of weak implicatures.
This article deals with the use or relative que with a circumstantial complement of time as an antecedent, and introducing clauses in which the relative, not preceded by a preposition, performs also in the subordinate clause the grammatical function of time adjunct. The adjunct that plays the role of an antecedent can be formed by a prepositional phrase, an adverb or a subordinate clause of time. In these cases, the relative has an adverbial function within the clause that introduces. These uses seem to continue that of undeclinable QUOD in Late Latin, which was used following a noun expressing time, to introduce a clause that indicates simultaneity. Although this type of clauses introduced by que is documented since medieval times, only when the antecedent is a time adverb or certain prepositional phrases its use is accepted in modern normative Spanish, while in other cases, specially if the antecedent is a time clause, there is a trend to reject it in the written linguistic norm, though it is still documented in colloquial speech.
The lexical database of the humanistic and baroque Czech MADLA covers vocabulary from 1500–1780. It contains about 750 000 hand-excerpted documents stemming from dictionaries, herbaria, chronicles and other literary documents from this historical period. Using this database, it is possible to monitor the development of meaning, word formation, paradigm and other grammatical categories, and it serves as a basis for further research on the vocabulary from the 16th–18th centuries. The text describes a representative sample of change in the meaning of the substantive kredenc („cupboard“), the meaning of which has shifted from the original „tasting“ to the contemporary „dresser“. Emphasis is placed on the proof and comparison of different meanings of this word.
Here we describe work on learning the subcategories of verbs in a morphologically rich language using only minimal linguistic resources. Our goal is to learn verb subcategorizations for Quechua, an under-resourced morphologically rich language, from an unannotated corpus. We compare results from applying this approach to an unannotated Arabic corpus with those achieved by processing the same text in treebank form. The original plan was to use only a morphological analyzer and an unannotated corpus, but experiments suggest that this approach by itself will not be effective for learning the combinatorial potential of Arabic verbs in general. The lower bound on resources for acquiring this information is somewhat higher, apparently requiring a a part-of-speech tagger and chunker for most languages, and a morphological disambiguater for Arabic.
The neological issue in the OSLL is a component of the wider framework of the research plan Creation of a Lexical Database of the Czech Language of the Beginning of the 21st Century (2005–2011, head – K. Oliva). The construction of the lexical collections (archives of lexical dynamics) is dealt with by the excerption section; the theoretical understanding and lexicographic treatment are assured by an independent working group of lexicographers. Thanks to the research plan being resolved, it was possible to ensure the continual complementation of the neological excerption, modify the method of the accumulation of the material in connection with the new tasks of the department and modernise the software equipment (in connection with that to make part of the neological material accessible to the wider public). These results are built on by the theoretical and practical activities of the neological working group, focusing on the treatment of new material (2002–2010).
The neological issue in the OSLL is a component of the wider framework of the research plan Creation of a Lexical Database of the Czech Language of the Beginning of the 21st Century (2005–2011, head – K. Oliva). The construction of the lexical collections (archives of lexical dynamics) is dealt with by the excerption section; the theoretical understanding and lexicographic treatment are assured by an independent working group of lexicographers. Thanks to the research plan being resolved, it was possible to ensure the continual complementation of the neological excerption, modify the method of the accumulation of the material in connection with the new tasks of the department and modernise the software equipment (in connection with that to make part of the neological material accessible to the wider public). These results are built on by the theoretical and practical activities of the neological working group, focusing on the treatment of new material (2002–2010).
Based on Kachru’s Three Circle Model on the spread of English in different parts of the world, I question how foreign university students migrating from Expanding Circle countries to Singapore deal with the “clash” of linguistic norms set by different circles. This research hence explores English speaking behavioural intentions of these foreign tertiary students in Singapore and how they account for their language behavioural plans. The collected data reveal that most students held a notion on the native English-Singlish dichotomy, seeing native English as standard and superior while regarding Singlish as improper and non-standard. Considering language behavioural intentions, most respondents claimed to adopt three main strategies: speech maintenance, adapting to the formality level of the communicative situation, and speech convergence. Looking into respondents’ accounts for these intended strategies, I argue that speakers orient their language use not only towards language perceptions but also the communicative situations they are in. However, the relative influences of these two factors on language use vary among different cases.
We investigate aspects of interoperability between a broad range of common annotation schemes for syntacto-semantic dependencies. With the practical goal of making the LinGO Redwoods Treebank accessible to broader usage, we contrast seven distinct annotation schemes of functor‐argument structure, both in terms of syntactic and semantic relations. Drawing examples from a multi-annotated gold standard, we show how abstractly similar information can take quite different forms across frameworks. We further seek to shed light on the representational ‘distance’ between pure bilexical dependencies, on the one hand, and full-blown logical-form propositional semantics, on the other hand. Furthermore, we propose a fully automated conversion procedure from (logical-form) meaning representation to bilexical semantic dependencies. †
This paper introduces the error corpus of Korean learner English and mal rules to detect the errors. Based on the corpus, we classified 42 error types. Our criteria for error classification are more general in order to enhance agreement rate and decrease errors. For generating mal rule, we testified two different grammars. One is Context Free Grammar (CFG) from Penn Treebank. The other is the typed feature structure grammars based on the Head-Driven Phrase Structure Grammar (HPSG), using Natural Language ToolKit (Bird et al. 2009). We advanced grammatical formalism from CFG to HPSG since CFG needs abundant phrasal markers causing over-generation, structural ambiguity and complexity.
This thesis introduces an experimental and quantitative approach to language through the study of the concept of soft constraints and its application to two phenomena of order in French: the position of the attributive adjective and the ordering of verbal complements occurring in postverbal position. Soft constraints are defined as affecting the acceptability rather than the grammaticality of the sentences. Our main hypothesis is that these constraints are properties of the language and thus must studied in syntax. These constraints raise a methodological issue: since they do not affect the grammaticality of the sentences, they cannot be investigated using the traditional tools of syntax (introspection and grammaticality judgment). It is therefore necessary to define tools for their description and analysis. The proposed methods are statistical analysis of corpus date, inspired by the work of Bresnan et al. (2007) and Bresnan & Ford (2010) and, to a lesser extent, psycholinguistic experiment. Regarding the position of the adjective, we test most of the constraints encountered in the literature and we propose a statistical analysis of the data extracted from the French Treebank corpus. We show the importance of the adjectival item and the nominal item with which it combines. Other constraints linked to the internal syntax of the adjectival phrase and the noun phrase also play a significant role in the choice of position. The work on the relative order of the verbal complements is conducted on a sample of sentences extracted from two newspapers corpora (French Treebank and Est-Républicain) and two corpora of spoken French (ESTER and C-ORAL-ROM). We show the significant influence of the constituent weight over the ordering: short before long order which is a feature of SVO languages like French, is observed in over 86% of cases. We also identify the important role of the verbal lemma associated with its semantic class (annotated with the dictionary of Dubois & Dubois-Charlier, 1997). Finally, building on the analysis of corpus data as well a two questionnaires eliciting acceptability judgments, it seems that nor animacy neither information structure (given/new, Prince, 1981) have a significant effect on the postverbal complement ordering.
The chapter presents a meta-search tool developed in order to deliver search results structured according to the specific interests of users. Meta-search means that for a specific query, several search mechanisms could be simultaneously applied. Using the clustering process, thematically homogenous groups are built up from the initial list provided by the standard search mechanisms. The results are more user oriented, as a result of the ontological approach of the clustering process. After the initial search made on multiple search engines, the results are pre-processed and transformed into vectors of words. These vectors are mapped into vectors of concepts, by calling an educational ontology and using the WordNet lexical database. The vectors of concepts are refined through concept space graphs and projection mechanisms, before applying the clustering procedure. Implementation details and early experimentation results are also provided.
This paper outlines a proposal for maritime English language teaching in public and private Nautical Schools and other maritime educational institutions and establishments in Italy, using a content and language integrated learning (CLIL) approach. The courses are addressed in particular to those students who would like to take up a marine career as officers, engineers or other crew members of the Merchant Navy, and thus require an adequate knowledge of seafaring terminology, but can also be interesting for those wishing to explore the origins and development of maritime language. In order to provide a more challenging environment and better opportunity for the learning of seafaring terms and expressions in English, students are supported by Mariterm, a lexical database, organized in semantic relations, available at the Institute for Computational Linguistics (ILC) of the National Research Council (CNR) in Pisa. A
In this paper we employ a most recent approach to Data Oriented Parsing (DOP), which has named Double-Dop, for Persian sentences. Like other DOP models, Double-Dop parser utilizes syntactic fragments of arbitrary size from a treebank to analyse new sentences, but it extracts a restricted yet representative subset of fragments. It uses only those which are encountered at least twice. The accuracy of Double-DOP is well within the range of state-of-the-art parsers currently used in other NLP-tasks, while offering the additional benefits of a simple generative probability model and an explicit representation of grammatical constructions. Heretofore there isn’t any standard parser for Persian language and this work try to employ Double-Dop Method for parsing Persian sentences.
Nowadays, the volume of information increases exponentially, forcing the corporations to keep their business information distributed under several heterogeneous sources such as relational databases, spread sheets, XML documents and Web pages, and stored under different structures and formats. Integrating heterogeneous sources is recently acknowledged as an important vision on semantic web research. The concept of heterogeneity arises at different levels: from the lexical level to the semantic or structural level. For discovering and consolidating the semantic relationships among the semantically related data present in different types of databases and files, this paper presents the enhancements obtained due to the use of available online large lexical databases, combined with lexical and structural similarity models and the available source metadata. Finally, we reveal the experimental results that demonstrate the applicability and usability of our approach.
Does international law's effectiveness require a clear distinction between law and non-law? This essay, which reviews Jean d'Aspremont's Formalism and the Sources of International Law, argues the answer is no. Ambiguity about the legal nature of international instruments has important benefits. Clarity in the law may encourage states to do the minimum necessary to comply, while some uncertainty about what the law requires may induce states to take extra efforts to ensure they are in compliance. Ambiguity in the law also promotes dynamic change, an important feature in rapidly developing areas of the law such as international environmental law and human rights. Most importantly, though, soft law — international instruments that have legal consequences but are not unambiguously 'law' — expands the range of instruments available to states when cooperating. Institutionalist theories of international law suggest that a larger menu of international instruments is valuable because it allows states to calibrate the level of their commitments more precisely, thereby expanding their ability to cooperate. Institutional theories, however, have heretofore not explained exactly how states communicate to each other the level of their commitment; that is, they have not explained how states mark an instrument as soft law and whether and how states distinguish between types of soft law commitments. A theory of law-identification based on linguistic norms, such as d'Aspremont proposes, offers a descriptive account of how states might signal levels of legal commitment beyond the dichotomy of 'binding' and 'non-binding' law. A communicative theory of international law — one based on the use of language in international instruments to signal relatively fine-grained variation in the level of commitment — thus would enrich our understanding of what soft law is, and when and how states use it.
The majority of fear conditioning studies in humans have focused on fear acquisition rather than fear extinction. For this reason only a few functional imaging studies on fear extinction are available. A large number of animal studies indicate the medial prefrontal cortex (mPFC) as neuronal substrate of extinction. We therefore determined mPFC contribution during extinction learning after a discriminative fear conditioning in 34 healthy human subjects by using functional near-infrared spectroscopy. During the extinction training, a previously conditioned neutral face (conditioned stimulus, CS+) no longer predicted an aversive scream (unconditioned stimulus, UCS). Considering differential valence and arousal ratings as well as skin conductance responses during the acquisition phase, we found a CS+ related increase in oxygenated haemoglobin concentration changes within the mPFC over the time course of extinction. Late CS+ trials further revealed higher activation than CS– trials in a cluster of probe set channels covering the mPFC. These results are in line with previous findings on extinction and further emphasize the mPFC as significant for associative learning processes. During extinction, the diminished fear association between a former CS+ and a UCS is inversely correlated with mPFC activity – a process presumably dysfunctional in anxiety disorders.
With english speaking population expanding rapidly due to increasing international communication, english is no longer spoken by native speakers alone. English is spoken in different regions around the world, developing into different variants reflecting local language and culture. When speakers in an international conference speak in non-native english, interpreters, unfamiliar to such varieties, may be challenged as unfamiliarity to specific variants is more likely to present difficulties in intelligibility and comprehensibility. Previous studies indicate that conference interpreters found unfamiliar accents challenging and difficult to interpret. This paper discusses stress factors for English-Korean interpreters presented by non-native english speakers, following the concept of ‘World Englishes,’ which refers to different varieties developed around the world reflecting cultural and linguistic norms of the region or society of speakers. This paper aims to identify previous researches that point to the challenges of interpreting performance in general, and world englishes, in particular, to develop a theoretical framework for analyzing specific difficulties experienced by English-Korean conference interpreters at the level of phonology, syntax, and processing effort
Domain adaptation is an important task in order for NLP systems to work well in real applications. There has been extensive research on this topic. In this paper, we address two issues that are related to domain adaptation. The first question is how much genre variation will affect NLP systems ’ performance. We investigate the effect of genre variation on the performance of three NLP tools, namely, word segmenter, POS tagger, and parser. We choose the Chinese Penn Treebank (CTB) as our corpus. The second question is how one can estimate NLP systems ’ performance when gold standard on the test data does not exist. To answer the question, we extend the parsing prediction model in (Ravi et al., 2008) to provide prediction for word segmentation and POS tagging as well. Our experiments show that the predicted scores are close to the real scores when tested on the CTB data.
This experimental study investigates the impact of affective attitudes on risk and return estimates of stocks. Participants rate well-known blue-chip firms on an affective scale and forecast risk and return of the firms' stock. We find that positive affective attitudes lead to a prediction of high return and low risk, while negative attitudes lead to a prediction of low return and high risk. This bias increases with participants' confidence in their ratings and decreases with financial literacy. Firm characteristics such as a firm's marketing expenditures and the strength of its brand have a positive impact on its affective rating.
The automatic recognition of the maximal-length noun phrase (MNP) helps to the shallow parsing. In this paper, automatic labeling of Chinese MNP is regarded as a sequential labeling task and Support Vector Machine model (SVM) is employed in the model. We propose a method which takes 2-phase hybrid approach which first identifies base chunk and then identifies MNP. Furthermore, the base chunk features can be exploited to improve performance of MNP recognition. In addition, both left-right and right-left sequential labeling were employed to identify Chinese MNP by bidirectional sequence labeling merging. The data set in the experiments is selected from Penn Chinese Treebank 5.0 Corpus, and split into train set, development set and test set according to the proportion of 4:4:1. Experimental result shows a high quality performance of 90.13% in F1-measure.
The Penn Discourse Treebank (PDTB) is a data set that promotes the advancement of dis- course analysis tools. Improving discourse analysis is useful for other areas of Natural Language Processing, but it still proves to be a challenging task. This research focuses on improving sense relation classification in the PDTB for implicit relations at all three levels. The features selected for classification are motivated by prior research and statistical testing in this research. The goal is not only to provide features that improve classification in PDTB, but also to select features which are broad enough to be effective beyond the scope of the PDTB. Moreover, these features are derived from a variety of categories such as Semantics, Syntax and Entity in order to ensure stronger results. Using these features, of which many are new for PDTB sense classification, and Naive Bayes, Maximum Entropy and SVM, this research shows improvement in a number of senses that prior research has had difficulty improving, mainly Comparison and Contingency.
If the sentences or phrases like Fast Food-Konzept in Chile, Leonardo goes Gastronomie, Der Claim „Inspiration for modern living“, WM-Countdown läuft; Deutschland ist die fünftgrößte Incoming-Destination weltweit und … are considered, the first impression is that this is a typical mixture of both English and German. To all intents and purposes, they support the main part of a definition of pidgin language. Communication plays a central role in tourism. As emphasized by numerous authors, it is a personification of tourism; conversely tourism could be said to totally encompass a system of communication. As it is a crucial component of the industry, tourist discourse not only serves as a medium for buying/selling tourist products, it also assumes the role of the product itself within the complexity of various economic, technological and political processes in tourism. If analyzed within linguistic norms, discussion of Anglo-American influence on the German tourist discourse should focus on the problem of erosion of the national language, due to the impact of these Ango-Americanisms. The German tourist discourse aims at a customer-friendly approach in order to attract potential buyers or guests, all of which results in a slightly pidginized version of the German tourist discourse.
The purpose of this paper is to investigate the distribution of the pronunciations (mainly /ju, u,?/) of orthographic 〈u〉 in British English. /ju, u,?/ have incurred many complicated problems regarding the status of /j/, the underling form of these vowels, etc. Chomsky and Halle (1968) argue that the pronunciations of 〈u〉 are derived from Middle English lax /?/; it becomes /ju/ when it occurs in an open syllable; it remains /?/ when it is preceded by one of the nonnasal labials /p, b, f, v/ and followed by /l,?,?/; elsewhere, that is, in a closed syllable, it becomes /?/. It is expected that the pronunciations of 〈u〉 and their contexts are reflected in the present-day pronunciations of 〈u〉. To check this I analyzed the monomorphemic words in the CELEX lexical database; [?] appears mostly in a closed syllable as expected, but it is hard to say that [ju] and [u] occur mostly in the expected contexts. It is also found that /j?/ occurs only in a word-medial unstressed open syllable and that the favorite onset of a /u/-syllable is /l, r/, the reason of which is to be studied in the future.
This paper is stimulated by the ideas of Karel Hausenblas, the work of Olga Müllerová and the research on the syntax of Czech dialects (J. Balhar, J. Chloupek, M. Šipková and others). It presents a catalogue of phenomena and means of expression which mark the syntax of spoken Czech and devotes attention above all to: a) special syntactic constructions b) the varying formation of transitions between syntactic units in spoken expression (sharply structured transitions) and in written expression (softer, less apparent transitions, couched or “stuck” with numerous redundant means with non-definite semantics); that is, differences in the degree and type of cohesion, connection, or glutination between written and spoken expression c) differences between condensed, constricted written syntax and the relaxed structure of syntactic units in spoken Czech (with the prevalence of parataxis and juxtaposition). The paper views the syntactic differences between written and spoken expression as stylistic differences. It is based on data from various corpora of spoken Czech (including the Prague Dependency Treebank of Spoken Czech) and on the comparison of written and spoken narrative by the same speaker/author.
It is said that Vietnamese is a language with highly ambiguous words. However, there has been no published Word Sense Disambiguation (WSD hereafter) research on this language. This current research is the first attempt to study Vietnamese WSD. Especially, we would like to explore the effective features for training WSD classifiers and verify the applicability of the ‘pseudoword’ technique to both investigating effectiveness of features and training WSD classifiers. Three tasks have been conducted, using two corpora which were built manually based on Vietnamese Treebank and automatically by applying pseudowords technique. Experiment results showed that Bag-Of-Word feature performs well for all three categories of words (verbs, nouns, and adjectives). However, its combination with POS, Collocation or Syntactic features can not significantly improve the performance of WSD classifiers. Moreover, the experiment results confirmed that pseudoword is a suitable technique to explore the effectiveness of features in disambiguation of Vietnamese verbs and adjectives. Furthermore, we empirically evaluated the applicability of the pseudoword technique as an unsupervised learning method for real Vietnamese WSD.
This paper is stimulated by the ideas of Karel Hausenblas, the work of Olga Müllerová and the research on the syntax of Czech dialects (J. Balhar, J. Chloupek, M. Šipková and others). It presents a catalogue of phenomena and means of expression which mark the syntax of spoken Czech and devotes attention above all to: a) special syntactic constructions b) the varying formation of transitions between syntactic units in spoken expression (sharply structured transitions) and in written expression (softer, less apparent transitions, couched or “stuck” with numerous redundant means with non-definite semantics); that is, differences in the degree and type of cohesion, connection, or glutination between written and spoken expression c) differences between condensed, constricted written syntax and the relaxed structure of syntactic units in spoken Czech (with the prevalence of parataxis and juxtaposition). The paper views the syntactic differences between written and spoken expression as stylistic differences. It is based on data from various corpora of spoken Czech (including the Prague Dependency Treebank of Spoken Czech) and on the comparison of written and spoken narrative by the same speaker/author.
Linguistic Thought of the Spanish Renaissance and HumanismThe present article examines the linguistic thoughts of the period of the Humanism and Renaissance in Spain comparing with the case of Italy.In Spain, in the end of the 15th century, the philologist Antonio de Nebrija introduced the humanisitic ideas born in Italy, and applied them not only to reform the education of the Latin or the studies of letters in general but also to create the grammar of a vernacular castillian language.In Italy, already in the 14th century, there were great works of literature written by the three most brilliant authors; Dante, Boccaccio and Petrarca.Then during the next two centuries the debates surround the language got intense, and the intellectuals discussed the question of how to establish linguistic norms and codify the language, that is to say which dialect or speech of Italy -like the Tuscan-should be standard of the whole Italian peninsula.This called questione della lingua"question of the language" no longer was a problem exclusively in Italy, rather than the problem of the whole Europe.In Spain, due to its proximity to Italy, that current of thought was introduced and developed pronto.But while then Italy was divided into a number of warring city-states that have distinctive dialects for each, in Spain, contrastively, there was a strong centralized government due to the union of Castille and Aragon in the latter half of the 15th century.This difference of political situations made different atmosphere in the debates about language in each of the two peninsulas.I will describe what was the "questions of the language" and how developed this both in Italy and in Spain.
The focus of this article is on the creation of a collection of sentences manually annotated with respect to their sentence structure. We show that the concept of linear segments—linguistically motivated units, which may be easily detected automatically—serves as a good basis for the identification of clauses in Czech. The segment annotation captures such relationships as subordination, coordination, apposition and parenthesis; based on segmentation charts, individual clauses forming a complex sentence are identified. The annotation of a sentence structure enriches a dependency-based framework with explicit syntactic informa- tion on relations among complex units like clauses. We have gathered a collection of 3,444 sentences from the Prague Dependency Treebank, which were annotated with respect to their sentence structure (these sentences comprise 10,746 segments forming 6,341 clauses). The main purpose of the project is to gain a development data—promising results for Czech NLP tools (as a dependency parser or a machine translation system for related languages) that adopt an idea of clause segmentation have been already reported. The collection of sentences with annotated sentence structure provides the possibility of further improvement of such tools.
Statistické jazykové modely jsou důležitou součástí mnoha úspěšných aplikací, mezi něž patří například automatické rozpoznávání řeči a strojový překlad (příkladem je známá aplikace Google Translate). Tradiční techniky pro odhad těchto modelů jsou založeny na tzv. N-gramech. Navzdory známým nedostatkům těchto technik a obrovskému úsilí výzkumných skupin napříč mnoha oblastmi (rozpoznávání řeči, automatický překlad, neuroscience, umělá inteligence, zpracování přirozeného jazyka, komprese dat, psychologie atd.), N-gramy v podstatě zůstaly nejúspěšnější technikou. Cílem této práce je prezentace několika architektur jazykových modelůzaložených na neuronových sítích. Ačkoliv jsou tyto modely výpočetně náročnější než N-gramové modely, s technikami vyvinutými v této práci je možné jejich efektivní použití v reálných aplikacích. Dosažené snížení počtu chyb při rozpoznávání řeči oproti nejlepším N-gramovým modelům dosahuje 20%. Model založený na rekurentní neurovové síti dosahuje nejlepších publikovaných výsledků na velmi známé datové sadě (Penn Treebank).
The study of the Tip of the Tongue phenomenon (TOT) provides valuable clues and insights concerning the organisation of the mental lexicon (meaning, number of syllables, relation with other words, etc.). This paper describes a tool based on psycho-linguistic observations concerning the TOT phenomenon. We've built it to enable a speaker/writer to find the word he is looking for, word he may know, but which he is unable to access in time. We try to simulate the TOT phenomenon by creating a situation where the system knows the target word, yet is unable to access it. In order to find the target word we make use of the paradigmatic and syntagmatic associations stored in the linguistic databases. Our experiment allows the following conclusion: a tool like SVETLAN, capable to structure (automatically) a dictionary by domains can be used sucessfully to help the speaker/writer to find the word he is looking for, if it is combined with a database rich in terms of paradigmatic links like EuroWordNet.
The explosion of information in the World Wide Web is overwhelming readers with limitless information. Large internet articles or journals are often cumbersome to read as well as comprehend. More often than not, readers are immersed in a pool of information with limited time to assimilate all of the articles. It leads to information overload whereby readers are trying to deal with more information than they can process. Hence, there is an apparent need for an automatic text summarizer as to produce summaries quicker than humans. The text summarization research on mobile platform has been inspired by the new paradigm shift in accessing information ubiquitously at anytime and anywhere on Smartphones or smart devices. In this research, a semantic and syntactic based summarization is implemented in a text summarizer to solve the overload problem whilst providing a more coherent summary. Additionally, WordNet is used as the lexical database to semantically extract the text document which provides a more efficient and accurate algorithm than the existing summary system. The objective of the paper is to integrate WordNet into the proposed system called TextSumIt which condenses lengthy documents into shorter summarized text that gives a higher readability to Android mobile users. The experimental results are done using recall, precision and F-Score to evaluate on the summary output, in comparison with the existing automated summarizer. Human-generated summaries from Document Understanding Conference (DUC) are taken as the reference summaries for the evaluation. The evaluation of experimental results shows satisfactory results.
This paper focuses on the links between contemporary literature and the various positions choosenchosenby authors facing the problematics of translation. Beginning with the observation that translation studies should develop from a theoretical point of view in Japan--an emblematic country for translations--this paper shows that currently, translation in Japan has to be considered as a cultural exportation trend and not only as the importation trend that dominated the cultural scene during the 20th century. For example, data on published translations in France show that since 2007, Japanese is the second most frequently translated language after American-English--due to the popularity of mangas in France. In the literary field, new phenomenons can also be observed in Japan. In this paper, four case studies are presented. The most remarkable case concerns Murakami Haruki's strategy, in which he, being an important translator of the Great American Novel, crosses the boundaries between countries and languages in order to represent a new kind of nationless writer, i.e. a global writer appreciated all over the world. On the other hand, Mizumura Minae mixes English and Japanese in her I novel from left to right, making it untranslatable into English. This for her represents the resistance of a minor language, Japanese, to the domination of English. Tawada Yôko, for her part, writes in two languages, Japanese and German, and in doing so tries to deconstruct both cultural and linguistic norms, enhancing translation as an impossible tool. Finally, the American-born Hideo Levy's three-piece band features Japanese, English and Chinese members, interconnected by the belief in translation as an ideal vector of communication. All these new streams contribute to the reshifting of Japanese literature in the world and induce a necessary renewal of the critical approaches.
Despite of rapid progress in Southern Africa in the direction of multifunctionality of lexical databases through the advent of generic lexicographic software, a considerable number of lexicographic projects — especially in Khoe and Saan languages — still use or have recently used a word processor with the sole objective of compiling a printed dictionary. Hence the present paper expounds on the case of the Khoekhoegowab Dictionary Project, how in the early 1990s some off-the-shelf DOS-based database software was configured as part of a "home-grown" custom-made dictionary writing system. It is demonstrated in a non-technical way that the use of a structured database with fully-fledged retrieval facilities allows for the far-reaching elimination of human error in a dictionary, for the automatisation of processes like language reversal and sorting, and, finally, for the significantly enhanced usability of the data for purposes other than fixed media dictionary compilation. Compiling a dictionary without extensive query facilities as offered by tabular databases, is argued to be a lost opportunity, as it should be possible to utilise lexicographic data for more than just lexicography. By 2010 the data was accommodated in open source software to ensure its optimal survival in digital form for future use. Keywords: automatisation; compilation software; data retrieval; database configuration; database report; flat-file database; form; information generation; khoekhoe; khoesaan dictionaries; lexicography; lookup facilities; multifunctionality; query facilities; retrieval facilities; software; tones
Alexithymia is a personality trait characterised by difficulties in identifying and describing one’s emotions, constricted imaginal processing, and an externally oriented cognitive style. Alexithymia is associated with psychopathology and interpersonal problems. The aim of the current study was to evaluate the psychometric properties of the most frequently used measure of alexithymia, the self-report 20-item Toronto Alexithymia Scale (TAS-20). Specifically, the study aimed to (1) cross-validate the hypothesised three-factor structure of the TAS-20, (2) determine whether the measure indirectly assesses the constricted imaginal thinking component of alexithymia, despite its absence of imagination items, (3) examine the overlap between the TAS-20 and measures of psychopathology, and (4) determine whether the TAS-20 assesses actual, rather than merely perceived, emotional understanding. Participants were 194 (138 female) university students and community members who completed an online survey. Confirmatory factor analyses showed mixed support for the hypothesised three-factor model of the TAS-20; however, this model provided a better fit to the data than either a one- or two-factor model. Inverse relationships were found between the TAS-20 and measures of perspective-taking and fantasy (although this relationship was only marginally significant for fantasy). There were moderate to large positive associations between the TAS-20 and measures of depression, anxiety, stress, and negative affectivity. Inverse relationships were found between the TAS-20 and objective measures of emotional ability; however, these relationships were no longer significant after the effects of negative emotions and affectivity were partialled out. Higher TAS-20 scores were also associated with more moderate affective valence ratings of emotion-evoking stimuli, and thus lower self-reported arousal. Together, these findings suggest problems with the TAS-20’s construct validity. Theoretical and practical implications are discussed.
Individuals who effectively regulate or mildly increase their systolic blood pressure (SBP) in response to an orthostatic challenge exhibit healthier affective status, cognitive functioning, and better quality of life. Thus, increased SBP in response to an orthostatic challenge serves as a proxy for several underlying changes. This study examined the relationship between SBP regulation and self-esteem in children. Data were collected from 92 boys and girls, aged 8–11 years. Systolic, diastolic, and pulse measurements were obtained after 5 minutes of remaining supine and again after 1 minute of standing. Children also provided affective ratings on the Children's Depression Inventory. The Negative Self-Esteem subscale was examined for this study. A multiple regression analysis revealed that poorer orthostatic regulation was associated with higher levels of negative self-esteem among children aged 8–11 years. Thus, orthostatic BP regulation may serve as a biological marker for poor self-esteem in children. This may have further implications for children's emotional functioning as low self-esteem may serve as a risk factor for future negative affective states.
Early-latency theories of emotional processing state that at least coarse monitoring of the emotional valence (a pleasure-displeasure continuum) of facial expressions should be both rapid and highly automated (LeDoux, 1995; Russell, 1980). Research has largely substantiated early-latency differential processing of emotional versus non-emotional facial expressions; however, the effect of valence on early-latency processing of emotional facial expression remains unclear. In an effort to delineate the effects of valence on early-latency emotional facial expression processing, the current investigation compared ERP responses to positive (happy and surprise), neutral, and negative (afraid and sad) basic facial expression photographs as well as to positive (happy-surprise), neutral (afraid-surprise, happy-afraid, happy-sad, sad-surprise), and negative (sad-afraid) morph facial expression photographs during a valence-rating task. Morphing manipulations have been shown to decrease the familiarity of facial patterns and thus preclude any overlearned responses to specific facial codes. Accordingly, it was proposed that morph stimuli would disrupt more detailed emotional identification to reveal a valence response independent of a specific identifiable emotion (Balconi & Lucchiari, 2005; Schweinberger, Burton & Kelly, 1999). ERP results revealed early-latency differentiation between positive, neutral, and negative morph facial expressions approximately 108 milliseconds post-stimulus (P1) within the right electrode cluster; negative morph facial expressions continued to elicit significantly smaller ERP amplitudes than other valence categories approximately 164 milliseconds post-stimulus (N170). Consistent with previous imaging research on emotional facial expression processing, source localization revealed substantial dipole activation within regions of the mesolimbic dopamine system. Thus, these findings confirm rapid valence processing of facial expressions and suggest that negative valence processing may continue to modulate subsequent structural facial processing.
New Irish speakers in Belfast play a crucial, complex part in the revitalization and change of both the city and Irish within Northern Ireland. This paper examines the role of new Irish speakers in transforming Belfast, whose emergence from a post-conflict period involves a reassessment of communal cultural expressions. Markers of ethno-national identity are bitterly contentious locally, and yet increasingly celebrated, in line with international trends, as high status cultural forms and potentially profitable tourist attractions. Irish in Belfast currently occupies an ambiguous position: divisive enough for a sign reading ‘Happy Christmas’ in Irish to be experienced as an insult by some city councillors, yet a secure enough part of the establishment for a neighbourhood to be officially rebranded as the Gaeltacht Quarter. <br/>When, how and where new Irish speakers use the language in Belfast has implications for the relationship of Irishness to the Northern Irish state and for the place of Belfast within regional frameworks across the UK, Ireland and Europe. Adult learners and young people exiting Irish medium education have an impact on life in Belfast beyond its small population of Irish speakers. Urbanisation fuelled by new speakers, which shifts the balance of Irish language resources and speakers away from traditional rural Gaeltacht areas and towards cities, also has implications for the language itself. Recent increase in new Irish speakers in Belfast is due to expansion in the Irish-medium sector as well as to adult learners, whose decisions contribute to the school expansion. <br/>Urbanisation, multilingualism and intergenerational shift combine in Belfast to produce new linguistic norms. Moreover, in a minority language community where hierarchies of ‘authenticity’ are weighted towards the rural and the native speaker, where the rural and the native have traditionally been conflated, and where indigeneity is a central concept to contested nationalisms, the emergence of a self-confident, youthful Irish speaking community in Northern Ireland’s biggest city involves a recalibration of the qualities signifying ‘gaelicness’. As students, professionals, hobbyists and activists, new Irish speakers in Belfast occupy a vital position at the crux of changing ideas about place, language and identity.<br/>
This dissertation explores the nature and extent of retroflex consonant harmony in South Asia. Using statistics calculated over lexical databases from a broad sample of languages, the study demonstrates that retroflex consonant harmony is an areal trait affecting most languages in the northern half of the South Asian subcontinent, including languages from at least three of the four major families in the region: Dravidian, Indo-Aryan and Munda (but not Tibeto-Burman). Dravidian and Indo-Aryan languages in the southern half of the subcontinent do not exhibit retroflex consonant harmony. In South Asia, retroflex consonant harmony is manifested primarily as a static co-occurrence restriction on coronal consonants in roots/words. Historical-comparative evidence reveals that this pattern is the result of retroflex assimilation that is non-local, regressive and conditioned by the similarity of interacting segments. These typological properties stand in contrast to those of other retroflex assimilation patterns, which are local, primarily progressive, and not conditioned by similarity. This is argued to support the hypothesis that local feature spreading and long-distance feature agreement constitute two independent mechanisms of assimilation, each with its own set of typological properties, and that retroflex consonant harmony is the product of agreement, not spreading. Building on this hypothesis, the study offers a formal account of retroflex consonant harmony within the Agreement by Correspondence (ABC) model of Rose & Walker (2004) and Hansson (2001; 2010). Two Indo-Aryan languages, Kalasha and Indus Kohistani, figure prominently throughout the dissertation. These languages exhibit similarity effects that have not been clearly observed in other retroflex consonant harmony systems; retroflexion is contrastive in both non-sibilant (i.e., plosive) and sibilant obstruents (i.e., affricates and fricatives), but harmony applies only within each manner class, not between them. At the same time, harmony is not sensitive to laryngeal features. Theoretical implications of these and other similarity effects are discussed.
The task of automatic machine translation (MT) is the focus of a huge variety of active research efforts, both because of the intrinsic utility of this difficult task, and the theoretical and linguistic insights that arise from modeling relationships between natural languages. However, MT systems that leverage syntactic information are only recently becoming practical, and in a typical system of this sort, syntactic information is generated by monolingual parsers; the task of explicitly modeling syntactic relationships between target and source languages is yet to be fully explored. This thesis investigates the problem of finding syntactic parse trees of target and/or source sentences that are more appropriate for use in a syntactic MT system. Two basic methodologies are explored. First, we present a sequence of two statistical models that leverage bilingual information to improve the linguistic quality of syntactic parses, as measured by their ability to replicate human-generated gold-standard annotations. The first model uses word to word alignments as an external source of information, while the second models the alignments jointly. These models are both quite effective at improving the intrinsic quality of the parse trees, and the second model additionally improves word alignment performance. However, while the two models achieve similar parsing improvements, we find that improving parses in conjunction with word alignments is much more helpful for the downstream machine translation task. In the next part of the thesis, we explore this finding further by investigating the effects on MT performance of agreement between parse trees and word alignments. We present a simple method for transforming input trees in a way that ignores gold-standard annotations, concentrating instead on improving syntactic agreement directly. In experiments, we find that though we obviously lose fidelity to more linguistically informed treebank annotation guidelines, this transformation-based approach yields the strongest improvements in syntactic machine translation.
We often use tactile-input in order to recognize familiar objects and to acquire information about unfamiliar ones. We also use our hands to manipulate objects and utilize them as tools. However, research on object affordances has mainly been focused on visual-input and, thus, limiting the level of detail one can get about object features and uses. In addition to the limited multisensory-input, data on object affordances has also been hindered by limited participant input (e.g., naming task). In order to address the above mention limitations, we aimed at identifying a new methodology for obtaining undirected, rich information regarding people’s perception of a given object and the uses it can afford without necessarily viewing the particular object. Specifically, 40 participants were video-recorded in a three-block experiment. During the experiment, participants were exposed to pictures of objects, pictures of someone holding the objects, and the actual objects and they were allowed to provide unconstrained verbal responses on the description and possible uses of the stimuli presented. The stimuli presented were lithic tools given the: novelty, man-made design, design for specific use/action, and absence of functional knowledge and movement associations. The experiment resulted in a large linguistic database, which was linguistically analyzed following a response-based specification. Analysis of the data revealed significant contribution of visual- and tactile-input in naming and definition of object-attributes (color/condition/shape/size/texture/weight), while no significant tactile-information was obtained for object-features of material, visual-pattern, and volume. Overall, this new approach highlights the importance of multisensory-input in the study of object affordances.
Medical discharge documents are summaries written by a physician about the patient’s condition and aim at transferring information to other health care personnel but also to the patient. According to the legislation, the patient should be able to understand the document. In practice, however, this has been shown to be problematic. This paper studies discharge documents from the patients’ perspective and examines how they fulfil the legislation’s demands on understandability. Concentrating on the vocabulary of the texts, we analyse the frequency of domainadapted terms, abbreviations and foreign words. The material consists of 23 528 heart patients’ discharge documents (5 747 126 words). The analysis is performed with the morphological analyser FinTWOL (http://www2.lingsoft.fi/cgi-bin/fintwol). Altogether, FinTWOL analyses 24% of the corpus as unknown or foreign words, abbreviations or medical terms. The most common category, unknown words, includes misspellings and medical terms, such as l.dex. Of these, 100 most common cover for 43% of the total. These terms thus seem to be relatively fixed. Of the words analysed as abbreviations, some are common also in standard language, but others are still very domain-specific, such as I.V. (intravenous). Also the used abbreviations are very fixed: the 100 most common ones cover for 94% of the total. This, however, does not help the patient who probably reads only one document. Similarly, even though misspellings are globally infrequent, they still occur more than once per document. In order to place the obtained results in a context, we performed a similar analysis on general Finnish university newspaper text from Turku Dependency Treebank. In comparison with the 24% obtained with the discharge documents, from the total of 10 687 words, 8,6% were given a special tag. The results show that that terms and abbreviations are considerably more used in discharge documents than in general newspaper text. It is clear that a text with such a vocabulary is domain-specific and distinct from the language that the patient is used to. Also e.g. the varying use of upper and lower case letters (dg and DG for diagnosis) emphasize the particularity of the language. In standard language texts such writing would not be acceptable. Standard writing would, however, help the patients to better understand the texts.