Papers reviewed and determined not to be word norm studies. Use the flag icon to report errors or suggest re-inclusion.
16504 papers
Quantitative linguistics (QL) is a discipline of linguistics, that, using real texts, studies languages with quantitative mathematical approaches, aiming to precisely describe and explain, with a system of mathematical laws, the operation and development of language systems. Later in this review, we will address the relationship between QL and computational linguistics. Quantitative Syntax Analysis is a recent work on QL by Reinhard Köhler that not only provides a comprehensive introduction to the work of QL on the syntactic level, but also sketches the theoretical grounds, the research paradigm, and the ultimate goals of quantitative linguistics in general.In the first chapter, Köhler points to the vital role of syntax in language: Syntax enables language users to code structures instead of ideas as wholes. A text embodies a complex cognitive formation and meets several basic requirements in human communication, which implies that language is not autonomous, but a dynamic communicative system used by human beings. Hence, the ultimate understanding and explanation of syntax (and the whole language system) depends on usage-based investigation of the cognitive basis and the functional requirements of language, which is somewhat neglected in many mainstream syntactic studies. Even in those cases where explanatory power is acknowledged as the ultimate goal of linguistic investigation, the necessary knowledge is still required as to what a scientific linguistic theory is and how such a theory may be built. So far, it is rare for quantitative means to be used in syntactic study. One reason is that many syntacticians are too addicted to the enshrined traditional paradigms that have been proven to be somewhat inadequate when it comes to processing real texts. And that is why in computational linguistics, “devout executors of the belief in strictly formal methods as opposed to statistical ones do not have any chance to succeed” (page 4).The second chapter, entitled “The Quantitative Analysis of Language and Text,” begins with an explanation of the difference between quantitative linguistics and the formal branches of linguistics that have once been widely used in computational linguistics: QL is concerned with the quantitative properties important for understanding the development and the operation of linguistic systems, whereas the formal branches of linguistics use only qualitative mathematical means and formal logics to model structural properties of language, overlooking, in most cases, the aspects of systems that exceed structure, viz., functions, dynamics, and processes. Köhler points out that the successes of modern natural sciences (the exact, testable statements, the precise predictions, and the copious applications) all derive from their instruments and their advanced models. This implies that these instruments and models, for which the quantitative parts of mathematics (probability theory and statistics, function theory, differential equations) are indispensable ingredients, are worth integrating into linguistics, which is the aim of QL.Chapter 3, entitled “Empirical Analysis and Mathematical Modeling,” reviews the important works of quantitative syntactic analysis. In the first section, Köhler gives a long list of the important syntactic units and properties defined within the frameworks of both phrase structure syntax and dependency syntax; this reflects the fact that researchers in both fields have been engaged in some fruitful quantitative studies. In Section 3.2, he defines quantitation of syntactic concepts as counting the objects under study, because syntactic analysis investigates only discrete objects. Section 3.4 is a detailed review of the important works on various syntactic phenomena within the frameworks of both phrase structure syntax and dependency syntax, including sentence length, probabilistic grammars and probabilistic parsing, Markov chains, Frumkina’s law on the syntactic level, distribution of dependency distance, and distribution of dependency types, and so on. These quantitative models, which have been empirically corroborated with real texts, or sometimes treebanks and dictionaries (of various languages), can be linguistically, cognitively, or functionally interpreted—a rare achievement in the past statistical investigations of language. Apart from the models concerning probabilistic grammars and Markov chains, which have already been widely used in computational linguistics, there are some other works that may also have practical applications in various fields. For example, the mathematical model of sentence length, which describes the probability of neighboring length classes as a function of the probability of the first of the two given classes, may contribute to practical applications such as text classification and the measurement of text comprehensibility, and so forth. The frequency studies of word and syntactic constructions have obtained many results useful for language teaching, the construction of parsing algorithms, and the estimation of effort of (automatic) rule learning, and more. The syntactic studies on Frumkina’s law, which is concerned with the number of text blocks with x occurrences of a given syntactic element or category, may benefit certain types of computational text processing if specific constructions or categories can be differentiated and found automatically by their particular distributions. One advantage of QL is that all its findings are mathematically formulated and linguistically interpreted, which at least makes it possible to be used in constructing models necessary for computational linguistics.In Chapter 4, Köhler introduces his efforts to build a real “linguistic theory,” for he believes that “there is not yet any elaborated linguistic theory in the sense of the philosophy of science” (page 21). Building such a theory begins with “plausible hypotheses,” which may become laws when sufficiently attested and may then be further integrated into a coherent system. This is the process of setting up a scientific theory, as succinctly summarized in the title of this chapter: “Hypotheses, Laws, and Theories.”The first section of this chapter shows the first step toward a scientific linguistic theory, the process in which “plausible hypotheses” are deduced, interpreted, and empirically attested before finally becoming laws. In Section 2, Köhler introduces the foundation of his synergetic linguistics, which views language as a dynamic, self-organizing, and self-regulating system where the so-called enslaving principle and order parameters are the crucial elements. On this basis, the author builds a synergetic syntactic model in Section 4.2.7 with certain modeling principles. Within the framework of phrase-structure syntax, eight properties of syntactic constructions and four inventories are chosen to build this model, which are linked together by laws resulting from the verified hypotheses and subject to the regulation of some order parameters.Quantitative linguistics, which is an unfamiliar field of study for many linguists, depends heavily on real texts and mathematical tools. Therefore we believe it is worthwhile to briefly clarify the differences and the relations between QL and corpus linguistics, on one hand, and between QL and computational linguistics, on the other hand. In comparison with QL, corpus linguistics is in fact more of a research methodology rather than an independent linguistic discipline, reflecting a shift of focus from competence to performance, from introspection to empirical study. This makes both the common ground and the difference between corpus linguistics and QL, which aims to quantitatively and mathematically explore, on the basis of real texts and treebanks, the fundamental laws governing the structure and evolution of language, and integrate them into a systematic theory capable of explanation and prediction.Computational linguistics (CL) is an interdisciplinary field investigating the structure of natural languages from a formal, mathematical, and computational point of view. Compared with QL, CL seems to be more interested in research that can have direct applications in such fields as language understanding, language generation, machine translation, and so forth, than in the explanation of language structure, operation, and evolution. Traditionally, the mathematical models used in CL are derived from the linguistic theories via the process of formalization. Due to the limitations of the qualitative mathematical models, however, the statistical models, most of which so far fail linguistic interpretations, have recently become dominant in the field of computational linguistics. That is perhaps why Shuly Wintner (2009, page 641), in a “Last Words” article in this journal, called for “the return of linguistics to computational linguistics,” implying that the advances in the field of CL may ultimately lie in the advances in the understanding of language itself.Wintner’s appeal does reflect the present situation of computational linguistics: a discipline in which many works are heavily oriented towards engineering and weakly grounded in linguistics. The formal linguistic models seem now outshone by the purely statistical paradigms, as the result of their inadequacy in processing real-world languages. Of course, this means no renouncement of the value of the traditional formal models. But it is obvious that considerable updating and enrichment are necessary for these formal models if they are to play significant roles in the future. In this regard, QL, which has a solid linguistic foundation, may help by providing quantitative cues (which are linguistically interpretable) to improve the performance in NLP, as has been illustrated in the case of probabilistic grammar that ingeniously integrates quantitative, statistical devices into qualitative linguistic models. QL has provided some useful models for CL. And it is reasonable to believe that it will continue to do so in the future, though some of its achievements seem now not ready to be directly used in CL.But perhaps QL can do more than that. Though QL has not yet drawn much attention from computational linguistics, it is potentially a proper answer to Wintner’s call for the return of linguistics to computational linguistics. CL aims to replicate in computers the patterns of human language behavior, whereas QL endeavors to mathematically and quantitatively reveal the laws and the principles that govern human language behavior—this is a relation between theory and practice. The success of a scientific discipline is usually based on precise models that are well-grounded in profound understanding of the object of study. Computational linguistics is no exception. Two features of QL are hence noteworthy. One is that it aims to explain, within a certain linguistic framework, the operation and the evolution of language by uncovering the systematic cognitive and functional regulations that underlie human languages. The other is that it tries to mathematically model, with systematic and precise quantitative laws, these regulations and the resulting mechanism of language. In view of these two features, we believe that the success of QL will somehow and somewhat boost the studies in the field of CL and that the communication between QL and CL is and will be not only possible but also mutually beneficial. This is why we hold that it is worthwhile to recommend this book to researchers in the field of computational linguistics, a book presenting a panorama of QL in general and discoveries on the syntactic level in particular.
The chapter presents a meta-search tool developed in order to deliver search results structured according to the specific interests of users. Meta-search means that for a specific query, several search mechanisms could be simultaneously applied. Using the clustering process, thematically homogenous groups are built up from the initial list provided by the standard search mechanisms. The results are more user oriented, as a result of the ontological approach of the clustering process. After the initial search made on multiple search engines, the results are pre-processed and transformed into vectors of words. These vectors are mapped into vectors of concepts, by calling an educational ontology and using the WordNet lexical database. The vectors of concepts are refined through concept space graphs and projection mechanisms, before applying the clustering procedure. Implementation details and early experimentation results are also provided.
This thesis introduces an experimental and quantitative approach to language through the study of the concept of soft constraints and its application to two phenomena of order in French: the position of the attributive adjective and the ordering of verbal complements occurring in postverbal position. Soft constraints are defined as affecting the acceptability rather than the grammaticality of the sentences. Our main hypothesis is that these constraints are properties of the language and thus must studied in syntax. These constraints raise a methodological issue: since they do not affect the grammaticality of the sentences, they cannot be investigated using the traditional tools of syntax (introspection and grammaticality judgment). It is therefore necessary to define tools for their description and analysis. The proposed methods are statistical analysis of corpus date, inspired by the work of Bresnan et al. (2007) and Bresnan & Ford (2010) and, to a lesser extent, psycholinguistic experiment. Regarding the position of the adjective, we test most of the constraints encountered in the literature and we propose a statistical analysis of the data extracted from the French Treebank corpus. We show the importance of the adjectival item and the nominal item with which it combines. Other constraints linked to the internal syntax of the adjectival phrase and the noun phrase also play a significant role in the choice of position. The work on the relative order of the verbal complements is conducted on a sample of sentences extracted from two newspapers corpora (French Treebank and Est-Républicain) and two corpora of spoken French (ESTER and C-ORAL-ROM). We show the significant influence of the constituent weight over the ordering: short before long order which is a feature of SVO languages like French, is observed in over 86% of cases. We also identify the important role of the verbal lemma associated with its semantic class (annotated with the dictionary of Dubois & Dubois-Charlier, 1997). Finally, building on the analysis of corpus data as well a two questionnaires eliciting acceptability judgments, it seems that nor animacy neither information structure (given/new, Prince, 1981) have a significant effect on the postverbal complement ordering.
This paper introduces the error corpus of Korean learner English and mal rules to detect the errors. Based on the corpus, we classified 42 error types. Our criteria for error classification are more general in order to enhance agreement rate and decrease errors. For generating mal rule, we testified two different grammars. One is Context Free Grammar (CFG) from Penn Treebank. The other is the typed feature structure grammars based on the Head-Driven Phrase Structure Grammar (HPSG), using Natural Language ToolKit (Bird et al. 2009). We advanced grammatical formalism from CFG to HPSG since CFG needs abundant phrasal markers causing over-generation, structural ambiguity and complexity.
We investigate aspects of interoperability between a broad range of common annotation schemes for syntacto-semantic dependencies. With the practical goal of making the LinGO Redwoods Treebank accessible to broader usage, we contrast seven distinct annotation schemes of functor‐argument structure, both in terms of syntactic and semantic relations. Drawing examples from a multi-annotated gold standard, we show how abstractly similar information can take quite different forms across frameworks. We further seek to shed light on the representational ‘distance’ between pure bilexical dependencies, on the one hand, and full-blown logical-form propositional semantics, on the other hand. Furthermore, we propose a fully automated conversion procedure from (logical-form) meaning representation to bilexical semantic dependencies. †
Based on Kachru’s Three Circle Model on the spread of English in different parts of the world, I question how foreign university students migrating from Expanding Circle countries to Singapore deal with the “clash” of linguistic norms set by different circles. This research hence explores English speaking behavioural intentions of these foreign tertiary students in Singapore and how they account for their language behavioural plans. The collected data reveal that most students held a notion on the native English-Singlish dichotomy, seeing native English as standard and superior while regarding Singlish as improper and non-standard. Considering language behavioural intentions, most respondents claimed to adopt three main strategies: speech maintenance, adapting to the formality level of the communicative situation, and speech convergence. Looking into respondents’ accounts for these intended strategies, I argue that speakers orient their language use not only towards language perceptions but also the communicative situations they are in. However, the relative influences of these two factors on language use vary among different cases.
The neological issue in the OSLL is a component of the wider framework of the research plan Creation of a Lexical Database of the Czech Language of the Beginning of the 21st Century (2005–2011, head – K. Oliva). The construction of the lexical collections (archives of lexical dynamics) is dealt with by the excerption section; the theoretical understanding and lexicographic treatment are assured by an independent working group of lexicographers. Thanks to the research plan being resolved, it was possible to ensure the continual complementation of the neological excerption, modify the method of the accumulation of the material in connection with the new tasks of the department and modernise the software equipment (in connection with that to make part of the neological material accessible to the wider public). These results are built on by the theoretical and practical activities of the neological working group, focusing on the treatment of new material (2002–2010).
Abstract Since the late 1800s, the Uruguayan Government has attempted to enforce cultural and linguistic norms along the border with Brazil through the prohibition of Portuguese, especially in schools, despite the fact that this is the heritage language of most border residents. This research focuses on the differential use of Spanish and Portuguese in Rivera, the largest city on the border. Using self-reported data and metalinguistic commentaries extracted from interviews with 63 Spanish–Portuguese bilinguals, the use of both languages in various domains (home, school, work spaces) and with diverse interlocutors (family, friends, co-workers, superiors) is analyzed. Quantitative and qualitative analysis reveals that Portuguese, which has been marginalized for decades, is more frequently used in the home with relatives and close friends. The use of Portuguese in more formal domains, including schools, is much less frequent. The results from this study corroborate a perception within the community that Portuguese lacks the prestige of Spanish and provide further evidence of its status as a primarily home language. The current research does not show a progressive shift toward Spanish in Rivera nor does it support claims by other researchers that this community is diglossic. Keywords: SpanishPortuguese portuñol language contactlanguage usebilingual education Acknowledgements This research would not have been possible without the generous support of a Tinker Foundation Grant from the Latin American and Iberian Institute of the University of New Mexico, which allowed me to conduct fieldwork in Rivera. I am also extremely grateful for the advice given to me by Ana Maria Carvalho, Adolfo Elizaincín, and Rena Torres-Cacoullos. Thank you also to the people of Rivera, who welcomed me into their community with open arms. Notes 1. Definitions of bilingualism should not be based on language use alone, however, since attitudes toward both languages also shape a bilingual's identity (Ben-Rafael, Olshtain, and Geijst Citation1998; Hoare Citation2001; Joseph Citation2004; Lawson and Sachdev Citation2004). In other words, bilingualism is not merely a linguistic phenomenon involving frequency of use, but rather a social one as well, which is defined largely by its role within the community. Consequently, these attitudes affect the bilingual's choice of one language or another. 2. Five of six of these speakers are of the second generation, while only one is of the third generation (speaker 15, female, professional occupation). The social attributes of each consultant, however, vary. Of the second generation, two males (one with a professional job and one with a nonprofessional job) did not complete the questionnaire. Likewise, the three women who did not complete the questionnaire cover both occupational classes (two professionals and one nonprofessional), thereby maintaining the diversity of the original composition of social characteristics for this generation. 3. Nonprofessionals differ from professionals in that the work they perform does not require formal academic training. These members of the community are taxi drivers, shopkeepers and their employees, hotel owners and their employees, waiters, bartenders, construction workers, etc. 4. Some of the younger consultants indicated percentages of language use with a spouse, which most clarified by writing in the margin novio or novia 'boyfriend' or 'girlfriend.' These percentages were not included in the rate calculations presented in Table 4 since these interlocutors do not belong strictly to home domains.
The aim of this paper is to analyse the productivity of the Old English weak verbs suffixed with -læcan. The main sources for this research are the lexical database of Old English Nerthus and The Dictionary of Old English Corpus. The assessment of productivity is based on the distinction between type-frequency (dictionary-based) and token-frequency (corpus-based). This work contributes to a methodology for assessing the productivity of a morphological process in a historical language as well as for dealing with very low indexes of productivity. The conclusion is reached that the type-frequency of -læcan is relatively high, whereas its productivity is considerably low. It may be affirmed then that a type-frequency higher than token-frequency is compatible with a rather unproductive affix. Finally, the analysis evidences that -læcan suffixed verbs are much more frequent in prose and glosses than in poetry.
The Noun Phrase (NP) is the dominant construct in natural language text. While base NPs (BNP) and maximal length NPs (MNP) are relatively easy to identified and extracted, the internal structure of NPs is rather a challenge in natural language processing. Penn Treebank leaves the BNPs flat as implicit right branching. Vadas and Curran added BNP internal structure to the Penn Treebank. But the results of the BNP structure are very often incorrect when it is considered within a longer complex NP (CNP). Structural ambiguity prevails in most CNPs and multilingual comparison may help improve disambiguation. We introduce a new NP annotation scheme, which is applicable to multilingual parallel corpora and discriminate genuine flat branching and right branching. Flat branching is preferred instead of binary branching wherever appropriate so as to achieve inter-lingual consistency. As a pilot task to build a gold standard corpus for structural and semantic analysis of CNPs, 381 document titles are extracted from the UN resolutions as typical examples of CNPs. Document titles in Chinese, English and Russian are manually annotated in XML format with the hope to help acquire rules for parsers or machine translators targeted at CNPs. The problems encountered are reported.
This paper focuses on the links between contemporary literature and the various positions choosenchosenby authors facing the problematics of translation. Beginning with the observation that translation studies should develop from a theoretical point of view in Japan--an emblematic country for translations--this paper shows that currently, translation in Japan has to be considered as a cultural exportation trend and not only as the importation trend that dominated the cultural scene during the 20th century. For example, data on published translations in France show that since 2007, Japanese is the second most frequently translated language after American-English--due to the popularity of mangas in France. In the literary field, new phenomenons can also be observed in Japan. In this paper, four case studies are presented. The most remarkable case concerns Murakami Haruki's strategy, in which he, being an important translator of the Great American Novel, crosses the boundaries between countries and languages in order to represent a new kind of nationless writer, i.e. a global writer appreciated all over the world. On the other hand, Mizumura Minae mixes English and Japanese in her I novel from left to right, making it untranslatable into English. This for her represents the resistance of a minor language, Japanese, to the domination of English. Tawada Yôko, for her part, writes in two languages, Japanese and German, and in doing so tries to deconstruct both cultural and linguistic norms, enhancing translation as an impossible tool. Finally, the American-born Hideo Levy's three-piece band features Japanese, English and Chinese members, interconnected by the belief in translation as an ideal vector of communication. All these new streams contribute to the reshifting of Japanese literature in the world and induce a necessary renewal of the critical approaches.
The explosion of information in the World Wide Web is overwhelming readers with limitless information. Large internet articles or journals are often cumbersome to read as well as comprehend. More often than not, readers are immersed in a pool of information with limited time to assimilate all of the articles. It leads to information overload whereby readers are trying to deal with more information than they can process. Hence, there is an apparent need for an automatic text summarizer as to produce summaries quicker than humans. The text summarization research on mobile platform has been inspired by the new paradigm shift in accessing information ubiquitously at anytime and anywhere on Smartphones or smart devices. In this research, a semantic and syntactic based summarization is implemented in a text summarizer to solve the overload problem whilst providing a more coherent summary. Additionally, WordNet is used as the lexical database to semantically extract the text document which provides a more efficient and accurate algorithm than the existing summary system. The objective of the paper is to integrate WordNet into the proposed system called TextSumIt which condenses lengthy documents into shorter summarized text that gives a higher readability to Android mobile users. The experimental results are done using recall, precision and F-Score to evaluate on the summary output, in comparison with the existing automated summarizer. Human-generated summaries from Document Understanding Conference (DUC) are taken as the reference summaries for the evaluation. The evaluation of experimental results shows satisfactory results.
The study of the Tip of the Tongue phenomenon (TOT) provides valuable clues and insights concerning the organisation of the mental lexicon (meaning, number of syllables, relation with other words, etc.). This paper describes a tool based on psycho-linguistic observations concerning the TOT phenomenon. We've built it to enable a speaker/writer to find the word he is looking for, word he may know, but which he is unable to access in time. We try to simulate the TOT phenomenon by creating a situation where the system knows the target word, yet is unable to access it. In order to find the target word we make use of the paradigmatic and syntagmatic associations stored in the linguistic databases. Our experiment allows the following conclusion: a tool like SVETLAN, capable to structure (automatically) a dictionary by domains can be used sucessfully to help the speaker/writer to find the word he is looking for, if it is combined with a database rich in terms of paradigmatic links like EuroWordNet.
Statistické jazykové modely jsou důležitou součástí mnoha úspěšných aplikací, mezi něž patří například automatické rozpoznávání řeči a strojový překlad (příkladem je známá aplikace Google Translate). Tradiční techniky pro odhad těchto modelů jsou založeny na tzv. N-gramech. Navzdory známým nedostatkům těchto technik a obrovskému úsilí výzkumných skupin napříč mnoha oblastmi (rozpoznávání řeči, automatický překlad, neuroscience, umělá inteligence, zpracování přirozeného jazyka, komprese dat, psychologie atd.), N-gramy v podstatě zůstaly nejúspěšnější technikou. Cílem této práce je prezentace několika architektur jazykových modelůzaložených na neuronových sítích. Ačkoliv jsou tyto modely výpočetně náročnější než N-gramové modely, s technikami vyvinutými v této práci je možné jejich efektivní použití v reálných aplikacích. Dosažené snížení počtu chyb při rozpoznávání řeči oproti nejlepším N-gramovým modelům dosahuje 20%. Model založený na rekurentní neurovové síti dosahuje nejlepších publikovaných výsledků na velmi známé datové sadě (Penn Treebank).
The focus of this article is on the creation of a collection of sentences manually annotated with respect to their sentence structure. We show that the concept of linear segments—linguistically motivated units, which may be easily detected automatically—serves as a good basis for the identification of clauses in Czech. The segment annotation captures such relationships as subordination, coordination, apposition and parenthesis; based on segmentation charts, individual clauses forming a complex sentence are identified. The annotation of a sentence structure enriches a dependency-based framework with explicit syntactic informa- tion on relations among complex units like clauses. We have gathered a collection of 3,444 sentences from the Prague Dependency Treebank, which were annotated with respect to their sentence structure (these sentences comprise 10,746 segments forming 6,341 clauses). The main purpose of the project is to gain a development data—promising results for Czech NLP tools (as a dependency parser or a machine translation system for related languages) that adopt an idea of clause segmentation have been already reported. The collection of sentences with annotated sentence structure provides the possibility of further improvement of such tools.
It is well known that accuracies of statistical parsers trained over Penn treebank on test sets drawn from the same corpus tend to be overestimates of their actual parsing performance. This gives rise to the need for evaluation of parsing performance on corpora from different domains. Evaluating multiple parsers on test sets from different domains can give a detailed picture about the relative strengths/weaknesses of different parsing approaches. Such information is also necessary to guide choice of parser in applications such as machine translation where text from multiple domains needs to be handled. In this paper, we report a benchmarking study of different state-of-art parsers for English, both constituency and dependency. The constituency parser output is converted into CoNLL-style dependency trees so that parsing performance can be compared across formalisms. Specifically, we train rerankers for Berkeley and Stanford parsers to study the usefulness of reranking for handling texts from different domains. The results of our experiments lead to interesting insights about the out-of-domain performance of different English parsers.
The majority of fear conditioning studies in humans have focused on fear acquisition rather than fear extinction. For this reason only a few functional imaging studies on fear extinction are available. A large number of animal studies indicate the medial prefrontal cortex (mPFC) as neuronal substrate of extinction. We therefore determined mPFC contribution during extinction learning after a discriminative fear conditioning in 34 healthy human subjects by using functional near-infrared spectroscopy. During the extinction training, a previously conditioned neutral face (conditioned stimulus, CS+) no longer predicted an aversive scream (unconditioned stimulus, UCS). Considering differential valence and arousal ratings as well as skin conductance responses during the acquisition phase, we found a CS+ related increase in oxygenated haemoglobin concentration changes within the mPFC over the time course of extinction. Late CS+ trials further revealed higher activation than CS– trials in a cluster of probe set channels covering the mPFC. These results are in line with previous findings on extinction and further emphasize the mPFC as significant for associative learning processes. During extinction, the diminished fear association between a former CS+ and a UCS is inversely correlated with mPFC activity – a process presumably dysfunctional in anxiety disorders.
Does international law's effectiveness require a clear distinction between law and non-law? This essay, which reviews Jean d'Aspremont's Formalism and the Sources of International Law, argues the answer is no. Ambiguity about the legal nature of international instruments has important benefits. Clarity in the law may encourage states to do the minimum necessary to comply, while some uncertainty about what the law requires may induce states to take extra efforts to ensure they are in compliance. Ambiguity in the law also promotes dynamic change, an important feature in rapidly developing areas of the law such as international environmental law and human rights. Most importantly, though, soft law — international instruments that have legal consequences but are not unambiguously 'law' — expands the range of instruments available to states when cooperating. Institutionalist theories of international law suggest that a larger menu of international instruments is valuable because it allows states to calibrate the level of their commitments more precisely, thereby expanding their ability to cooperate. Institutional theories, however, have heretofore not explained exactly how states communicate to each other the level of their commitment; that is, they have not explained how states mark an instrument as soft law and whether and how states distinguish between types of soft law commitments. A theory of law-identification based on linguistic norms, such as d'Aspremont proposes, offers a descriptive account of how states might signal levels of legal commitment beyond the dichotomy of 'binding' and 'non-binding' law. A communicative theory of international law — one based on the use of language in international instruments to signal relatively fine-grained variation in the level of commitment — thus would enrich our understanding of what soft law is, and when and how states use it.
Nowadays, the volume of information increases exponentially, forcing the corporations to keep their business information distributed under several heterogeneous sources such as relational databases, spread sheets, XML documents and Web pages, and stored under different structures and formats. Integrating heterogeneous sources is recently acknowledged as an important vision on semantic web research. The concept of heterogeneity arises at different levels: from the lexical level to the semantic or structural level. For discovering and consolidating the semantic relationships among the semantically related data present in different types of databases and files, this paper presents the enhancements obtained due to the use of available online large lexical databases, combined with lexical and structural similarity models and the available source metadata. Finally, we reveal the experimental results that demonstrate the applicability and usability of our approach.
In this paper we employ a most recent approach to Data Oriented Parsing (DOP), which has named Double-Dop, for Persian sentences. Like other DOP models, Double-Dop parser utilizes syntactic fragments of arbitrary size from a treebank to analyse new sentences, but it extracts a restricted yet representative subset of fragments. It uses only those which are encountered at least twice. The accuracy of Double-DOP is well within the range of state-of-the-art parsers currently used in other NLP-tasks, while offering the additional benefits of a simple generative probability model and an explicit representation of grammatical constructions. Heretofore there isn’t any standard parser for Persian language and this work try to employ Double-Dop Method for parsing Persian sentences.
This paper outlines a proposal for maritime English language teaching in public and private Nautical Schools and other maritime educational institutions and establishments in Italy, using a content and language integrated learning (CLIL) approach. The courses are addressed in particular to those students who would like to take up a marine career as officers, engineers or other crew members of the Merchant Navy, and thus require an adequate knowledge of seafaring terminology, but can also be interesting for those wishing to explore the origins and development of maritime language. In order to provide a more challenging environment and better opportunity for the learning of seafaring terms and expressions in English, students are supported by Mariterm, a lexical database, organized in semantic relations, available at the Institute for Computational Linguistics (ILC) of the National Research Council (CNR) in Pisa. A
Abstract. In this paper we present an adaptation of two Czech syntactic analyzers Synt and SET for Slovak language. We describe the transformation of Slovak morphological tagset used by the Slovak development corpora skTenTen and r-mak-3.0 to its Czech equivalent expected by the parsers and modifications of both parsers that have been performed partially in the lexical analysis and mainly in the formal grammars used in both systems. Finally we provide an evaluation of parsing results on two datasets – a phrasal and dependency treebank of Slovak.
We present a new dependency parsing algorithm based on the decomposition of large sentences into smaller units such as clauses and intraclausal coordinations. For the identification of these units, new methods combining machine learning techniques and heuristic rules were developed. The algorithm was evaluated on the Slovene dependency treebank text corpus. Compared to the MSTP parser, currently the most accurate for Slovene, parsing accuracy was improved by 1.27 percentage points, which equals 6.4 % relative error reduction.
Recent approaches for building syntactic language models include the combination of Probabilistic Tree Substitution Grammars (PTSGs) and Bayesian learning methods. While PTSGs have appealing features for syntax modeling, Bayesian methods provide a framework for inducing compact grammars that do not overfit the training corpus. In this paper, we apply these approaches to learn syntactic language models from a Brazilian Portuguese treebank. © 2012 Springer-Verlag.
We present a mentod to combine the maximum spanning tree(MST) algorithm and the deterministic algorithmfor Chinese dependency parssing.We introduce the results and the dependency degree of Nivre parser into MST parser.Our system achieves the accuracy of 86.49% using 10-fold cross-validation on the Penn Chinese Treebank Corpus,which is a significant improvmentin the parsing accuracy.
In the past few years, much attention has been paid on extending phrase-based statistical machine translation with syntactic structures. In this paper we introduce a novel syntax encapsulated phrase(SEP) model, in which treebank tag sequences are employed to decorate the bilingual phrase pairs. We use tag sequences, instead of phrase pairs, to train the lexicalized reordering model. Since the number of treebank tags is much smaller than the number of words, the tag sequence based reordering model is smaller and more accurate than the phrase based reordering model. Experiments were carried out on four types of models: the phrase model, the hierarchical phrase model, the POS tag encapsulated phrase(PTEP) model and the syntactic tag encapsulated phrase(STEP) model. The STEP model obtained higher BLEU-4 score than other models on NIST 2005 MT task.
The degree of translation adequacy and full-value depends on its compliance with the existing general linguistic norms. Vocabulary potential of a translator is determined by the proficiency of language to translate into. Key moments in course of transferring means of another language text are those three main features: context, word-collocations, the knowledge of ethnic specifications. Meanings of words and sentences and even whole abstracts are not autonomous, and depend on the general distributions and surroundings.
The degree of translation adequacy and full-value depends on its compliance with the existing general linguistic norms. Vocabulary potential of a translator is determined by the proficiency of language to translate into. Key moments in course of transferring means of another language text are those three main features: context, word-collocations, the knowledge of ethnic specifications. Meanings of words and sentences and even whole abstracts are not autonomous, and depend on the general distributions and surroundings.
This dissertation asks whether and to extent linguistic meaning is conventional. By combining Davidson's theory of meaning with Nancy's conception of community it develops a model of communication in which the meaning of words is not determined by a fixed set of norms, but is constantly negotiated through a multiplicity of concrete communicative events. By presenting linguistic norms as fluid and contested, the dissertation undercuts the idea that linguistic communities are unified in the way that the idea of convention suggests, thus exposing the extent to which the boundaries of linguistic communities are subject to constant political and cultural negotiation. Chapters 1 and 2 present Davidson's and Nancy's position that is nothing but a pattern of relations between the observable behaviors and practices of speakers and interpreters. Accordingly, the set of conventions we call language is supervenient upon actual occasions of interpretation, and is only a part of the variety of means we employ in order to interpret one another and ascribe meaning to utterances. So while shared linguistic norms are often employed in interpretation, they are not a precondition of communicative success. Chapters 3 and 4 argue that if conventions are secondary to actual interpretation, then the interpreter's first task is to determine constitutes linguistic behavior, i.e. when a pattern of behavior justifies the attribution of intentions to a creature. The question of speakerhood, of makes one take another being to be a creature whose behavior is potentially meaningful and warrants interpretation, thus emerges as a crucial issue. It is argued that speakerhood cannot be reduced to a biological or cognitive fact, but is subject to constant social negotiation, which is concealed by linguistic norms that reify meaning and make the answer to the question what and who is meaningful? seem more settled than it is. By posing the question of nonhumans' participation in communication, chapter 5 calls attention to the need to articulate and examine the various constraints that can prevent a creature (whether human or not) from being considered a speaker, and thus receiving a fair opportunity to participate in a linguistic community.
Discourse is coming in from the cold. After years of being ignored by researchers in other areas of computational linguistics and language technology, many of these same researchers are beginning to think that their own work could benefit from treating text as more than just a bag of sentences. That is, they are beginning to think that discourse offers some low-hanging fruit—achievable improvements in system performance that exploit either aspects of text structure or the context that text establishes and uses for efficient referring and/or predicational expressions.This new monograph on Discourse Processing by Manfred Stede both reflects this new zeitgeist and provides an introduction to discourse for researchers in computational linguistics or language technology with little or no background in the area. This clear and timely monograph consists of a brief introduction to discourse, a meaty chapter on each of the three aspects of discourse processing that hold most promise for language technology, and a brief conclusion on where discourse research might go in the future. I will go through the three major chapters, and then make some general remarks.Chapter 2Chapter 2 addresses two distinct types of large-scale discourse structure: structure that follows from a text belonging to a particular genre, and structure that follows from the topic (or topic mix) of a text. The genre of a text affects features such as style and register. What is relevant here is structure that genre may confer on a text. Stede suggests that some, but not all, texts inherit large-scale structure from their genre, calling some unstructured, some structured, and some semi-structured. As a reader, I did not find this distinction useful, because all text that belongs to a genre seems to get some large-scale structure from it. On the other hand, all or part of this structure might simply not be manifest in the kind of lexico-syntactic features that automated systems regularly rely on for text segmentation. As a case in point, although Stede offers the text Suffering (used as a running example throughout the book) as an example of unstructured text, like other instances of Comments in the Talk of the Town section of the New Yorker magazine, its large-scale structure comprises a “hook” aimed at getting the reader's attention, followed by a short essay that concludes with a serious point. Although ways of attracting a reader's attention may not have specific lexico-syntactic features, it might still be possible to recognize the transition between “hook” and essay, and essay structure itself is what ETS's eRater system (Burstein and Chodorow 2010) aims to recognize and evaluate.This first half of Chapter 2 focuses on the genre-based structure of scientific texts and of film reviews. Here researchers have already shown that language technologies such as information extraction and sentiment analysis benefit from taking such structure into account, so this is entirely appropriate for the book's target audience. More on genre-based functional structure and its use in producing structured biomedical abstracts can be found in the recent survey of research on discourse structure and language technology by Webber, Egg, and Kordoni 2012.The second half of Chapter 2 discusses large-scale discourse structure associated with patterns of topics. Such structure is often found in expository writing such as encyclopedia articles and travel pieces. Here, changing patterns of content words correlate well with changes in topic, rendering them useful for the many approaches to text segmentation that are well-described in this half of the chapter. Because the discussion here of probabilistic models for topic segmentation is rather short, the reader who wants to know more should consult the excellent survey of topic segmentation methods by Purver (2011).Chapter 3Chapter 3, entitled Coreference Resolution, addresses more than this, dealing with the resolution of other expressions whose reduction is licensed by the discourse context, such as bridging reference and “other” reference, which Halliday and Hasan 1976 call comparative reference because it occurs with comparative forms such as “larger fish” and “a more impressive poodle,” as well as with “other,” “another,” and “such.” Stede justifies inclusion of this chapter for two reasons—the close connection between coreference resolution and topic segmentation and the benefits to text analysis provided by having its pronouns resolved. But another reason must be the link mentioned earlier between text and context: Discourse creates the context in which context-reduced expressions make sense, so it falls naturally within the tasks of discourse processing to resolve them, either through modeling context explicitly or through the use of proxies.The chapter starts with an overview of coreference and anaphora that covers both their forms and their functions. This is followed by an important section on corpus annotation (Section 3.2), included because (as Stede notes) what has been annotated and why it has been annotated strongly determines what expressions are resolved and how. This section identifies many of the problems in coreference annotation that have been raised in the literature, but recognizes that research has to make use of the resources that exist and not just the resources it wants. Several of these are indicated at the end of the section, reminding one that it would have been useful to have some pointers in Chapter 2 to corpora available for genre-based segmentation (such as Liakata's ART corpus)1 or for topic-based segmentation.Stede then links the current chapter to the previous one through a discussion of entity-based coherence (Section 3.3) and then discusses how to identify when a pronoun or definite noun phrase should be treated as anaphoric (Section 3.4) as groundwork for discussion of anaphora resolution (Sections 3.5–3.7). Missing from the discussion of detecting non-anaphoric (pleonastic) pronouns is mention of Bergsma's recent system NADA for doing this (Bergsma and Yarowsky 2011).2The discussion of anaphora resolution covers rule-based approaches to resolving nominal anaphora (Section 3.5) and then supervised machine learning methods for anaphora resolution (Section 3.6). The latter follows the structure (albeit not the content) of Ng's survey 2010, in discussing mention-pair models, and then entity-mention models. Whereas Ng then discusses ranking models, including his cluster ranker (Rahman and Ng 2009), which is conceptually similar to the Lappin and Leass 1994 approach described in Section 3.5, Stede discusses a range of more recent models, most of which are subsequent to Ng's survey.Section 3.8 surveys methods evaluating coreference resolution and some of the known problems in doing so. A good complement to this is Byron's too-little-known discussion of problems in the consistent reporting of such results (Byron 2001). Chapter 3 concludes with a section on Recent Trends, which would also have been useful in Chapter 2.Chapter 4The fourth and longest chapter deals with semantic or pragmatically oriented coherence relations that hold between adjacent text spans or discourse units. Whereas the previous two chapters were essentially theory-neutral, the presentation in Chapter 4 largely reflects the perspective of Rhetorical Structure Theory (Mann and Thompson 1988). RST takes a text to be a sequence of elementary discourse units that comprise the leaves of a tree structure of coherence relations between recursively defined discourse units. RST also assumes that one of the arguments to a coherence relation may be more important to the speaker's purpose than the other, calling the former the nucleus and the latter, the satellite.This RST framework dictates the structure of the chapter: Following an introductory section that explains and motivates coherence relations, each subsequent section considers the next task in an RST analysis—segmenting a text into elementary discourse units (Section 4.2), recognizing which (adjacent) units stand in a coherence relation and what (single) relation holds between them (Section 4.3), and finally, inducing the overall tree structure of coherence relations that hold between recursively defined discourse units (Section 4.4). All these tasks are well described, both from a theoretical perspective and in terms of automated procedures for carrying them out. Coverage of relevant work is very high.Where the reader may get confused, however, is that a good proportion of the more recent work on identifying coherence relations does not fall within the framework of RST, and thus doesn't adhere to several of its assumptions—in particular, that a text is divisible into a covering sequence of elementary discourse units, that only one relation can hold between discourse units, that the arguments to a coherence relation must be adjacent, that one argument to a coherence relation may intrinsically convey information that is more important to the speaker's purpose than the other, and that coherence relations impose an overall tree structure on a text in terms of recursively defined discourse units.Although Chapter 4 discusses the Penn Discourse TreeBank (Prasad et al. 2008) and its “somewhat modest annotations” (page 126), the discussion is framed in terms of RST tasks, whereas the assumptions underlying the Penn Discourse TreeBank reflect its concerns with a quite different set of tasks involved in recognizing coherence relations. The first task requires finding evidence for a coherence relation (in the form of a discourse connective such as a coordinating or subordinating conjunction or a discourse adverbial, or in the form of sentence adjacency) and then determining (1) if the evidence does indeed signal a coherence relation, given that evidence is often ambiguous; (2) if it does, what constitutes its arguments; and (3) what is its sense. Although Chapter 4 covers some of this work (Dinesh et al. 2005; Wellner and Pustejovsky 2007; Elwell and Baldridge 2008; Pitler and Nenkova 2009; Prasad, Joshi, and Webber 2010), its appearance within the context of a discussion of RST-tasks may lead to some confusion.Chapter 4 concludes with a brief discussion of some important open issues regarding coherence relations, including problems with associating a large text span with a single recursive structure of coherence relations and problems with inter-annotator agreement.SummaryFor its intended audience, this monograph will serve as a compact, readable introduction to the subject of discourse processing. The relevant phenomena are presented clearly, as are many of the computational methods for dealing with them. What readers won't get is criteria for choosing among the methods or an understanding of what each method is good for. This problem may reflect the absence of comparable performance results and useful error analyses in the original publications, however.Also missing from the monograph is discussion of applications of discourse processing, and pointers to more of the resources available to researchers interested in discourse structure. This is where the additional resources I have mentioned may prove complementary.Finally, a plea to the series editor: Monographs such as this one really need an index. Some monographs in the series have one, whereas others (like this one) don't. Because the series appears in both electronic and physical format, one could excuse the former not having an explicit index, since in most cases, one can get away with the basic search facility in the Adobe Reader. Nothing similar is available for the nicely sized physical monographs. Their authors should be strongly encouraged to provide them.
The choice,process and product of translation are directly decided by the core subjective element-translator,by whom the external cultural norms and internal linguistic norms can be concretized in the target texts.The paper adopts Pierre Bourdieu's concepts of habitus and field to study the relationship between norms and translator's subjectivity in literary translation system in order to justify translator's translation activity.
Development of information and communication technologies is accompanied by a variety of views surrounding new forms of language use, literary practices and general communication norms. The authors prove that development of information technologies serves an indicator for important changes in linguistic norms development.
There is increasing interest in the nature of the emotion recognition deficit in Huntington9s disease (HD). Recognition of all negative emotions tends to be impaired in HD, particularly in the facial domain. This study aims to evaluate brain9s reactivity to emotional pictures by recording event-related potentials and assessing their relationship to evaluative measures of affect in a cohort of 20 early HD patients, compared to 24 age and sex matched controls. Fourty-two colour slides were chosen from the International Affective Picture System. Each subject was requested to attribute a valence and an arousal rating. (Centre for the Study of Emotion and Attention, 1995). We placed 19 scalp electrodes, for the recording of slow late positive potential (LPP) in the 400–800 ms following pictures presentation. In HD patients the valence and arousal rating were within normal ranges for pleasant, neutral and unpleasant pictures. The amplitude of LPP was slightly reduced during unpleasant pictures viewing. We found a positive correlation between LPP amplitude and functionality scales. In the early phase of HD, the selective processing of emotional stimuli seemed not to be seriously impaired, and may indicate a preservation of high cortical functions in the initial stages of the disease.
We describe our participation in the MTPIL Hindi Parsing Shared Task-2012. Our system achieved the following results: 82.44% LAS/90.91% UAS (auto) and 85.31% LAS/92.88% UAS (gold). Our parser is based on the linear classification, which is suboptimal as far as the accuracy is concerned. The strong point of our approach is its speed. For parsing development the system requires 0.935 seconds, which corresponds to a parsing speed of 1318 sentences per second. The Hindi Treebank contains much less different part of speech tags than many other treebanks and therefore it was absolutely necessary to use the additional morphosyntactic features available in the treebank. We were able to build classifiers predicting those, using only the standard word form and part of speech features, with a high accuracy.
This article, drawing upon Juraj Dolnik’s book on the theory of standard language with regard to standard Slovak (2010), concentrates on the question of the sources of standard variety and the problem of objectivity of scientific knowledge. Reconsidering Dolnik’s concept of norm critically, it places emphasis on the fact that linguistic norms, as a part of social norms, are constituted in interactions, which helps to explain their indexicality. It also argues that language users are actors in social processes who hold specific social roles, which corresponds to their differing power (and vice versa). Referring to Language Management Theory, the article concludes with some more general arguments in favor of qualitative methodology in the research on linguistic norms and the standard variety.
Starting from the definition of treebanks and considering that treebanks are theory dependent, we propose an annotation scheme for Romanian using several approaches ranging from phrase structure to dependency grammars and property grammars. The annotation has its starting point in a generative grammar study of the Romanian AP and validates the data of the linguistic study using an annotation scheme consisting of a constraint based approach.
This paper describes the methods used for the parsing the Sinica Treebank for the bakeoff task of SigHan 2012. Based on the statistics of the training data and the experimental results, we show that the major difficulties in parsing the Sinica Treebank comes from both the data sparse problem caused by the fine-grained an-notation and the tagging ambiguity. 1
In this article, we first present the overall structure of the Pralex lexical database, the work with the data entry form and Praled’s functions. We then focus on the general principles of the elementary processing of database entries, which are subsequently specified according to the individual word classes and entry types. The article ends with a specific example of the processing of an entry in the Pralex lexical database.
<h3>Objective</h3> Although language and culture are different in each area in the world, how does culture affect the recognition of nonverbal emotional vocalisations? A previous study on non-verbal emotional vocalisations has shown a cross-cultural effect in Western and African participants. However, nobody has ever investigated the cross-cultural differences between Japanese and Caucasian participants in their emotional response to non-verbal vocal sounds. In the present study, we aimed to investigate cross-cultural effect between Caucasian subjects and Japanese subjects when the subjects listen to nonverbal affect bursts. <h3>Method</h3> 30 Japanese subjects (15 males) participated in this study. The data of Japanese subjects were compared with data from 30 Canadian subjects (15 males). Subjects listened to the Montreal Affective Voices (MAVs), which consist of a database of nonverbal affect bursts portrayed by Canadian actors. Each voice was evaluated using three criteria: perceived emotional intensity in each of the eight emotions (Anger, Disgust, Fear, Pain, Sadness, Surprise, Happiness, and Pleasure), perceived valence, and perceived arousal. To investigate cross-cultural differences between Japanese and Canadian participants, mixed 2×8 ANOVAs with Group (Japanese, Canadian) and Emotion (eight emotions) as factors were calculated on ratings of Intensity, Valence, and Arousal. <h3>Results</h3> Significant Group × Emotion interactions were observed for ratings of intensity, valence and arousal (intensity: F(5.5, 313.5) = 9.137, p<0.001, valence: F(4.3, 244.3) = 25.101, p<0.001, arousal: F(4.4, 250.5) = 8.955, p<0.001). Post-hoc tests showed that intensity ratings from Japanese listeners were significantly higher than ratings from Caucasian listeners for angry, disgusted, fearful, surprise, and pleased (p<0.05/8). Further, valence ratings from Japanese listeners were significantly higher than ratings from Caucasian listeners for angry, disgusted, fearful, painful, surprised (p<0.05/9), whereas valence rating of pleased in Japanese listener was significantly lower than in Caucasian listener (p<0.05/9). Arousal ratings of sad vocalisations by Japanese listeners were significantly higher than by Caucasian listeners (p<0.05/9). <h3>Conclusion</h3> This study demonstrates important cross-cultural differences in the perception of non-verbal affect bursts and extends recent observations by showing that these cross-cultural differences are also found for negative emotions.
American linguistics in the 20th century witnessed two totally opposite but equally powerful lines of thought: the former was represented by the mentalist direction in anthropological linguistics, initiated by Franz Boas and carried on by Edward Sapir and Benjamin Lee Whorf, while the latter was Leonard Bloomfield's structuralism. Mentalist descriptivists considered that the language of a people could only be studied in close connection to the culture of that specific people, because their language structures reflected the speakers' mentality (or particular way of seeing the world), influencing it, at the same time, by transfer of linguistic norms into the domain of experience. At the opposite pole, structuralist linguists (Bloomfield, Zellig Harris, Noam Chomsky), despised any reference of language analysis to mental processes, since they strongly believed this approach was highly unscientific, and made a purpose for themselves to completely ignore meaning.Leonard Bloomfield is considered to have been for American linguistics what de Saussure was, in his time, for the European one. He set out to take a fully scientific, empirical approach to language, which meant exclusively direct observation of visible phenomena. He thought variability in human behavior was due solely to the complexity of the human organism, especially that of the human nervous system. In this line of thought, any of man's activities could be described in purely scientific terms. For example, a word like salt can be described very accurately because we can give it a scientific definition: NaCl or sodium chloride. In contrast, words like love or hate cannot be defined yet, since we lack the necessary knowledge at present, but the future might change that.As for human language in its totality, Bloomfield considered it to be a system of conditioned reflexes triggered by verbal stimuli. This is how he explains the act of speech: a speaker produces a noise (stimulus) that triggers a reaction (response) in the nervous system of the interlocutor (receptor). Any reference to ideas is left out in the description, because Nonlinguists /.../ constantly forget that a speaker is making noise, and credit him, instead, with the possession of impalpable. It remains for the linguist to show, in detail, that the speaker has no and that the noise is sufficient [1].Judging things from a wider perspective, we notice that American structuralism is well impregnated with behaviorist values. Indeed, the psychology of the same early 20th century was dominated by behaviorists. In Russia, the principles of classical conditioning had been firmly established by Ivan Pavlov and his experiments that demonstrated how an originally neutral stimulus (a sound), if repeatedly associated with a reflex (salivation before food), induced a reflex response. In the U.S.A., psychologists like John Watson and B. F. Skinner founded their research on the direct observation of animal behavior, ruling out any reference to concepts like mind or psychic phenomena. Any behavior originated in learning to associate a stimulus and a response.The fundamental laws of learning were extracted from this type of study, but they were applied on man as well. This is a very bizarre extrapolation from the mechanic nature of animal kingdom activities to the infinite complexity of human activities. Speech, for example, displays features of utmost sophistication, yet all behaviorists could see in it just another type of behavior (called verbal behavior) that could be accounted for in (the same) terms of stimulusresponse.We can now parallel the two phenomena described so far: on the one hand, psychologists trying to formulate a psychological theory without any reference to the central concept of mind, and, on the other hand, linguists trying to formulate a language theory without any reference to the central concept of meaning. Since meaning lies in the mind of the speaker, everything becomes clearer if we detect the fundamental idea that overwhelmingly influenced the scientific thinking of that time: Darwin's hypothesis on evolution. …
The present study examined factors relating to momentary emotional response (i.e., arousal and valence ratings of affective stimuli) in 38 individuals with schizophrenia (SZ) and 53 healthy individuals (HC). Participants completed an affective picture stimuli ratings task to assess in-the-moment positive, negative, and neutral emotional experiences, and the factors examined were social (human) content and symptom presentation. \n\nThe results indicated that slides with social content were rated as more arousing and more negative than slides without social content, by all participants, but that SZ rated slides as more arousing than HC. Further, within negative stimuli, HC were more aroused by social than nonsocial slides, whereas SZ displayed the opposite pattern. Also within negative stimuli, SZ with high negative symptoms (SHNS) rated social stimuli as less arousing than SZ with low negative symptoms (SLNS), but SHNS and SLNS did not differ in their arousal ratings of nonsocial stimuli.\n\nThese findings exhibit a heightened level of arousal in schizophrenia, which is dampened by negative symptom severity, but only for negative slides with social content. This suggests that social content and symptom presentation may indeed play a role in momentary emotional experiences in schizophrenia.
Multisensory integration depends on the temporal proximity of events in different modalities. Recent studies have shown that multisensory temporal binding may be related to individual traits (Foss-Feig et al., 2010; Stevenson et al., 2012). Here we show that positive moods in observers enhance the temporal binding of audiovisual multisensory integration. Twenty-five healthy participants observed two identical visual disks moving toward each other, coinciding, and moving away. The two disks were perceived as either streaming through or bouncing off each other (stream/bounce display), and a belief sound around the visual coincidence facilitated bouncing perception (Sekuler et al., 1997; Watanabe and Shimojo, 2001). We asked the participants to report whether the two disks appeared to stream through or bounce off while listening to either exhilarating music of their own choice or a neutral pink noise. The results showed that the participants listening to exhilarating music reported bouncing percept more frequently. The proportion of bouncing percepts was correlated with the valence rating rather than the arousal rating during the experiment. These results suggest that positive moods enhance the temporal binding process in audiovisual integration.
This qualitative study examined the psychological and emotional effects of contemporary western society's standards of beauty on college-age African American women. Various studies on perception of beauty have explored body image perception from a middle class Caucasian perspective; and as a result, body image conceptualization from the perspective of African American women has gone unnoticed (Spurgas, 2005, Hatcher, 2007, Hall 1995). Previous literature suggests that African American women define self-perception and beauty differently from mainstream definitions. This study involved conducting a focus group with eighteen students between the ages of 18 and 25 who were enrolled at a Historically Black College in the Southeast. Students were asked a series of demographic background questions followed by a mixture of unstructured, open-ended questions that explored their perception of beauty standards and how such beauty standards have made an impact throughout their lives. In addition, the body image rating scale was given to each participate to rate based on their desirability. Major findings from this present study suggest that college-age African American women continue to be challenged by contemporary American standards of beauty in various ways. These findings also reveal how the family, media, and society's idealized beauty standards have a great influence in a woman's overall self-esteem and body image.
Due to the explosive growth of the Web, the domain of Web personalization has gained great momentum both in the research and commercial areas. One of the most popular web personalization systems is recommender systems. In recommender systems choosing user information that can be used to profile users is very crucial for user profiling. In Web 2.0, one facility that can help users organize Web resources of their interest is user tagging systems. Exploring user tagging behavior provides a promising way for understanding users' information needs since tags are given directly by users. However, free and relatively uncontrolled vocabulary makes the user self-defined tags lack of standardization and semantic ambiguity. Also, the relationships among tags need to be explored since there are rich relationships among tags which could provide valuable information for us to better understand users. In this paper, we propose a novel approach for learning tag ontology based on the widely used lexical database WordNet for capturing the semantics and the structural relationships of tags. We present personalization strategies to disambiguate the semantics of tags by combining the opinion of WordNet lexicographers and users' tagging behavior together. To personalize further, clustering of users is performed to generate a more accurate ontology for a particular group of users. In order to evaluate the usefulness of the tag ontology, we use the tag ontology in a pilot tag recommendation experiment for improving the recommendation performance by exploiting the semantic information in the tag ontology. The initial result shows that the personalized information has improved the accuracy of the tag recommendation.
The Slovene language is often presented as a national element. Even in the 19th century, which saw the Spring of Nations and the United Slovenia project, the Slovene language was a constitutive element of the Slovene nation. In the meantime, the Slovene language was positioning itself as an all-Slovene language, trying to be supra-regional. By the end of the 19th and early 20th centuries, the Slovene written language had stabilized, while at the same time the spoken language had only begun to assert itself. During this time, the prevailing principle was to "speak the way the language is written." In the mid-20th century, the theoretical idea of a literary language that is based on the central Slovene-speech (i.e. the speech of Ljubljana) came to dominate. In the third millennium, the question is whether a regionally-defined speech can be used as the basis for a Standard language. Another central question is what this "suitable" regionally-conditioned speech would be like. The principle of how important, decision-wise, the centre of a nation is, when it comes to questions of linguistic norms, may seem very attractive and, to a certain extent, logical. However, even examples of historically and linguistically comparable languages do not support the theory of creating the norm for the Standard Slovene language, based on the contemporary speech of Ljubljana, as claimed by Toporišič in Slovenska slovnica and, later, in Slovenski pravopis. Within Slovenia, the Standard Slovene language is tied to written language, which has proven, in the past, to be a suitable way of setting the norm. Regressing back to the principles of standardising a language, based on regional variants, would be unproductive, would introduce needless discord, and would cause problems with everyday, public communication. Contemporary research of actual speech, a portion of which is also presented within this article, confirms the all-Slovene and regionally-independent character of the Slovene Standard language.
This article explains why XML format has become established as the standard format for multilevel hierarchical structuring of linguistic databases and how an XML Schema can be used to manage the formal structure and content of elements in a dictionary database. Various aspects that must be taken into account when structuring complex dictionary databases in XML format are presented: the lexicographic or content aspect, the practical aspect, and the technical aspect. Decision-making is illustrated with the example of designing an XML Schema for the Dictionary of Slovenian Synonyms.
The article demonstrates how generic parsers in a minimally supervised information extraction framework can be adapted to a given task and domain for relation extraction (RE). For the experiments, two parsers that deliver n-best readings are included: (1) a generic deep-linguistic parser (PET) with a largely hand-crafted head-driven phrase structure grammar for English (ERG); (2) a generic statistical parser (Stanford Parser) trained on the Penn Treebank. It will be shown how the estimated confidence of RE rules learned from the n-best parses can be exploited for parse reranking for both parsers. The acquired reranking model improves the performance of RE in both training and test phases with the new first parses. The obtained significant boost of recall does not come from an overall gain in parsing performance but from an application-driven selection of parses that are best suited for the RE task. Since the readings best suited for the successful extraction of rules and instances are often not the readings favoured by a regular parser evaluation, generic parsing accuracy actually decreases. The novel method for task-specific parse reranking does not require any annotated data beyond the semantic seed, which is needed anyway for the RE task.
Abstract Emine Sevgi Özdamar is a representative author of German-Turkishliterature, which has been growing in size and importance during the last decades.The present article assumes that authors who use norm-deviant language(diverging from established linguistic practices as well as from conventional andmodern literary forms of expression) are particularly capable of making innovativecontributions to contemporary German literature. In Özdamar’s narrativeworks uncertain linguistic attitudes and deficient or “odd” verbalizations canproduce aesthetically valuable estrangement effects –; a conception recallingBertolt Brecht’s literary theory, as well as Ernst Jandl’s experimentations withwhat he called “heruntergekommene Sprache” (degenerate language). Due,among other factors, to Özdamar’s decision to use predominantly autobiographicalmaterial for her narrations, her exophonic writing accompanies in someworks childlike-naive forms of perception, and diction that can produce similareffects in respect to linguistic norms and habits. We can observe numerousmetaphorical expressions that reveal a playful animistic attitude towards reality.A look at Özdamar’s three-volume autobiographical sequence suggests that theestranging/innovative quality of her literary work decreases with the her gradualintegration into the German linguistic and literary community. Viewed in thislight, the struggle for literaricity would consist of maintaining and renewing theessential “linguistic strangeness” (Marcel Proust) of her style. Such an attitude isdifferent from a globalized (or globalizing) type of literature, which intends toproduce universal linguistic and narrative schemata.