Papers reviewed and determined not to be word norm studies. Use the flag icon to report errors or suggest re-inclusion.
18265 papers
In this paper, we present a document visualization technique for data analysis based on the semantic representation of text in the form of a directed graph, referred to as semantic graph. It is derived using natural language processing as follows. Firstly subject– verb – object triplets are automatically extracted from the Penn Treebank parse tree obtained for each sentence in the document. Secondly, the triplets are further enhanced by linking them to their corresponding co-referenced named entity, by resolving pronominal anaphors as well as attaching the associated WordNet synset. Starting from the document's semantic graph and the list of extracted triplets we automatically generate the document summary, for which we also derive the semantic representation.
The existence of graded structure in fruit and flower odour categories and its stability in different cultures is examined. Groups of students from France, the United States, and Vietnam performed a typicality rating task, a similarity judgment task, a membership verification task, a recognition memory task, a familiarity rating task, and a free identification task using a set of 40 odorants (20 fruit odorants and 20 flower odorants). Overall, our results demonstrate that fruit and flower odour categories possess graded structure. Moreover, principal component analyses of the data revealed the implication of typicality in a variety of cognitive tasks where typical odours receive a preferential processing compared to atypical ones. Finally, our results suggest that typicality can be predicted to a certain extent by experiential knowledge but that other determinants play a role in odour category structure. Altogether, this study confirms that graded structure is a universal property of categories and suggests that universals and cultural specifics can both constrain the emergence of odour category structures.
Abstract We investigated how viewing positive and negative emotional stimuli influences the functional field of view. Two types of emotional pictures were used, representing either a negative emotion such as disgust or fear, or a positive emotion such vigour or excitement. The participants' task was to detect and identify a digit presented in any of the four corners of a picture on the display, while discriminating a letter in the centre of the display. We used two stimulus onset asynchronies (SOAs), of 500 and 3000 ms, between the picture and the digit. Performance was poorer in the negative condition than in the positive and non-emotional conditions. Performance in the 500 ms SOA condition was poorer than in the 3000 ms SOA condition. These results suggest that the functional field of view (FFOV) becomes narrower when people view negative emotional stimuli, whereas it does not change when viewing positive or neutral emotional stimuli. Keywords: Emotional stimuliThe functional field of viewSOA Acknowledgements This work was supported by a grant to the first author from the Research Fellowships of the Japan Society for the Promotion of Science for Young Scientists and Grant no. 15330156 from the Japan Society for the Promotion of the Science to the second author. The authors thank Sachio Nakamizo for his helpful comments on an earlier draft of this manuscript. Notes 1The IAPS slide numbers were as follows: positive, 4608, 4660, 5470, 8030, 8170, 8179, 8185, 8186, 8200, and 8490; negative, 1050, 1525, 2811, 3215, 3500, 3550, 6021, 6022, 6300, 6312, 6350, and 9921; neutral, 2102, 2393, 2396, 2579, 2850, 5530, 5731, 7034, 7179, 7205, 7490, and 7710. An ANOVA was performed on the pleasant ratings, arousal ratings, and the number of bytes of the compressed image file sizes with one factor of emotion (positive, negative, and neutral). In the pleasant ratings, the main effect of emotion was significant, F(2, 33) = 560.63, MSE=0.13, p<.001. A post hoc multiple comparison test using Ryan's method indicated that the positive picture (M=7.28) was higher than the neutral (M=5.25) or the negative (M=2.54) picture (p<.05), and the neutral picture was higher than the negative picture (p<.05). In the arousal ratings, the main effect of emotion was significant, F(2, 33) = 257.17, MSE=0.19, p<.001. A post hoc multiple-comparison test using Ryan's method indicated that the positive (M=6.64) or the negative picture (M=6.47) was higher than the neutral (M=3.04) picture (p<.05). In the number of bytes of the compressed image file sizes, a main effect of emotion was significant, F(2, 33) = 8.74, MSE=3641.56, p<.001. A post hoc multiple comparison test using Ryan's method indicated that the negative picture (M=112.83) was lower than the positive (M=187.00) or the neutral (M=211.33) pictures (p<.05).
We present a cost effective strategy for the creation of a mid-size fine-grained dependency treebank of surface- and deep-syntactic structures as defined in the Meaning-Text Theory for Spanish. The strategy starts from a small seed dependency corpus, the AnCora corpus, whose annotation is considerably more coarse-grained than our target annotation. We show that this discrepancy can be bridged largely by automatic means, relying upon contextual information and leaving thus minimal work to the annotators. This allows us to develop the resources with limited human effort within a limited period of time. We also propose a preliminary evaluation of the actual amount of work that the annotation process requires. 1
German genitive attributes are usually tagged as such in treebanks. However, it is well known that this information is not sufficient for determining the type of relation between head nouns and attributes, as genitive attributes can express many different semantic relations. Various linguistic classifications have been worked out, but to my knowledge, nobody has so far proposed to apply this linguistic knowledge to a corpus. The challenge here is to come up with a classification that is both easy to verify and sufficiently fine-grained. Using earlier linguistic approaches as guidelines, I propose in this paper a detailed annotation scheme for German genitive attributes based on readily identifiable noun features. First insights from its application to the Smultron Treebank show that it is easy to distinguish between the proposed classes and that my classification of genitive attributes can be related to a more general semantic annotation level.
We present a semi-supervised method to improve statistical parsing performance. We focus on the well-known problem of lexical data sparseness and present experiments of word clustering prior to parsing. We use a combination of lexicon-aided morphological clustering that preserves tagging ambiguity, and unsupervised word clustering, trained on a large unannotated corpus. We apply these clusterings to the French Treebank, and we train a parser with the PCFG-LA unlexicalized algorithm of (Petrov et al., 2006). We find a gain in French parsing performance: from a baseline of F1=86.76% to F1=87.37% using morphological clustering, and up to F1=88.29% using further unsupervised clustering. This is the best known score for French probabilistic parsing. These preliminary results are encouraging for statistically parsing morphologically rich languages, and languages with small amount of annotated data.
This paper describes the MulTra project, aiming at the development of an efficient multilingual translation technology based on an abstract and generic linguistic model as well as on object-oriented software design. In particular, we will address the issue of the rapid growth both of the transfer modules and of the bilingual databases. For the latter, we will show that a significant part of bilingual lexical databases can be derived automatically through transitivity, with corpus validation.
One of the biggest challenges in compiling a dictionary of a minority language is managing the large quantity of lexical data. Decisions about the format and content of the dictionary or the orthography typically evolve over the years that such projects usually take. This results in inconsistencies between older and newer entries. Revising the data for publication as a dictionary introduces further inconsistencies as does having multiple contributors and/or editors. Proofreading a lexical database takes a great deal of time and the richer its structure the more this is the case. The tools described in this presentation significantly reduce this effort. Tools developed for checking the consistency of the lexical database in the Iu Mien—Chinese—English dictionary project have proven extremely helpful. Two basic approaches are used: 1) use of a program written to check for likely errors that scans the lexical database and produces an error report that is used by a lexicographer to make appropriate corrections. 2) outputting the lexical data in alternate forms that make it easier for the lexicographer to spot problem areas. These alternative forms include the reverse indexes and views structured according to semantic domains. The Iu Mien—Chinese—English dictionary project, like many minority language dictionary projects, uses SIL's Toolbox software. It is very flexible software but its capabilities to enforce consistency are quite limited. Some parts of the approach described here are specific to MDF (Multi-Dictionary Formatter) lexical databases in Toolbox but will be equally useful for other MDF databases. Other parts are specific to each of the three languages involved but will be useful for non-Toolbox lexical databases. Every dictionary is unique and this applies not only to content of the entries but also the decisions about how entries should be arranged to suit the languages involved. Other decisions about the structure are likely to be made differently even in other dictionaries of the same languages. It is the way that each dictionary combines themes that are found in many dictionaries that makes them unique, e.g. to be root based or not, to have include subentries. Therefore our approach is to use a toolkit based approach to curating lexical databases. This allows checking techniques to be mixed and matched to suit the unique aspects of a lexical project. The checking software is written in Python and relies on the toolbox module in NLTK (The Natural Language Toolkit http://nltk.sourceforge.net).
This paper presents a structural statistical machine translation (SSMT) model to deal with the data sparseness problem that occurs as a result of the necessarily small corpus to translate Chinese into Taiwanese Sign Language (TSL). A parallel bilingual corpus was developed, and linguistic information from the Sinica Treebank is adopted for Chinese sentence analysis. The synchronous context free grammar (SCFG) was adopted to convert a Chinese structure to the corresponding TSL structure and then extract a translation memory which comprises the thematic relations between the grammar rules of both structures. In structural translation, the statistical MT (SMT) approach was used to align the thematic roles in the grammar rules and the translation memory provides the reference templates for TSL structure translation. Finally, the agreement information for TSL verbs was labeled for enriching the expressiveness of the translated TSL sequence. Several experiments were conducted to evaluate the translation performance and the communication effectiveness for the deaf. The evaluation results demonstrate that the proposed approach outperforms a baseline statistical MT system using the same small corpus, especially for the translation of long sentences.
This paper describes log-linear models for a general-purpose sentence realizer based on dependency structures. Unlike traditional realizers using grammar rules, our method realizes sentences by linearizing dependency relations directly in two steps. First, the relative order between head and each dependent is determined by their dependency relation. Then the best linearizations compatible with the relative order are selected by log-linear models. The log-linear models incorporate three types of feature functions, including dependency relations, surface words and headwords. Our approach to sentence realization provides simplicity, efficiency and competitive accuracy. Trained on 8,975 dependency structures of a Chinese Dependency Treebank, the realizer achieves a BLEU score of 0.8874.
Automatic syllabification of words is challenging, not least because the syllable is not easy to define precisely. Consequently, no accepted standard algorithm for automatic syllabification exists. There are two broad approaches: rule-based and data-driven. The rule-based method effectively embodies some theoretical position regarding the syllable, whereas the data-driven paradigm tries to infer "new" syllabifications from examples assumed to be correctly syllabified already. This article compares the performance of several variants of the two basic approaches. Given the problems of definition, it is difficult to determine a correct syllabification in all cases and so to establish the quality of the "gold standard" corpus used either to evaluate quantitatively the output of an automatic algorithm or as the example-set on which data-driven methods crucially depend. Thus, we look for consensus in the entries in multiple lexical databases of pre-syllabified words. In this work, we have used two independent lexicons, and extracted from them the same 18,016 words with their corresponding (possibly different) syllabifications. We have also created a third lexicon corresponding to the 13,594 words that share the same syllabifications in these two sources. As well as two rule-based approaches (Hammond's and Fisher's implementation of Kahn's), three data-driven techniques are evaluated: a look-up procedure, an exemplar-based generalization technique, and syllabification by analogy (SbA). The results on the three databases show consistent and robust patterns. First, the data-driven techniques outperform the rule-based systems in word and juncture accuracies by a very significant margin but require training data and are slower. Second, syllabification in the pronunciation domain is easier than in the spelling domain. Finally, best results are consistently obtained with SbA.
The systematic presentation of collocations is increasingly recognized as a very useful addition to specialized reference works. However, few dictionaries or terminological databases actually include this kind of data. More surprisingly still, no method has been designed yet to allow efficient access to and retrieval of specific specialized collocations from electronic reference tools. This article presents two new search paths for accessing and extracting collocations from an English-French specialized lexical database. The paths have been designed according to two specific user-defined situations: (1) translation from L1 to L2; and (2) text production in L2. We exploit a formal semantic encoding of collocations based on Lexical Functions (LFs). LFs allow us to establish an equivalence relationship between collocations that convey the same meaning in different languages without having to link the collocations formally. They also allow us to extract sets of collocations associated with specific meanings.
Within the STEVIN1 project Large Scale Syntactic Annotation of written Dutch (LASSY), a manually corrected treebank of 1 million words is constructed. Lassy is part of a series of annotation projects for modern written and spoken Dutch. More specifically, it is an extension of the D-Coi and CGN projects,2 and constitutes the core of SoNaR, a 500 million words reference corpus of modern written Dutch.3 One of the goals of the latter project is to enrich the corrected treebank produced in Lassy4 with several semantic layers. For a general overview of the relations between D-Coi, Lassy and SoNaR, cf [19]. In this paper we will concentrate on the semantic layers of SoNaR core: (1) named entity labeling, (2) annotation of co-reference relations, (3) semantic role labeling and (4) annotation of spatial and temporal relations. Of these (2) originates from the STEVIN-project COREA,5 (3) and (4) from D-Coi, whereas (1) is a new area within STEVIN.
We adapt a semantic role parser to the domain of goal-directed speech by creating an artificial treebank from an existing text tree-bank. We use a three-component model that includes distributional models from both target and source domains. We show that we improve the parser's performance on utterances collected from human-machine dialogues by training on the artificially created data without loss of performance on the text treebank.
Collocations constitute a subclass of multi-word expressions that are particularly problematic for machine translation, due 1) to their omnipresence in texts, and 2) to their morpho-syntactic properties, allowing virtually unlimited variation and leading to long-distance dependencies.Since existing MT systems incorporate mostly local information, these are arguably ill-suited for handling those collocations whose items are not found in close proximity.In this article, we describe an integrated environment in which collocations (and possibly their translation equivalents) are first identified from text corpora and stored in the lexical database of a translation system, then they are employed by this system, which is capable of dealing with syntactic transformations as it is based on a deep linguistic approach.We compare the performance of our system (in terms of collocation translation adequacy) with that of two major MT systems, one statistical, and the other rule-based.Our results confirm that syntactic variation affects translation quality and show that a deep syntactic approach is more robust in this sense, especially for languages with freer word order (e.g., German) and richer morphology (e.g., Italian) than English.
We present an approach for smoothing treebank-PCFG lexicons by interpolating treebank lexical parameter estimates with estimates obtained from unannotated data via the Inside-outside algorithm. The PCFG has complex lexical categories, making relative-frequency estimates from a treebank very sparse. This kind of smoothing for complex lexical categories results in improved parsing performance, with a particular advantage in identifying obligatory arguments subcategorized by verbs unseen in the treebank.
The present study compared gender differences in directly reported and indirectly derived career preferences and tested the hypothesis that individuals' implicit preferences would show less gender-biased occupational choices than their directly elicited ones. Two hundred sixty-six visitors to a career-related Internet site were asked to (a) list 5 to 10 suitable occupations (the directly reported list) and (b) report their preferences in terms of 31 career-related aspects. The latter were used to produce a short list of promising occupational alternatives (the indirectly derived list), using the occupational database of an Internet-based career planning system. Each occupation in the database rated for sex dominance. The findings indicated that the sex dominance ratings of the occupations on the directly reported list accorded with the participants' gender for both men and women: Men's lists included mostly “masculine” occupations, whereas women's lists included mostly “feminine” occupations. This gender bias was significantly lower for the implicit lists. The difference between the directly reported and the indirectly derived lists was larger for women than for men, suggesting that the impact of stereotypes is more pronounced in women's than in men's directly reported career preferences.
BACKGROUND: Memory impairment and verbal learning are the most common cognitive deficits associated with schizophrenia. Hopkins Verbal Learning Test (HVLT) is considered to be the most reliable test to asses memory and verbal learning in this mental illness. AIMS: to create one form of the HVLT which would suit our linguistic and cultural context and to study the characteristics of this test in a group of healthy subjects. METHODS: The HVLT consists of a list of 12 words belonging to 3 semantic categories and which are read orally to the subject with an immediate and differed recall. The first part of this work was to select words from a lexical database in order to create the list of the HVLT. The test was then administered to 103 subjects aged from 17- to 45-years-old (mean=27,4; SD =7,3) and having between 1 and 20 years of education ( mean=12,2; SD=5,3). RESULTS: No statistical difference was found within performances of the HVLT across gender and sex. Whereas, years of education was found to have an impact on performances. Although statistically difference was found across level of education. CONCLUSION: Our study permitted us to create one form of the HVLT which well suits our Tunisian context and which we could use to evaluate memory functions among people suffering from schizophrenia.
Abstract. Finding information about companies on multiple sources on the Web has become increasingly important for business analysts. In particular, since the emergence of the Web 2.0, opinions about companies and their services or products need to be found and distilled in order to create an accurate picture of a business entity. Without appropriate text mining tools, company analysts would have to read hundreds of textual reports, newspaper articles, forums’ postings and manually dig out factual as well as subjective information. This paper describes a series of experiments to assess the value of a number of lexical, morpho-syntactic, and sentiment-based features derived from linguistic processing and from an existing lexical database for the classification of evaluative texts. The paper describes experiments carried out with two different web sources: one source contains positive and negative opinions while the other contains fine grain classifications in a 5-point qualitative scale. The results obtain are positive and in line with current research in the area. Our aim is to use the result of classification in a practical application that will combine factual and opinionated information in order to create the reputation of a business entity. 1
In this paper we present a question answering system supported by semantic graphs. Aside from providing answers to natural language questions, the system offers explanations for these answers via a visual representation of documents, their associated list of facts described by subject – verb – object triplets, and their summaries. The triplets, automatically extracted from the Penn Treebank parse tree obtained for each sentence in the document collection, can be searched, and we have implemented a question answering system to serve as a natural language interface to this search. The vocabulary of questions is general because it is not limited to a specific domain, however the questions's grammatical structure is restricted to a predetermined template because our system can understand only a limited number of question types. The answers are retrieved from the set of facts, and they are supported by sentences and their corresponding document. The document overview, comprising the semantic representation of the document generated in the form of a semantic graph, the list of facts it contains and its automatically derived summary, offers an explanation to each answer. The extracted triplets are further refined by assigning the corresponding co referenced named entity, by resolving pronominal anaphors, as well as attaching the associated WordNet synset. The semantic graph belonging to the document is developed based on the enhanced triplets while the document summary is automatically generated from the semantic description of the document and the extracted facts.
Impersonating Identity in Spaces of Difference is an ALTERnative discourse of dis/ruption, decolonization, deconstruction....Writing on the b/orders of theories, disciplines, genres, cultures....this re/search weaves together personal, familial, and societal stories of silence and silencings. By subverting conventional academic texts and hegemonic frameworks of Canadian "MultiCULTural institutions and Canadian "MultiCULTural society," this subaltern re/search claims a space for marginalized voices. Conventional ways of knowing, being, becoming are generatively disRupted in order to create awareness of the continuing legacies of colonialism, modernity, patriarchy...and to highlight the urgent need to provide genuine spaces of "belonging" that inclusively honour and respect the gifts of all individuals. Entangled in the in/visible hierarchical realities of "Canadianness," this in/quiry articulates decolonizing resistance. The layers of the impure academic textual body perform the multiple fragmented intertextual layers of the improper Indocanadian/ Can-indian body creating em(bodi)ed re-imag(e)inings of epistemology and pedagogy. This tense, restless landscape of multi-languages, multi-genres, multimeanings, multi-truths is an inter-Ruption of predetermined b/orders and predetermined bodies of predetermined purity, in a world of ever-changing multiplicity. Drawing on "difference" in multiCULTuralism, language, voice, and identity, this performative work travels in and out of questions of absence, "hybridity," foreignness, loss, displacement, marginalization, patriarchy, colonialism, modernity....In this powerful and liberating form of non-traditional in/quiry, meaning making takes precedence over conventional stylistic or pre-established structures and acknowledges personal "ethnic" experience as a valuable form of reliable "academic" knowledge. This trans-disciplinary, transformative, transcultural...in/quiry disRupts traditional hegemonic narratives and challenges the conventional notions of re/search and writing through its form and content. Through the braided weaving of English, French and Punjabi, personal stories, familial narratives, prose, letters, e-mails, cross cultural conversations, visual imagery, historical documents, "subtexts," "surtexts," intertexts, collaborative texts...undermine and upset hegemonic linguistic norms. Fiction, fantasy, History, herstory, theirstories, memory...are juxtaposed through mixed non-linear genres and codes in protest of violent acts of com(form)ity, exclusion and censorship. Stories of India and Canada find themselves interwoven unexpectedly, betraying the lies and the truths of patriarchy, colonialism, modernity, multiCULTuralism, transCULTuralisms...disrupting clean, linear readings of writing and of research. This experiment with nonstructure and typography attempts to actively decolonize, deconstruct/re-construct imposed academic and social identities in a more meaningful way and provide readers with a sense of living in the "transculturality" of the "diasporic in-between." Only by provoking a critical, cross-cultural "INNERstanding" of History, POWer, systemic marginalization, colonialism and modernity can we begin to relationally take on the collective response-abilty for social justice in policy and practice ~ in the word and in the wor(l)d.
This paper presents a comparative study of Judgment and Assessing frames in English and Portuguese. The aim is to verify the possibility of using the FrameNet frames to construct a lexical database for Brazilian Portuguese. The research corpus is composed by 50 legal documents, totalizing 1.055,535 tokens and 39,108 types. Through a contrastive method the Judgment and Assessing frames were selected and translation equivalents for the English lexical units were established. The points considered in this research were the polysemy and the semantic relations of words. The polysemy is the main difficulty in applying FrameNet frames for Portuguese description.
In the article we compare the role of the dictionary and the lexical database, and address the issue of language register and correctness in dictionaries. We then deal with various types of sense distribution in dictionaries, the history of the word, and the principles of selection of dictionary headwords. We cite the corpus as an essential source for the treatment of meaning, collocation and syntagmatics, and investigate ways of interpreting corpus data – corpus profiling of headwords. We conclude with the thought that a dictionary represents the central language standard, whereby all of the expressed linguistic opinions contained in it must be based on corpus evidence.
In this paper, we present our in-progress research tasks for building lexical database of the verb valences in the Arabic Quran using FrameNet frames. We study the verbs in their context in the Quran, and compare that with matching frames and frame evoking verbs in the English FrameNet. We analyze the gaps and make appropriate amendments to the FrameNet by adding new frame elements and relations.
Due to the lack of descriptive mechanism for syntactic functional structure,the dependency grammar can not express all complex syntactic structures explicitly.In addition,few parsing model takes the restrictive information of modifiers' nesting level into account,though it is a common sense in pragmatics.To solve these problems,a generative binary combinational grammar (BCG) parsing model is proposed which incorporates the restrictive information.In this model,the construction of sentence is regarded as the combination of adjacent chunks according to their headwords.Moreover,the symbolic local priorities between the adjacent binary relations and the modifiers' nesting levels are used to constrain the generation of parsing trees.The BCG parsing model is constructed by converting the dependency treebank to the BCG form.Then,the syntactic relations,the local priorities and the parameters of the model are induced automatically.Experimental results show that the proposed model improves the parsing accuracy.
A study of Guilln's critical work through a comparative analysis of Lenguaje y poesa essays concerning the mystic poetry of saint John of the Cross and the visionary lyrics of G.A. Bcquer.The obscure difficulties of these poets challenge Guilln's vision of literature as a conscious art of language.Symbolic interpretation requires the use of different strategies, basically a philological approach and the methods of several new criticisms stemming from the Romantic hermeneutics, concerning problems such as the alternance of litteral and spiritual senses, the poem as an imaginative transgression of linguistic norm, or the shifts between textual and intentional messages.Guilln's answer to the question of unintelligibility leads to a consideration of the way in which texts are determined and the borders between poetry and literary criticism.
This paper proposes a novel query expansion method using Local Context Analysis(LCA) based concept tree pruning.It extracts the suggested terms from initial documents which are retrieved for the original query by LCA method,and uses these terms to prune the concept tree built by lexical database,to add new terms,and to recalculate the weight of expanding terms.Experiments show that the combined expansion yields significant improvements over existing algorithms under the same experimental condition.
Lexicon-Grammar tables are a very rich syntactic lexicon for the French language. This linguistic database is nevertheless not suitable for use by computer programs, as it is incomplete and lacks consistency. Our goal is to adapt the tables, so as to make them usable in various Natural Language Processing (NLP) applications. We describe the problems we encountered and the approaches we followed to enable their integration into a parser.
This dissertation examines the relationship between gender and language in Japanese through the often ignored lens of sexuality. Although linguists are increasingly examining these issues for American gay, lesbian, and bisexual speakers, little similar research has been done in Japan. Lesbians, in particular, are relatively invisible in Japanese society. Examining these women, who do not fit neatly into the hegemonic gender ideology, illuminates how speakers can project a specific identity by displaying or rejecting prescriptive gender-specific linguistic norms of Japanese.I analyzed data recorded from interviews with both Japanese lesbian/bisexual and heterosexual women, looking for differences in frequency and range of use of pronouns and sentence-final particles and for phonetic differences in terms of average pitch height and width. I also considered the results of a perception experiment undertaken to investigate the effect of pitch height and width on Japanese speakers' perceptions of sexuality.Although Japanese speakers were generally unable to identify a cohesive lesbian stereotype, especially in terms of language use, the perception experiment indicated that both average pitch height and width significantly affect judgments on whether a voice sounds lesbian or heterosexual. Tokens judged to be lesbian were also judged to be more masculine and less emotional than those judged to be heterosexual. Analysis of the interview data showed that lesbian participants produced an average pitch height that was significantly lower than that of heterosexual participants. In terms of gendered morphemes, lesbians were significantly more likely to use masculine morphemes than heterosexual women, both for sentence-final particles and first-person pronouns, and were significantly less likely to use the feminine first-person pronoun atashi. Finally, correlations showed that speakers who instantiate gender through the use of gendered-morphemes also do so through manipulations of pitch.Although Japanese lesbians are still fairly closeted and interviewees maintained that there are no cultural stereotypes for this group, significant differences in pitch and gendered-morpheme usage were still apparent. These lesbian/bisexual women did not appear to be mimicking men's language, but instead seemed to be rejecting hegemonic femininity and many of the cultural and linguistic stereotypes that accompany it.
We present several algorithms for assigning heads in phrase structure trees, based on different linguistic intuitions on the role of heads in natural language syntax. Starting point of our approach is the observation that a head-annotated treebank defines a unique lexicalized tree substitution grammar. This allows us to go back and forth between the two representations, and define objective functions for the unsupervised learning of head assignments in terms of features of the implicit lexicalized tree grammars. We evaluate algorithms based on the match with gold standard head-annotations, and the comparative parsing accuracy of the lexicalized grammars they give rise to. On the first task, we approach the accuracy of hand-designed heuristics for English and inter-annotation-standard agreement for German. On the second task, the implied lexicalized grammars score 4% points higher on parsing accuracy than lexicalized grammars derived by commonly used heuristics.
In the decade of the sixties, the main concern of the experts in Basque language was how to unite the speech community. However, after forty years of Standard Basque, the objective is to develop a language that is well adapted to each register. The lexis is the key in the functional development that the adaptation to each register requires, and that is exactly what we shall study in this work. We shall analyse, from the point of view of functional diversity, the methodology of the Standard Basque Dictionary, the purpose of which is to organize the standard Basque lexicon. We reach the conclusion that Standard Basque has taken certain steps to get closer to the functional variability desideratum. However, we have also detected methodological practices that may be an obstacle in this functional development. As a result, we draw up the outline of the methodological changes that would be required to accommodate linguistic norms to the variation at the register level.
In his later work in the philosophy of language Davidson analyses communicative exchanges and arrives at the startling and weighty conclusion that linguistic norms and conventions are entirely inessential to linguistic meaning. I argue that this inference is flawed: if we place the account of communication against the backdrop of Davidson's own views about radical interpretation, then it becomes evident that linguistic norms are an essential feature of the Davidsonian picture.
Sciendo provides publishing services and solutions to academic and professional organizations and individual authors. We publish journals, books, conference proceedings and a variety of other publications.
We present in this paper an approach to assessing student paraphrases in the intelligent tutoring system iSTART. The approach is based on measuring the semantic similarity between a student paraphrase and a reference text, called the textbase. The semantic similarity is estimated using knowledge-based word relatedness measures. The relatedness measures rely on knowledge encoded in Word-Net, a lexical database of English. We also experiment with weighting words based on their importance. The word importance information was derived from an analysis of word distributions in 2,225,726 documents from Wikipedia. Performance is reported for 12 different models which resulted from combining 3 different relatedness measures, 2 word sense disambiguation methods, and 2 word-weighting schemes. Furthermore, comparisons are made to other approaches such as Latent Semantic Analysis and the Entailer.
It is a commonplace, by now, to refer to the recent explosive growth in the power and availability of computers as an information revolution. The most casual of computer users, linguists included, have at their fingertips an enormous amount of computing power. Tasks such as writing a document or playing
This paper reports on-going work on building a large automatically tree-aligned parallel treebank in the context of a syntax-based machine translation (MT) approach. For this we develop a discriminative tree aligner based on a log-linear model with a rich feature set. We incorporate various language-independent and language-specific features taking advantage of existing tools and annotation. Our initial experiments on a small hand-aligned treebank show promising results even with small amounts of training data. The performance of our approach is well above unsupervised techniques reported elsewhere. This enables us to quickly create training material and alignment models for additional language pairs. In recent work, we aligned more than one million sentence pairs and started our experiments with the extraction of transfer knowledge for our example-based machine translation system.
The Arabic language has a very rich/complex morphology. Each Arabic word is composed of zero or more prefixes, one stem and zero or more suffixes. Consequently, the Arabic data is sparse compared to other languages such as English, and it is necessary to conduct word segmentation before any natural language processing task. Therefore, the word-segmentation step is worth a deeper study since it is a preprocessing step which shall have a significant impact on all the steps coming afterward. In this article, we present an Arabic mention detection system that has very competitive results in the recent Automatic Content Extraction (ACE) evaluation campaign. We investigate the impact of different segmentation schemes on Arabic mention detection systems and we show how these systems may benefit from more than one segmentation scheme. We report the performance of several mention detection models using different kinds of possible and known segmentation schemes for Arabic text: punctuation separation, Arabic Treebank, and morphological and character-level segmentations. We show that the combination of competitive segmentation styles leads to a better performance. Results indicate a statistically significant improvement when Arabic Treebank and morphological segmentations are combined.
A novel series of trifluoromethyl-containing quinazoline derivatives with a variety of functional groups was designed, synthesized, and tested for their antitumor activity by following a pharmacophore hybridization strategy. Most of the 20 compounds displayed moderate to excellent antiproliferative activity against five different cell lines (PC3, LNCaP, K562, HeLa, and A549). After three rounds of screening and structural optimization, compound 10 b was identified as the most potent one, with IC<sub>50</sub> values of 3.02, 3.45, and 3.98 μM against PC3, LNCaP, and K562 cells, respectively, which were comparable to the effect of the positive control gefitinib. To further explore the mechanism of action of 10 b against cancer, experiments focusing on apoptosis induction, cell cycle arrest, and cell migration assay were conducted. The results showed that 10 b was able to induce apoptosis and prevent tumor cell migration, but had no effect on the cell cycle of tumor cells.
Dolgozatomban, amint arra a címből is lehet következtetni, az 1996 és 2005 között adatolható magyarországi börtönszlenget mutatom be, azt a csoportnyelvet, amelynek az átfogó tanulmányozása hazánkban eddig még nem történt meg. Munkámban büntetés-végrehajtási intézeteink fogvatartottjainak belső, informális nyelvhasználatának általános kérdéseivel foglalkozom, és a mai magyar börtönszleng szó- és kifejezéskészletének szótárba foglalásán túl kísérletet teszek a vizsgált csoportnyelv nyelvi-szociolingvisztikai leírására. \n \nCélkitűzésemet, a magyar börtönszleng átfogó tanulmányozását az indokolta, hogy a kutatás első éveiben olyan mennyiségű és minőségű, a nyelvtudományban eddig még nem tárgyalt adatokra bukkantam, amelyek érdemesnek mutatkoztak arra, hogy egy mélyebb, megtervezett szlengkutatás irányuljon erre a területre. \n \nElsősorban célom volt a magyar börtöszlenget feltárni, bemutatni keletkezését, funkcióját, működését, a szlenghasználó közösségben betöltött szerepét. Célom volt továbbá rámutatni nyelvi előzményeire, összevetni a már létező bűnözői nyelvi adatbázissal, vagyis a tolvajnyelv elemeivel, egyben definiálni helyét a magyar szlengkutatás területén. Kutatásom során mindvégig azt tartottam szem előtt, hogy hol van az ember a szlengben, így célom volt annak leírása is, mikor, milyen körülmények között motiváltak a vizsgált csoport tagjai szlenghasználatra. Ennek kiderítéséhez a nyelvi adatok feltárásán és rendszerezésén túl, a zárt közeg csoportjainak vizsgálatára is ki kellett terjeszteni a kutatást, megfigyelve a csoportszerveződés lehetőségeit és okait a börtöntársadalomban. \n \nAs already the title has suggested, my dissertation presents Hungarian prison slang as attested from 1996 to 2005. This group language has never seen an overall study in Hungary up to now. I deal with general questions of the internal and informal language use of the prisoners of penal institutions in my work and, together with rendering the words and expressions into a dictionary, I attempt at the linguistic and sociolinguistic description of the group language examined. \n \nMy objective, the overall study of Hungarian slang, is justified by the data being of such quantity and quality and having never been dealt with in linguistics that seemed worth to be examined by a deeper and planned slang research. \n \nMy main objective was to explore Hungarian prison slang, to present its origins, functions and operation as well as its role within slang user communities. My aim also was to show its linguistic predecessors, to compare it to an existing linguistic database of criminals, that is, to the elements of cant, and, together with it, to define its place in the area of Hungarian slang research. During my studies, I always kept man’s role in slang in mind so my objective was the description of the conditions among which the members of a researched group are motivated for slang usage. In order to learn about it, I had to extend my research to the examination of the groups of closed space, observing the possibilities and reasons within prison community.
In this paper, we propose a modular cascaded approach to data driven dependency parsing. Each module or layer leading to the complete parse produces a linguistically valid partial parse. We do this by introducing an artificial root node in the dependency structure of a sentence and by catering to distinct dependency label sets that reflect the function of the set internal labels vis-a¿-vis a distinct and identifiable linguistic unit, at different layers. The linguistic unit in our approach is a clause. Output (partial parse) from each layer can be accessed independently. We applied this approach to Hindi, a morphologically rich free word order language using MST parser. We did all our experiments on a part of Hyderabad Dependency Treebank. The final results show an increase of 1.35% in unlabeled attachment and 1.36% in labeled attachment accuracies over state-of-the-art data driven Hindi parser.
We study the influence that image features may have on music tension and liveliness perception. 72 music excerpts from different genres and periods were selected, and 72 still shots were taken from different animation features little known to the subjects. 62 subjects rated the isolated images for tension and liveliness, 37 subjects rated the isolated music excerpts for tension and liveliness, and 153 subjects rated the music excerpts combined with the images for music tension and liveliness, and for music-image congruence. There is a significant variation of tension and liveliness of the music as a function of the tension and liveliness of the pairing image, showing a transfer of mood from image to music. The significance of ANOVA tests showed that 40% of music excerpts were image-sensitive for liveliness and 32% for tension. The transfer of mood was dependent on congruence: music excerpts with high congruence with the image had a higher correlation in tension and liveliness rating deviations with the image ratings. For low congruence, the liveliness correlation was not significant and the tension deviation was negatively correlated with the image tension. Feature transfer from image to music depends on the image-music congruence rated by each subject.
Query expansion is a widely studied technique for improving information retrieval effectiveness. In this paper we proposed a new query expansion technique using the comprehensive thesaurus WordNet and its semantic relatedness measure modules. Word sense disambiguation are performed on original query sentence, yielding the concept of each term in the query. Based on those recovered concepts, expanded query terms are generated from WordNet lexical database. The proposed method has been evaluated in document retrieval on the Web using query sentence. Our extensive experimental results demonstrate a 7% precision improvement over retrieval methods not employing query expansion techniques.
In recent years, the specter of litigants turning to religious or customary sources of law as authoritative guides to regulate their behavior, alongside or in lieu of secular norms, has risen to the forefront of politics in many countries worldwide. In this essay, we draw upon citizenship theory and comparative constitutional jurisprudence to identify two different categories of judicial response to religious-based claims for recognition, accommodation, and exemption: 1) 'diversity as inclusion;' and 2) 'non-state law as competition.' As long as legal claims for accommodation are not seen by courts as challenging the lexical superiority of the constitutional religion itself ('diversity as inclusion'), they stand a fair chance of success. Contrast that with the unyielding reluctance of legislatures and judiciaries to accept as binding or even cognizable any potentially competing legal order that originates in sacred or customary sources of identity and authority. This pattern of clamping down and refusing to accept any alternative sources of regulation becomes particularly visible where the legal challenge at issue is interpreted as raising doubts regarding which set of norms and institutions, or what set of high priests, should have the final word in authoritatively resolving legal disputes within a given society ('non-state law as competition'). This is a challenge that no secular legal order, no matter how tolerant and otherwise open to providing exemptions and accommodations to religious believers, can accept with indifference. For what perceived to be at stake here is the very authority and source of legitimacy of the accepted civil religion. We demonstrate these claims by focusing on recent jurisprudence from Canada and South Africa, two polities that represent the most difficult cases for our argument; if there is any place we would expect to find recognition by secular countries of religious or customary sources of law and authority, it would be in these diverse societies that have made an explicit constitutional commitment to promote their citizens’ freedom to preserve and enhance their multitude of backgrounds and distinctive cultural, linguistic and religious heritages as part of their 'mosaic' (Canada) or 'rainbow nation' (South Africa) conceptions of citizenship. Although operating in different contexts, the South African Constitutional Court and the Supreme Court of Canada seem to have made every effort to subject traditional legal regimes to general principles of constitutional law. By so doing, they have erected a new wall of separation that places noncompliance with the values of the civil religion beyond the pale of accepted accommodation, offering to those who espouse them the potential to either bring these alternative legal domains under the general rule of constitutional law or encounter the wrath of state fiat.