Papers reviewed and determined not to be word norm studies. Use the flag icon to report errors or suggest re-inclusion.
18265 papers
‘Lexical bundles’ as a category of word combinations are words which follow each other more frequently than expected by chance. This corpus-based study attempts to compare the frequencies of three- and four-word lexical bundles in research articles of three disciplines: physics, computer engineering, and applied linguistics. Moreover, it aims to scrutinize them between native and nonnative research articles of applied linguistics to see whether Iranian authors who publish articles in English, use lexical bundles in the same way as native authors. To this end, three native corpora and a non-native corpus of research articles were collected, each including approximately one million words. All the analyses were conducted through Wordsmith Tools (Scott, 2010) and Hyland’s (2008) taxonomy of most frequent academic lexical bundles. The results show that there are relatively significant differences between the frequencies of the lexical bundles employed across the disciplines. In addition, they differ significantly between the native and nonnative articles of applied linguistics. It is also revealed that lexical bundles are realized differently across different disciplines and that non-natives do not follow the norms of natives appropriately. Findings can be used to improve writing in different disciplines and create more cohesive and coherent texts.
Grote verzamelingen van vertaalde teksten – zogenaamde parallelle corpora - worden vaak automatisch op zins- en woordniveau gealigneerd om automatische vertaalsystemen op te trainen. Soms voegt men ook automatisch syntactische bomen aan de zinnen toe om meer taalkundige informatie eruit te kunnen halen. Als die bomen aan beide kanten verschijnen en de boomknopen ook worden gealigneerd, is er sprake van een parallelle treebank. De beste vertaalsystemen zijn bijna of helemaal puur statistisch, maar in recente jaren ontstond er een grotere nadruk op de integratie van meer taalkundig gemotiveerde data, waaronder ook het gebruik van parallel treebanks. Ze zijn echter alleen op een zeer grote schaal bruikbaar, omdat er door zo een systeem veel te leren is van hoe een taal typisch naar een andere moet worden omgezet. Daarom onderzoeken we technieken om automatisch de boomknopen accuraat te aligneren. Een bijkomend motief is het feit dat parallel treebanks ook voor andere applicaties bruikbaar zijn en als taalbronnen zelf van wetenschappelijk belang zijn. Het hele proces van het aligneren van knopen noemen wij tree alignment. Wij vinden dat een combinatie van statistiche en regelgebaseerde technieken met relatief weinig trainingsgegevens en weinig features zeer accurate alignments kan produceren. Ten slotte vinden we dat, wanneer wij alignments die relatief heel veel knopen aligneren – al zijn sommigen soms fout – op een syntactisch gebaseerde systeem toepassen, dat tot verbeterde automatische vertaling leidt, in vergelijking met hetzelfde systeem die op minder maar meer accurate alignments getrained is.
In the field of constituency parsing, there exist multiple human-labeled treebanks which are built on non-overlapping text samples and follow different annotation standards. Due to the extreme cost of annotating parse trees by human, it is desirable to automatically convert one treebank (called source treebank) to the standard of another treebank (called target treebank) which we are interested in. Conversion results can be manually corrected to obtain higher-quality annotations or can be directly used as additional training data for building syntactic parsers. To perform automatic treebank conversion, we divide constituency parses into two separate levels: the part-of-speech (POS) and syntactic structure (bracketing structures and constituent labels), and conduct conversion on these two levels respectively with a feature-based approach. The basic idea of the approach is to encode original annotations in a source treebank as guide features during the conversion process. Experiments on two Chinese treebanks show that our approach can convert POS tags and syntactic structures with the accuracy of 96.6 and 84.8 %, respectively, which are the best reported results on this task.
This paper reports an effort to annotate modality in the Penn Chinese Treebank. We introduce the modals and features that were annotated, and describe the phases of our working process. Along with this, we address the issues in the preparation of annotation guidelines, and present the preliminary results of the first pass. Finally, we analyze the types of disagreement, and propose directions to improve consistency. 1
We present a novel method ("waste") for the segmentation of text into tokens and sentences. Our approach makes use of a Hidden Markov Model for the detection of segment boundaries. Model parameters can be estimated from pre-segmented text which is widely available in the form of treebanks or aligned multi-lingual corpora. We formally define the waste boundary detection model and evaluate the system's performance on corpora from various languages as well as a small corpus of computer-mediated communication.
The present paper focuses on ways in which the pragmatic (functional) meaning that arises from various contextual features, known in corpus linguistics as semantic prosody, can become an integral part of lexicographical descriptions as they are represented in the Slovene Lexical Database (SLD). This is particularly important for the treatment of phraseology and idiomatics. First, the theoretical background is provided, with the focus on the prototype theory and its practical implications for monolingual lexicography. A parallel is drawn with the model of meaning analysis in the SLD. The second part begins with a brief introduction to semantic prosody and continues with an analysis of monolingual meaning descriptions in the SLD against a number of authentic corpus examples, investigating how their pragmatic components have been identified. The analysis of corpus data shows that pragmatics is an important contributor to the process of sense discrimination in works of lexical and lexicographic relevance.
With the growing interest in statistical parsing, special attention has recently been devoted to the problem of comparing different treebanks to assess which languages or domains are more difficult to parse relative to a given model. A common methodology for comparing parsing difficulty across treebanks is based on the use of the standard labeled precision and recall measures. As an alternative, in this article we propose an information-theoretic measure, called the expected conditional cross-entropy (ECC). One important advantage with respect to standard performance measures is that ECC can be directly expressed as a function of the parameters of the model. We evaluate ECC across several treebanks for English, French, German, and Italian, and show that ECC is an effective measure of parsing difficulty, with an increase in ECC always accompanied by a degradation in parsing accuracy.
English periphrastic causative constructions, i.e. constructions where a causative verb like make or get controls a non-finite complement clause, have been the subject of many studies representing different theoretical frameworks, among which generative grammar (e.g. Kastovsky 1973), the universal-typological theory (e.g. Wierzbicka 1998), cognitive linguistics (e.g. Hollmann 2006) and construction grammar (e.g. Stefanowitsch 2001). Most of the time these studies have focused on the way periphrastic causative constructions are used (or should be used) by native speakers of English. Fewer studies have considered the use of these constructions by non-native speakers of English (cf. Ziegeler & Lee 2009, Gilquin 2012). In this presentation, I adopt a constructionist approach to investigate the use of periphrastic causative constructions in two non-native varieties of English, namely English as a Foreign Language (EFL) and English as a Second Language (ESL). While both of these varieties correspond to L2s that are acquired in addition to the L1, the settings of acquisition are different (mainly an instructional setting for EFL and mainly a natural one for ESL), which could lead to some differences in the way the causative construction behaves in the two varieties. The present study is based on corpus data coming from the International Corpus of Learner English for EFL and from the International Corpus of English for ESL, and representing different L1 populations among the two varieties. Relying on a corpus of native English as a reference, I examine the well-formedness of causative constructions in EFL and ESL, but also their idiomaticity, which is measured through a collostructional analysis (Stefanowitsch & Gries 2003) of the lexemes occurring in the non-finite verb slot. For EFL, this investigation reveals, among others, that learners sometimes use non-standard patterns like [X cause Y Vprp] or [X make Y Vto-inf], and that they tend to produce certain infelicitous constructions, which display lexical preferences different from those of native speakers (e.g. make their norms legalised). These findings are compared with the results of the ESL corpus analysis. This study provides insights into the impact of the acquisitional setting on the behaviour of causative constructions, and hence helps to bridge the paradigm gap that exists between EFL and ESL (cf. Sridhar & Sridhar 1986, Mukherjee & Hundt 2011). More generally, it demonstrates the viability of construction grammar as a theoretical framework to conduct a corpus-based study of interlanguage since, given the right level of abstraction, this framework provides a tertium comparationis for the contrastive analysis of varieties that may not necessarily follow the same norms. The study also underlines the relevance of the collostructional method to perform a contrastive interlanguage analysis, by showing that in both native and non-native varieties words interact with constructions (though sometimes in different ways). Such considerations, hopefully, will contribute to a rapprochement between the constructionist approaches and second language acquisition.
The article presents the outcome of research on 30 books of Quranic interpretations for sura al- Ghasyiah, verses 17-26, which are strongly assumed to contain pedagogic meanings, concepts, and values that can be formulated into an instructional model. The research was conducted by analyzing the keywords of the verses lexically, contextually, and hermeneutically. Then, the meanings gained were categorized, compared, contrasted, and abstracted, so that they were eventually synthesized into a main idea as a hypothetical model, termed M-3 Model. The model consists of three main instructional activities, represented in the terms munazharah, mudzakarah, and muhasabah. The three activities are a mutually completing and supporting cycle for the achievement of various instructional objectives, ultimately to improve the ability and skills of students to think systematically, logically, creatively, and innovatively through the development of potentials and fi trah (human norm). Specifi cally, munazharah activity is expected to result in cognitivistic knowledge (ainal yaqin), mudzarakah to develop knowledge, experience, and values into faith-based knowledge (‘ilm al-yaqin), and muhasabah to encourage the achievement of knowledge and values whose truths have been proven (haqqul yaqin), so that they will be the driving force for various activities based on law, moral, and ethics. Because the model taught by God to human beings is still hypothetical and theoretical in nature, it is suggested that the model be empirically tested to be more valid. Keywords: Instructional Model, Munazharah, Mudzakarah, Muhasabah
In the light of the overall current strategies and directions of translation (orientation on the language, text and culture of the original, or on the language and cultural context of the target language), the author of the article provides a comparative analysis of the translations of F.M. Dostoevsky's novel Demons (chapter At Tikhon) into German, made by E.K. Rahsin and S. Geier. They are a part of the history of German-language translations of Dostoevsky's novels and are sampled for analysis as playing a significant role in the German reception of Demons in the 20th century. The translation by Rahsin was a result of teamwork. The issues of translation of the novel were discussed in the salon of Merezhkovsky. According to Rahsin, a good translator of Dostoevsky should be: 1) a chemist who finds the right words, 2) an engineer who reconstructs the sentences, 3) an artist who creates the arrangement of the action, pays attention to the shade of sound, rhythm, etc., and 4) a critic, an expert in the German language able to judge whether and which bold solutions / innovations are appropriate or not. S. Geier's approach to the artistic text and its translation was defined by the sound, so she sought to give the German translation the sound and syntax of the Russian original. The translations in the paper are compared on the lexical, syntactic and stylistic levels. It is stated that Rahsin and Geier try to find German equivalents of the original words and expressions in different ways. Rahsin gives explanations in the text of the novel, and Geier, trying not to disturb the sound of the original, often gives a fairly extensive explanation in the notes. Unlike Geier, Rahsin often orders words by the rules of the neutral norm of the German language, thus depriving them of stylistic coloring, expressiveness, smoothing its characteristic roughness. It is concluded that Geier successfully managed to bring the German translation to the original text. She did not seek to correct Dostoevsky's text or make it easier to read for the German public (which, the author believes, was the purpose of the translation by Rahsin), but gave it a rough feeling of the fresh and polyphonic sound of the Russian original.
In this paper, we propose a method for au-tomatic clause boundary annotation in the Hindi Dependency Treebank. We show that the clausal information implicitly encoded in a dependency structure can be made explicit with no or less human interven-tion. We exercised the proposed approach on 16,000 sentences of Hindi Dependency Treebank. Our approach gives an accuracy of 94.44 % for clause boundary identifica-tion evaluated over 238 clauses. The resul-tant corpus has varied usages and can be utilized for developing a statistical clause boundary identifier. 1
We present a new collection of treebanks with homogeneous syntactic dependency annotation for six languages: German, English, Swedish, Spanish, French and Korean. To show the usefulness of such a resource, we present a case study of crosslingual transfer parsing with more reliable evaluation than has been possible before. This ‘universal ’ treebank is made freely available in order to facilitate research on multilingual dependency parsing. 1 1
Bootstrap Effect Sizes (bootES; Gerlanc & Kirby, 2012) is a free, open-source software package for R (R Development Core Team, 2012), which is a language and environment for statistical computing. BootES computes both unstandardized and standardized effect sizes (such as Cohen’s d, Hedges’s g, and Pearson’s r) and makes easily available for the first time the computation of their bootstrap confidence intervals (CIs). In this article, we illustrate how to use bootES to find effect sizes for contrasts in between-subjects, within-subjects, and mixed factorial designs and to find bootstrap CIs for correlations and differences between correlations. An appendix gives a brief introduction to R that will allow readers to use bootES without having prior knowledge of R.
Sylvain Kahane est Professeur en Sciences du Langage à l'Université Paris Ouest - Nanterre et membre du laboratoire Modyco (CNRS UMR 7114). Support de présentation de Sylvain Kahane: PDF Podcast: Résumé de l'intervention: Nous présenterons les différentes couches d'annotation du treebank Rhapsodie, un corpus de français parlé richement annoté. Le corpus contient plusieurs niveaux de segmentation indépendants: en unités illocutoires pour la macrosyntaxe, en unités rectionnelles pour la mi...
Artiklis tulevad vaatluse alla 17. sajandi kiriklike teoste tõlkija ja keelenormi kujundaja Heinrich Stahli tekstides kasutatud kaheksa haruldast tüvisõna ja seitse tuletist, mis (1) esinevad tema teostes vaid ühe korra (nn hapax legomenon’id), (2) esinevad põhjaeesti kirjakeeles esimest korda just Stahli teostes, (3) ei ole tänapäeva kirjakeeles sellises vormis ja/või tähenduses kasutusel, (4) ei ole otselaenud (alam)- saksa keelest. Käsitluse eesmärgiks on välja tuua Stahli haruldast sõnavara, mis pole tänapäeva eesti keeles enam kas tüve, moodustusviisi või tähenduse poolest läbipaistev. Niisugused lekseemid võimaldavad täpsemat pilguheitu Stahli teoste eripärasele sõnavarale ja selle kasutamise motiividele. Artiklis käsitletud tüvisõnad peegeldavad arhailist sõnavarakihti, mis on enamjaolt rahvakeelne ja võib osaliselt pärineda varasematest, Stahli-eelsetest allikatest. Tuletised avavad lisaks ka mehhanisme, kuidas Stahl on kasutanud teksti vajadustest lähtudes produktiivseid tuletusvõimalusi või toetunud analoogiamallidele. Selline täpseid esinemissagedusi arvestav uurimus on võimalikuks saanud pärast Stahli tekstide korpuse lemmatiseeritud kuju valmimist 2013. aastal.Rare words from the works of Heinrich Stahl. The article takes a look at 8 stem words and 7 derivations used in the works of Heinrich Stahl, the 17th century translator of religious texts and shaper of language norms. The stem words and derivations discussed in the article (1) only occur once in his texts, (2) first appear in Literary North-Estonian in Stahl’s texts, (3) are not used in Modern Estonian in the same form and/or meaning, (4) are not direct loans from (Low) German. These lexemes reveal the peculiarity of the lexicon of Stahl’s works. Because of their rarity, the lemmatising of such units while coding the corpus of Stahl’s texts has been problematic. These archaic words are not transparent in stem, form or meaning in Modern Estonian. The lexical stems discussed in the article reflect an archaic layer of the lexicon which is mostly vernacular and may partly originate in earlier, pre-Stahl sources. Derivations, in addition to revealing archaic lexicon, also reveal the mechanisms of how Stahl used productive derivation or analogy patterns depending on the demands of the text.
In recent years much progress has been made in developing systematic protocols for finding linguistic metaphors in authentic language data. The description of conceptual structures, however, has not been placed on equally firm footing. One existing proposal, known as the five-step method, introduces systematicity to the process of determining conceptual structures of metaphors in discourse. However, it does not take sufficient steps to minimize intuition and to maximize transparency. This paper seeks to reduce these weaknesses by introducing the systematic use of dictionaries and a lexical database. The result is a more transparent and constrained method.
Slovene Lexical Database was created between 2008 and 2012 and represents a comprehensive syntactic and semantic description of a selected set of Slovene words. The description was based exclusively on the analysis of reference corpora of Slovene. The database is structured as a network of interrelated semantic and syntactic information about a particular word. Semantic level represents the top level in the hierarchy with the lexical unit as its core element. This includes all senses of the headwrd, multi-word expressions and phraseological units. Each sense is described with a short semantic indicator and/or whole-sentence definition which includes typical syntactic environment of the headword with the relevant number, form and semantic types in a valency frame (semantic frame). These are also reflected in a number of syntactic structures and corresponding collocations. All the higher types of information are confirmed by a selection of corpus examples. Multi-word expressions and phraseological units are treated independently from particular senses of the headword and have their own internal structure which requires the same types of information as single-word entries or senses.
Abstract for the 25th Scandinavian Conference of Linguistics<br/><br/>Some remarks on wordformation in Danish<br/><br/>Some Danish word formation phenomena pose a problem for the linguist, being a predicament for analysis. In Danish a train leaves the station when it afgår ‘leaves’, while a minister may gå af ‘resign’, whereas a Swedish minister may resign by (att) avgå ‘(to) resign’. Especially tricky are pairs like afholde ‘arrange, organise’ and holde af ‘like’ because of their abstract, but different, meanings, and because the phrasal verb also differs from concrete meanings of holde ‘hold’. In general, there are some patterns for these Danish compounds concerning their internal semantics, in that the same lexical items may be used for different purposes depending on whether they are formed as a straightforward linear sequence (a word formation) or a reversed sequence (a phrase). The problem is (i) how the two kinds of combinations should be analysed, and (ii) what patterns emerge from the potential combinations, and (iii) why there are differences between closely related languages like Danish and Swedish?<br/><br/>It seems to have to do with the semantics of the combinations and not with the basic lexical materials, and that raises the question how to explain the combinatorial patterns by a specific approach in semantics.<br/><br/>The problem may be illustrated by Danish deadjectival nominal conversions like (en) døvstum ‘(a) deaf-mute’. They may be considered copulatives (dvandvas) or may be regarded as appositional compounds depending on whether you focus on their extensional or their intensional meanings. As a copulative (deadjectival noun) døvstum denotes an entity (a person) that represents the union set of the properties (attributes) døv and stum (in that the person represents both all the people constituting the set of the deaf and all the people constituting the set of the mute; i.e. the sum of all entities with either of those properties), whereas as an appositional (deadjectival adjective) compound the expression døvstum denotes the intersection of the sets of the properties (attributes) døv and stum respectively; i.e. individuals with both properties. This kind of analysis may be controversial, but the basic claim is that a primitive set-theoretical notion may be a way of handling adjectival combinations like these.<br/><br/>This kind of approach may also be appropriate when dealing with the formation vs phrase problem illustrated above (afgå vs gå af), in that specific combinations seem to be based on special semantic perceptions of the language users – which can be explained set-theoretically – and in that one may invoke a particular notion called “normative”. If “formative” is the Chomskyan notion of an articulated expressions (in a sentence or phrase) then one might propose a technical term for expressions found in parallel in related languages (like Danish and Swedish) and, crucially, mutually understandable (a minister may ‘gå af’ or ‘afgå’ in both languages and be understood) but with different norms regulating what is licensed in each language. The term ‘normative’ may be suggested for this phenomenon.<br/><br/>The presentation will elaborate on the theoretical and the analytic problems of the approach, and illustrate this by a fair number of excerpts and examples.<br/>
In this paper, we discuss our efforts to anno-tate nominals in the Hindi Treebank with the semantic property of animacy. Although the treebank already encodes lexical information at a number of levels such as morph and part of speech, the addition of animacy informa-tion seems promising given its relevance to varied linguistic phenomena. The suggestion is based on the theoretical and computational analysis of the property of animacy in the con-text of anaphora resolution, syntactic parsing, verb classification and argument differentia-tion. 1
The recent success of statistical parsing methods has made treebanks become important resources for building good parsers. However, constructing highquality annotated treebanks is a challenging task. We utilized two publicly available parsers, Berkeley and MST parsers, for feedback on improving the quality of part-of-speech tagging for the Vietnamese Treebank. Analysis of the treebank and parsing errors revealed how problems with the Vietnamese Treebank influenced the parsing results and real difficulties of Vietnamese parsing that required further improvements to existing parsing technologies. 1
Information Structure (IS) determines the “communicative” segmentation of the meaning of an utterance, which makes it central to the semantics‐syntax‐ intonation interface and therefore also to NLP. Despite this relevance, IS has not received much attention in the context of the majority of the reference treebanks for data-driven NLP that already contain a semantic and syntactic layers of annotation. We present our work in progress on the annotation of the Penn TreeBank with the thematicity dimension of the IS as defined in the Meaning-Text Theory. We experiment with tagging and transitionbased parsing techniques. Especially the latter achieve acceptable accuracy with even very small training samples, which is promising for languages with scarce resources.
Workload capacity, an important concept in many areas of psychology, describes processing efficiency across changes in workload. The capacity coefficient is a function across time that provides a useful measure of this construct. Until now, most analyses of the capacity coefficient have focused on the magnitude of this function, and often only in terms of a qualitative comparison (greater than or less than one). This work explains how a functional extension of principal components analysis can capture the time-extended information of these functional data, using a small number of scalar values chosen to emphasize the variance between participants and conditions. This approach provides many possibilities for a more fine-grained study of differences in workload capacity across tasks and individuals.
We set forth to show that lexical connectivity plays a role in understanding early word learning. By considering words that are learned in temporal proximity to one another to be related, we are able to better predict the words next learned by toddlers. We build conditional probability models based on data from the growing vocabularies of 77 toddlers, followed longitudinally for a year. This type of conditional probability model outperforms the current norms based on baseline probabilities of learning given age alone. This is a first step to capturing the interaction between a child’s productive vocabulary and their learning environment in order to understand what words a child might learn next. We also test different types of variants of this conditional probability and find that not only is there information in words that are learned in proximity to one another but that it matters how models integrate this information. The application of this work may provide better cognitive models of acquisition and perhaps allow us to detect children at risk for enduring language difficulties earlier and more accurately.
Abstract This study assesses the effects of a parent-child reading project on the development of a variety of French prereading skills in Innu-speaking Kindergartners over the course of a school year. Phonological memory, expressive and receptive lexical knowledge, morphosyntactic knowledge, and basic arithmetic concepts were tested in two subgroups of a single cohort, one composed of participants in a parent-child reading project (experimental group) and the other composed of non-participants (control group). The results show that both groups made gains against French mother tongue age-level norms over the course of the year. The experimental group, whose members started the year with higher skill levels in a number of areas, improved more and on a greater variety of tasks than the control group. While the actual role the reading project played in the children’s gains cannot be determined because of intertwining of factors, bringing books into homes and informing parents about the importance of reading likely had a positive effect on project participants. Résumé Cette étude examine l’impact d’un projet de lecture parent-enfant sur le développement des habiletés préalables à l’apprentissage de la lecture chez les enfants innus inscrits à la maternelle. Nous avons comparé les résultats obtenus par deux groupes d’enfants—un dont les membres participaient au projet de lecture (groupe expérimental) et un dont les membres ne participaient pas (groupe de contrôle)—à une variété de tâches en utilisant un protocole expérimental pré-test, intervention, post-test. Alors que les deux groupes ont fait de bons progrès au cours de l’année scolaire, rattrapant une partie de leur retard initial par rapport aux normes francophones, le groupe expérimental a fait des gains plus importants. Le rôle exact joué par la lecture dyadique dans ces gains n’a pas pu être mesuré avec précision à cause d’un croisement de facteurs, mais certaines indications nous laissent croire que l’introduction de livres dans les foyers des participants a eu des retombées très positives.
In assignment initially theorethical context of linguistic expressions is defined.First part continues to place the theory of language stratification, started in 1932 in Prague Linguistic Circle and in the second half of 20th century assumed in Slovene linguistic.First part also discusses about creation of slovene literary language and literary norm, continuing with arrangement of expressions in slovene theory of language stratification by Toporii (1971).Next to this contemporary definition of social stratification by Andrej E. Skubic ( 2005) is presented, concentrating on speeches of social groups (sociolects).Second part of assignment -based od research about elements of sociolects in discourse of slovene literature by Skubic (2006) and the concept of script by Roland Barthes (1971)places analised literary works, Fuinski bluz (Skubic, 2001) and efurji raus! (Vojnovi, 2008) into the fifth degree of development script as speech (Barthes, 1971).After presentation of literary works the analysis of communicative situations, where non-literary elements are used in speeches of all five literary figures, is discussed.At the beginning of the third part the treoretical description of reflextion of prague theory of linguistic stratification in SSKJ is presented.Moreover the presentation of activities during the time when SSKJ was published to SP 2001 is described, concluding with discussion about inclusion of linguistic stratification and sociolects in SP 2001.In the second, practical part linguistic analysis of elements of sociolects and speeches of literary figures are presented.The analysis is concentrated on lexical linguistic level and try to determine the relation between literary and non-literary language in works of contemporary slovene urban prose.
The article aims at drawing the attention of language teachers to a huge number of phraseologisms which exist in every language and which are traditionally rarely used while teaching and learning languages and cultures. The fact that phraseology shows the features of folk culture is now widely accepted. The subject of the research is the experience of the author, namely, the investigation of phraseologisms related with lexicology, the stylistics of lexis and Latvian language for practical uses. These days a wide range of investigation is characteristic of linguistics. Lingvo-culturologic viewpoint to the learning of units of speech takes an important part in the investigations. The question of interaction between language and culture is nowadays relevant in our society, which experiences the growth of global problems, therefore, it is becoming essential to consider the versatility and particularity of behaviour of different nations. Looking at relations between different nations, it is important to foresee potential cultural misunderstandings. It is also important to determine cultural values which form the basis for communicational behaviour. In higher education institutions these skills are obligatory for students who, for various reasons, get into different cultural environments. Until 1990 students were encouraged to memorize word forms and to unpack the meaning of words (usually by means of translation) and only at the end of the 90’s the semantic and practical aspects of speech were highlighted in the process of teaching Latvian as a foreign language. The main unit of lingvo-culturologic viewpoint is lingvocultureme. Lingvocultureme belongs both to language and culture, as it unites the meaning of language and culture which exists outside the boundaries of the language system. Lingvocultureme may be the unit of both lexis and syntax: a word, a phrase, a sentence, a text (Gavrilina, Vulane, 2008, pp. 21). According to lingvo-culturologic language research, linguistic analysis allows to divide units of language into three types: words and sayings which totally coincide in the languages compared; words and sayings which partially coincide in the languages compared; words and languages which do not coincide in the language compared. Since 2005 various aspects of lingvo-culturology have also been the subject of the project which is carried out by the European Society of Phraseology. The aim of the project is to discover the similarities rather than differences, i.e. to find the part of phraseology which is common for European languages. The results of the previous project show that identical or similar phraseologisms can be found in nearly 50 languages. Idioms with similar lexical and semantic structure can be found even in languages which are not genetically connected and whose areas of usage are distant from each other. It is important to note that each nation has its own cultural vision of the world and a cultural-historic way. In the phraseologisms of each language the culture, the way of thinking and values of each nation are conveyed. Phraseologisms in texts encourage students to search for culturologic information, through which students can develop their communicative, language, socio-cultural and learning competences. These opportunities are important in the cases when students who get into different cultural environment for various reasons and who have different nationalities study in one group (in homogeneous cultural environment these opportunities are formal). Various problems are possible while learning phraseologisms: different theories on phraseologisms; students do not know phraseologisms; students know phraseologisms, but do not use them; it is impossible to translate the figurative sense of some words literally; different associations (e.g. sun); the same phraseologisms are used in different contexts (e.g. as brave as a lion, as angry as a lion); students need to learn the expressiveness of phraseologisms. While learning lingvoculturemes new opportunities are created: to develop lexis; to get familiarized with the heritage of your own language and culture; to know more about different cultural environment (the values, stereotypes, norms of behaviour, speech etiquette, customs, way of living, etc. of each nation); enrich intercommunion paying respect to cultural heritage; to motivate language users to take interest in linguistic and extralinguistic research. Such information will enrich both sides, as language users who share their experience learn from the cultural traditions of other nations. DOI: http://dx.doi.org/10.7220/2335-2027.2.9
This study aims to explore National Palace Museum (NPM)'s English texts for its exhibits. NPM, a treasure vault of valuable ancient Chinese cultural artifacts, has endeavored to enter the global arena in recent years. NPM's ambition can be clearly seen on its home page, which provides a great variety of languages. If fact, the vast majority of international visitors rely on its English texts to access knowledge of NPM's exhibits. As such, NPM's English texts play a critical role in making their exhibits understandable to international visitors. However, the English texts of NPM's exhibits are mostly verbatim translation from their original Chinese texts. Its lexical choices and syntactic structures as well as its rhetorical organizations are all highly circumscribed by Chinese norms of language and thinking. In other words, NPM's English texts are the result of using formal equivalence translation (Niad, 1969). Such Chinese-circumscribed English texts, with a low degree of comprehensibility, are ”exotic” and distant to the vast majority of international visitors. To date, there has been a paucity of research addressing the issue of this translation strategy. The current case study thus attempts to explore the reason behind and the influence of such a strategy by NPM. Meanwhile, this study hopes to serve as a reference for NPM translators-to take into account the naturalness of their English texts, thereby enhancing the comprehensibility of their exhibits for international visitors. All things considered, this factor would actually be of utmost importance in NPM's pursuing its goal to enter the global arena.
We investigate statistical dependency parsing of two closely related languages, Croatian and Serbian.As these two morphologically complex languages of relaxed word order are generally under-resourced -with the topic of dependency parsing still largely unaddressed, especially for Serbian -we make use of the two available dependency treebanks of Croatian to produce state-of-the-art parsing models for both languages.We observe parsing accuracy on four test sets from two domains.We give insight into overall parser performance for Croatian and Serbian, impact of preprocessing for lemmas and morphosyntactic tags and influence of selected morphosyntactic features on parsing accuracy.
In this paper, we provide a quantitative analysis of non-projective constructions attested in the Ancient Greek Dependency Treebank (AGDT). We consider the different types of formal constraints and metrics that have become standardized in the literature on non-projectivity (planarity, wellnestedness, gap-degree, edge-degree). We also discuss some of the linguistic factors that cause non-projective edges in Ancient Greek. Our results confirm the remarkable extension of non-projectivity in the AGDT, both in terms of quantitative incidence of non-projective nodes and for their complexity, which is not paralleled by the corpora of modern languages considered in the literature. At the same time, the usefulness of other constraint (especially well-nestedness) is confirmed by our researches. 1
This paper sets out to study the letters of Gaston B., a French prisoner of war held in captivity in the camp of Münster (Germany) from the beginning of the First World War until its end. These letters make possible a relativisation of linguistic macrohistory through microhistory, by focussing on the grassroots level and by using as sources the traces of people with no significant name or identity. They shed important light on how a member of a lower class acquired the prescriptive linguistic norm through his schooling at the end of the nineteenth century and how this affected his subsequent linguistic behaviour. An individual is exposed to the political and social dimension of language planning, and his language reflects its level of success, but also reveals what grammatical tools and rules have been focused on during his schooling.
В статье рассматривается проблема нормы и нормативного подхода к языку в диахроническом плане.Определяется специфика нормативного похода к языковым средствам в различных лингвистических традициях и выявляются основные характеристики лингвистической нормы.В статье указывается, что на каждом этапе развития языка складываются свои нормы как резуль
In this paper, we investigate errors in syntax annotation with the Turku Dependency Treebank, a recently published treebank of Finnish, as study material. This treebank uses the Stanford Dependency scheme as its syntax representation, and its published data contains all data created in the full double annotation as well as timing information, both of which are necessary for this study.
Nowadays, advertising is becoming an integral part of our daily life and is playing an increasingly significant role in modern society. It appears we are living in an advertising world. Many studies have been carried out in this field, and among them the study of advertising language has attracted particular attention from social linguists. As a way to promote the sales of products, advertisements must conform to the AIM principle—to grab readers’ attention, arouse their interest, and construct their memory to achieve the ultimate goal of triggering their action. Thus, the advertisers seek for attention-attracting strategies. The application of language deviation technique is an efficient way. Deviation refers to the special or unusual expression that deviates from normal norms and it appears in various forms such as deviation of phonology, lexicon and grammar. This paper attempts to give a description of language deviations in English advertising including phonological, graphological, lexical, and grammatical deviation.
Abstract Semantic substitution errors (slips of the tongue) naturally occurring in Russian normal speech were analyzed for word frequency, word length, target-error cooccurrence strength, and word association norms. Target word frequencies were found to be significantly lower than error word frequencies; besides, there is a very significant positive correlation between target and error frequency values. Contrary to the view that the frequency effect is located at the stage of phonological encoding, the results suggest that frequency is coded at an earlier stage of lexical selection. Word length is a significant variable that determines the outcome of the error for non-cohyponym target-error pairs but not for cohyponym pairs. At the same time, cohyponym target-error pairs are characterized by much higher cooccurrence measures and stronger associative links compared to non-cohyponym pairs. Theoretical implications of these findings are discussed
The Treebanks as the sets of syntactically annotated sentences, are the most widely used language resource in the application of Natural Language Processing. The occurrence of errors in the automatically created Treebanks is one of the main obstacles limiting the using of these resources in the real world applications. This paper aims to introduce an statistical method for diminishing the amount of errors occurred in a specific English LTAG-Treebank proposed in Basirat and Faili (2013). The problem has been formulated as a classification problem and has been tackled by using several classifiers. The experiments show that by using this approach, about 95% of the errors could be detected and more than 77% of them could successfully be corrected in the case of using Adaboost classifier. In addition, it has been shown that the new treebank could reach a high of 76% F-measure which is 8% higher than the original treebank.
National audience
This paper discusses the extension of a sys-tem developed for automatic discovery of tree-bank annotation inconsistencies over an entire corpus to the particular case of evaluation of inter-annotator agreement. This system makes for a more informative IAA evaluation than other systems because it pinpoints the incon-sistencies and groups them by their structural types. We evaluate the system on two corpora- (1) a corpus of English web text, and (2) a corpus of Modern British English. 1
During the project “Development of Communicative Competence in the Early Croatian Language Discourse”, a research on the mastering of Croatian as a second language was carried out among the children of Croatian emigrants to Germany. It needs to be stressed that the participants start learning Croatian systematically only within the program of the Croatian tuition abroad, when they encounter the standard Croatian idiom which is more or less different than their local Croatian idiom which the speakers were exposed to in their families. The input language is the individual organic idiom of the Croatian language, while the target language refers to the standard Croatian language. The process of the acquisition of the organic or first language (L1) differs from the process of the other or second language (L2) learning. This is due to the fact that in the process of the non-mother tongue mastering an interlanguage is created in which elements of the first and the second language interfere. In this process non-mother tongue learning can have a double meaning: learning a completely new and unfamiliar language system (foreign language), or acquiring a language idiom which the speakers had already partially mastered, usually in the early childhood. In that case we speak of the heritage language, the language of their parents, their cultural circle, and their national identity. Although all pupils who participated in this research had been born in the Federal Republic of Germany and have the German language as their dominant idiom, most of them consider Croatian to be their mother tongue. However, this research and the communicative practice have confirmed that the German language competence of the participants is higher than their Croatian language competence. Despite their personal attitudes towards Croatian as their mother tongue, the truth is that they learn Croatian as special kind of second language (heritage). The paper brings out the results of 150 pupils participated in the research, ranging from 6-18 years of age, and attending Croatian classes abroad in the German province Baden-Wurttenberg. A test of communicative competence was conducted and the written works of pupils were analyzed in order to examine their language competences in grammatical and lexical level of the Croatian language, according to age, cognitive development, communication, language exposure and language foreknowledge. The analysis of questionnaires was to determined the attitudes of pupils and teachers in Croatian tuition abroad (motivation, purpose and learning needs, socio-cultural environment ).The data were analyzed with the SPSS statistical analysis software. The method used was Pearson’s correlation coefficients to show correlation between the dependent variable (mastery of the language competences) and the independent variable (years they had spent learning Croatian). The Mean value, standard deviation, median-central value, and minimal and maximal score were used in the description and comparison of deviations from the standard language norm. Kolmogorov Smirnov z test was used to test the normal distribution, Kruskal Wallis test for differences between the different age groups of students and teachers, and the Mann Whitney test was used to assess the differences in the attitudes of teachers with regard to the degree and gender. Research results will be used for the analysis of language competences in pupils whose second (heritage) language is Croatian and for assessing the level of acqusition of Croatian language teaching abroad.
Dans cette thèse, nous menons une réflexion sur les notions de norme(s) et d usage(s) langagiers dans le cadre du domaine du contrôle aérien. Ce domaine offre l exemple parfait de la mise en œuvre d une norme langagière: la "phraséologie" aéronautique, le langage opératif censé permettre des communications sûres et efficaces entre pilotes et contrôleurs, lors des situations les plus courantes. Lorsque la phraséologie ne suffit pas, ces derniers ont recours à une forme langagière plus naturelle, le "plain language". De nombreux problèmes relatifs à la mise en œuvre de ce dernier ont rapidement vu le jour et ont suscité des interrogations, notamment chez les professionnels de l enseignement de l anglais de l aviation. Pour répondre aux besoins spécifiques de l ENAC, cette thèse dresse un panorama des usages faits de la langue anglaise par les contrôleurs français et les pilotes étrangers lors de leurs communications radiotéléphoniques. Notre méthode d analyse consiste en une étude comparative entre un corpus de référence, représentant la norme, et un corpus de communications réelles, représentant les usages. Cette analyse comparative nous permet de repérer, de décrire et de catégoriser les formes langagières employées lors de situations routinières de la navigation aérienne. Certaines différences sont ainsi repérées entre les deux corpus. En fonction de la situation, les pilotes et les contrôleurs peuvent procéder à des variations d ordre lexical, sémantique, syntaxique et discursif. Alors que certains semblent subir l influence de la langue naturelle, d autres semblent mettre en œuvre une stratégie communicationnelle pour tenter d "humaniser" ou de modaliser le contenu de leurs messages.
Reviewed by: Cultural Capital, Language and National Identity in Imperial Spain by Lucia Binotti Ryan Prendergast Binotti, Lucia. Cultural Capital, Language and National Identity in Imperial Spain. Woodbridge, Suffolk: Tamesis, 2012. 207 pp. Lucia Binotti’s study approaches the question of national identity by way of language, and, more specifically, how language and theories about language shaped the cultural underpinnings of sixteenth-century Spain and became the principal mechanism for constructing cultural identity in the Spanish Renaissance. Beginning with the understanding that there was a competition with the main tenets of the Italian Renaissance, Cultural Capital, Language and National Identity in Imperial Spain offers readers the opportunity to better understand how Spain’s intellectuals and elite both imitated and departed from the linguistic and cultural models [End Page 572] that were the basis for the Italian Renaissance. Finally, the book examines how authors used texts to construct and project the image of Spain as an imperial power. Within these parameters, Binotti reflects on the inextricably linked issues of courtly society, book production, patronage, literacy, translation, and canon formation, and she clearly articulates the complex coordinates that framed and amplified the field of cultural inquiry in the Renaissances of both Spain and Italy. Cultural Capital is divided into two parts and has five chapters. Part one, “Patronage, Audiences and Cultural Markets,” encompasses the first three chapters. In chapter one, Binotti focuses on how the Italian tradition appropriated Spanish sentimental fiction, in particular Diego de San Pedro’s Cárcel de amor and La historia de Grisel and Mirabella by Juan de Flores. The editorial choices for the Italian translations of these two texts speak to the changing balance of power between “courtly society and the new reality brought about by the Spanish domination of Italy” (18). This, in turn, constructed new social meaning for the texts and the genre. Binotti convincingly demonstrates how there are echoes of Carcel de amor in Baldassare Castiglione’s Il cortegiano and how Spanish sentimental fiction influenced Ludovico Ariosto’s Orlando furioso. She goes on to argue that these appropriations by Italians serve as an example of how “the commercialization of the book became a powerful factor in the drastic changes in the habits of consumption … that occurred in Europe at the middle of the [sixteenth] century” made the book “an agent of class distinction” and redefined what constituted cultural capital (50). The second chapter takes up the issue of literature and canon formation, using as a point of departure the idea that by the mid-sixteenth century the printing press had turned into a much more full-fledged enterprise that expanded the reading public. As a result, printed texts had the immense power to influence and define the ideology of the era. Binotti carefully studies the editorial project of Alfonso de Ulloa, who translated to Italian from Spanish and vice versa. She asserts quite convincingly that Ulloa, like most sellers of literature at the time, was involved in the process of defining and creating a “national literary profile” by “shaping cultural authority and cultural capital” (53). To illustrate these ideas, Binotti examines Ulloa’s 1553 translation of Orlando furioso into Spanish, noting how it was packaged as a classic. She suggests that it was meant for Italians who wanted to enjoy the poem while learning Spanish, but it also acted as a conduit for “idealized social and linguistic norms” to Spain (59). Chapter three examines Luis de Góngora’s Fábula de Polifemo y Galatea, in which Binotti analyzes erotic and vulgar imagery within an ekphrastic framework. She points out how the poem casts the reader in the role of a viewer by “transposing complicated tropes into visual images” that ultimately elicit both aesthetic and erotic responses (101). Through some adroit and suggestive close reading, Binotti also elucidates how Góngora, taking his cue from the conventions of painting, is able to weave together poetry and painting to achieve his erotic end, though often veiled in innuendos, puns, and double entendres. She maintains that the Polifemo is an aesthetic manifesto while also being sophisticated pornography, asserting that the “titillation of desire” is linked to the text’s “masked linguistic titillation” (101). By [End Page...
Modern communication environments have changed the cognitive patterns of individuals, who are now used to the interaction of information encoded in different semiotic modalities, especially visual and linguistic. Despite this, the main premise of Corpus Linguistics is still ruling: our perception of and experience with the world is conveyed in texts, which nowadays need to be studied from a multimodal perspective. Therefore, multimodal corpora are becoming extremely useful to extract specialized knowledge and explore the insights of specialized language and its relation to non-language-specific representations of knowledge. It is our assertion that the analysis of the image-text interface can help us understand the way visual and linguistic information converge in subject-field texts. In this article, we use Frame-based terminology to sketch a novel proposal to study images in a corpus rich in pictorial representations for their inclusion in a terminological resource on the environment. Our corpus-based approach provides the methodological underpinnings to create meaning within terminographic entries, thus facilitating specialized knowledge transfer and acquisition through images.
Many countries use national-level surveys to capture student opinions about their university experiences. It is necessary to interpret survey results in an appropriate context to inform decision-making at many levels. To provide context to national survey outcomes, we describe patterns in the ratings of science and engineering subjects from the UK’s National Student Survey (NSS). New, robust statistical models describe relationships between the Overall Satisfaction’ rating and the preceding 21 core survey questions. Subjects exhibited consistent differences and ratings of “Teaching”, “Organisation” and “Support” were thematic predictors of “Overall Satisfaction” and the best single predictor was “The course was well designed and running smoothly”. General levels of satisfaction with feedback were low, but questions about feedback were ultimately the weakest predictors of “Overall Satisfaction”. The UK’s universities affiliated groupings revealed that more traditional “1994” and “Russell” groups over-performed in a model using the core 21 survey questions to predict “Overall Satisfaction”, in contrast to the under-performing newer universities in the Million+ and Alliance groups. Findings contribute to the debate about “level playing fields” for the interpretation of survey outcomes worldwide in terms of differences between subjects, institutional types and the questionnaire items.
We present an empirical study on constructing a Japanese constituent parser, which can output function labels to deal with more detailed syntactic information.Japanese syntactic parse trees are usually represented as unlabeled dependency structure between bunsetsu chunks, however, such expression is insufficient to uncover the syntactic information about distinction between complements and adjuncts and coordination structure, which is required for practical applications such as syntactic reordering of machine translation.We describe a preliminary effort on constructing a Japanese constituent parser by a Penn Treebank style treebank semi-automatically made from a dependency-based corpus.The evaluations show the parser trained on the treebank has comparable bracketing accuracy as conventional bunsetsu-based parsers, and can output such function labels as the grammatical role of the argument and the type of adnominal phrases.
We present a reformulation of the word pair features typically used for the task of disambiguating implicit relations in the Penn Discourse Treebank. Our word pair features achieve significantly higher performance than the previous formulation when evaluated without additional features. In addition, we present results for a full system using additional features which achieves close to state of the art performance without resorting to gold syntactic parses or to context outside the relation.