Papers reviewed and determined not to be word norm studies. Use the flag icon to report errors or suggest re-inclusion.
16504 papers
The paper refers to the specifics of the French-Polish automated translation. The author summarizes briefly the rules of creating electronic lexical databases in the context of the Object Oriented Approach. This particular approach, created for the purpose of translation, especially for automated translation, by Wiesław Banyś, offers a specific approach towards the language description with the help of object classes seen through their semantic relations. The purpose of this description is to determine correctly the meaning of the words and polysemic expressions of the source language in the translation into the target language. Therefore, the main question of the paper concerns the role of frames and/or scripts criterion in the process of the natural language words’ disambiguation for automated translation. By analysing the example of the French word collier the Author shows how it is possible to solve the problem of near synonyms in automated translation with the use of the above-mentioned criterion.
The vocative is a residuary case in most Indo-European languages, mirroring a particular Proto-Indo-European status. Its syntactical function is preserved in the descendant languages, but the morphological aspects are strongly simplified. In Latin, not unlike the cognate languages, the general tendency is toward a formal overlapping with the nominative case. The Romanian vocative is, in the Romance frame, surprisingly multifarious. It displays four distinct variants: desinence and intonation; desinence, intonation and prolongation of the final vowel; intonation and vowel prolongation; solely intonation. Old Romanian texts attest the tendency of gradually replacing the vocative form with the nominative form, perceived as more expressive. On the other hand, there is an observable development of the formal marks specific to this syntactical function; these marks are only partially inherited from Latin. In nowadays Romanian language the formal specificity of the vocative case is not diminishing – on the contrary, some colloquial vocative forms (not yet acceptable in the frame of the linguistic norm) emphasize an unambiguous linguistic will to maintain this case, while the general tendency is to reduce as much as possible the differences between the actual two cases of the Romanian language, nominative-accusative and genitive-dative.
Bilingual dependency parsing aims to improve parsing performance with the help of bilingual information. While previous work have shown improvements on either or both sides, most of them mainly focus on designing complicated features and rely on golden translations during training and testing. In this paper, we propose a simple yet effective translation constrained reranking model to improve Chinese dependency parsing. The reranking model is trained using a max-margin neural network without any manually designed features. Instead of using golden translations for training and testing, we relax the restrictions and use sentences generated by a machine translation system, which dramatically extends the scope of our model. Experiments on the translated portion of the Chinese Treebank show that our method outperforms the state-of-the-art monolingual Graph/Transition-based parsers by a large margin (UAS).
Ontology is a formal specification of knowledge by a set of concepts with in a domain and their semantic relationship. Different Ontologies are devolved by developer for a same domain differently. Furthermore, ontology tools using different language. Ontology merging in semantic web is used to resolve the heterogeneity problem among the source ontology. The language level and ontology level mismatches are the heterogeneity problem in the ontology merging. Lexical matching strategies resolve the syntax based on Jaro distance. The word net is a lexical database with synset which solve the problem of semantic mismatches. A knowledge base is constructed by using Semantic Web Rule Language (SWRL) for ontology level mismatches. A merging framework identifies the similarities and dissimilarities of source ontologies are merged to resolve the heterogeneity problem. Keywords: Ontology, Heterogeneity problem, word net, Lexical, Semantic, SWRL
Objectives: Risky behavior seriously impacts the life of adult patients with attention deficit hyperactivity disorder (ADHD). Such behaviors have often been attributed to their exaggerated reward seeking, but dysfunctional anticipation of negative outcomes might also play a role. Methods: The present study compared adult patients with ADHD (n=28) with matched healthy controls (n= 28) during anticipation of monetary losses versus gains while undergoing functional magnetic resonance imaging (fMRI) and skin conductance recording. Results: Skin conductance was higher during anticipation of losses compared to gains in both groups. Affective ratings of predictive cues did not differ between groups. ADHD patients showed increased activity in bilateral amygdalae, left anterior insula (region of interest analysis) and left temporal pole (whole brain analysis) compared to healthy controls during loss versus gain anticipation. In the ADHD group higher insula and temporal pole activations went along with more negative affective ratings. Conclusions: Neural correlates of loss anticipation are not blunted but rather increased in ADHD, possibly due to a life history of repeated failures and the respective environmental sanctions. Behavioral adaptations to such losses, however, might differentiate them from controls: future research should study whether negative affect might drive more risk seeking than risk avoidance.
The parsing of non-projective dependency has been the bottleneck problem for the accuracy of dependency parsing system. There is extensive existence of non-projective dependency in Mandarin Chinese. Based on the U-penn Chinese Treebank,the paper investigates the distributions of edge-crossing dependency,an example of nonprojective dependency, in Mandarin Chinese. It is found that when edge crossings produced by clause sharing and circumposition are included,edge-crossing dependency can be found in 2849 sentences,about 16. 42 percent of the 15162-sentence corpus; when the range is limited to clauses,1712 sentences are identified,accounting for 11. 29 percent of the overall sentences; and apart from edge crossings caused by coordination,other usages can be divided into two categories with nine sub-categories.
Abstract. The eXtensible Markup Language (XML) has quickly become the de facto standard for data exchange via web. An XML document can be viewed as an ordered tree that has at least one node. Each node must be labeled by using a scheme approach to describe the XML data structure. There are two famous existing encodings, namely Dewey and Inteval Encodings. In this paper, ORDPATH encoding based on Dewey together with the two other encodings are empirically demontrated on dblp, nasa, and treebank datasets. The results show that while a new node was inserted into the tree, Dewey and Interval have to relabel the inserted node’s siblings and modify the interval number of the sibling nodes, respectively. Whereas, the ORDPATH eliminates this problem by adding an even number used as a caret for the new insertion node.Keywords: Ordered Tree, XML, ORDPATH. Abstrak. EXtensible Markup Language (XML) terus menjadi standard untuk penukaran data melalui web. Sebuah dokumen XML dapat ditinjau menjadi tree terurut yang berisikan sedikitnya satu node. Setiap node harus dilabelkan menggunakan sebuah algoritma pelabelan untuk mendeskripsikan struktur data XML tersebut. Ada dua algoritma encoding yang terkenal selama ini, Dewey dan Interval encoding. Pada tulisan ini, metode ORDPATH yang berbasiskan Dewey bersama-sama dengan Dewey dan Interval didemontrasikan secara empiris dengan menggunakan dataset dblp, nasa, dan treebank. Hasil menunjukkan bahwa ketika node baru dimasukkan ke dalam tree, Dewey dan Interval harus melakukan pelabelan kembali dan memodifikasi interval sibling node. Akan tetapi, ORDPATH dapat mengatasi masalah ini dengan memberikan angka genap yang digunakan sebagai penanda untuk node baru.Kata Kunci: Ordered Tree, XML, ORDPATH.
In the Polish language there are expressions like pierwszy i drugi, taki lub inny which consist of adjectives, numerals or pronouns constituting determiners of nouns (e.g. ten czy ów zawodnik). Some of these sequences are distributional equivalents of the compound subject owing to the reference and syntax. Such a sequence, like series subject, relates to at least two referents in the extralinguistic reality and allows the occurrence of ad formam and/or ad sensum adjustment of verbs, for example: Pierwszy i drugi zawodnik mają / ma dobry czas. The habit of adjusting the person form in the Polish language is often unstable.The article discusses chosen issues concerning features of structure and syntax requirements among the sequences which function in sentences similarly to nominal series in the function of a subject.Models of syntax adjustment of verb in various types of sequences in the modern Polish were created on the basis of quantitative and qualitative analyses. Such a formulation of the rule may be directly provided in the correctness publications which codify the modern linguistic norm.
The French Lexical Network (fr-LN) is a global model of the French lexicon presently under construction. The fr-LN accounts for lexical knowledge as a lexical network structured by paradigmatic and syntagmatic relations holding between lexical units. This paper describes how morphological knowledge is presently being introduced into the fr-LN through the implemen-tation and lexicographic exploitation of a dynamic morphological model. Section 1 presents theoretical and practical justifications for the approach which we believe allows for a cogni-tively sound description of morphological data within semantically-oriented lexical databases. Section 2 gives an overview of the structure of the dynamic morphological model, which is constructed through two complementary processes: a Morphological Process—section 3—and a Lexicographic Process—section 4. 1
The aim of this paper is to develop an analysis of Hanging Topic and Hanging Topic Left Dislocation in Brazilian Portuguese. To do so, we adopt the framework of Segmented Discourse Representation Theory (SDRT), because it allows one to describe the rhetorical relations, and therefore to study the use contexts of some marked constructions. The corpus which is the basis of analysis of this research are the interviews published in the VEJA magazine in the 1970s, as well as Dialogues between Informant and Documenter (DID) provided by the NURC-SP project (Cultured Urban Linguistic Norm – São Paulo). From the data analysis, it was possible to notice a clear distinction between the uses of the constructions under investigation in the two corpora. We have noticed that hanging topics are discursive units per se, which relate to the context by a subordination relation and to the comment by Frame or Attribution. Hanging topics seem therefore to be pragmatically indistinguishable from frame-setting topics, once both exert the function of inaugurating complex discursive units, so as to preserve textual progression.
Most standardized languages have a corresponding geo-political center, where the language of society offers a location to the linguistic norms and standards of the language in question. The purpose of this study is to identify the linguistic center for Spanish used in Spanish-language media in the United States. Some have dubbed this neutral accent, “Walter Cronkite Spanish” and argue that it is the discrete Mexicanization of the Spanish language (Ahrens 2004). The first is a reference to the legendary anchorman who was one of the first in broadcast media to attempt to eliminate all traces of an identifiable regional accent. The second nickname addresses Mexico’s heavy hand in Spanish-language media, as well as its large diaspora living in the U.S. These popular notions have raised the questions, is it possible for Walter Cronkite Spanish to have a geo-political center outside its country of use, in this case, Mexico? Or does a particular region in the United States make claim to a unique dialectal variety of Spanish? Or can we locate it somewhere else entirely?
A positive change in perceived self-efficacy expectation is considered an important goal of exposure-based treatments. In line with this, several studies have demonstrated a positive association between self-efficacy and therapy outcome. While those studies primarily focused on changes in self-efficacy expectation that are achieved through therapy, the present study sought to examine whether changes in self-efficacy prior to treatment have an influence on therapy outcome. To this end, 48 healthy subjects completed a differential fear conditioning task. After the fear acquisition phase, half of the subjects received a positive verbal feedback aimed at increasing self-efficacy (experimental group) whereas others received no feedback (control group). Our results not only show that self-efficacy beliefs can be enhanced through verbal feedback but also point to an enhanced extinction of conditioned fear in the experimental group relative to the control group, evident on the implicit (skin conductance responses) and explicit (valence rating) level. Our results may have clinical implications for exposure-based treatments in anxiety disorders.
We describe a technique to minimize weighted tree automata (WTA), a powerful formalisms that subsumes probabilistic context-free grammars (PCFGs) and latent-variable PCFGs. Our method relies on a singular value decomposition of the underlying Hankel matrix defined by the WTA. Our main theoretical result is an efficient algorithm for computing the SVD of an infinite Hankel matrix implicitly represented as a WTA. We provide an analysis of the approximation error induced by the minimization, and we evaluate our method on real-world data originating in newswire treebank. We show that the model achieves lower perplexity than previous methods for PCFG minimization, and also is much more stable due to the absence of local optima.
The paper examines the attitudes and language conceptions of dominating and non-dominating language communities of pluricentric languages which differ in many ways. It is shown that in dominating nations the one-nation-one-language concept is an undisputed basic concept which is shared by most speakers of these varieties and brings about a clear distinction between linguistic „standard“ and „nonstandard-forms“. Contrary to this speakers of non-dominating varieties find themselves under pressure to legitimate their „deviating“ language behavior, even though they might use the standard variety of their country. The situation also leads to a kind of „diglossia“ between the own national norms and the exogene norms of the dominating nation. The paper also looks at the psychological effects of this situation and at possible ways to overcome them. It is shown that the non-dominating varieties face a dilemma as a thorough codification of their actual linguistic norms sooner or later leads to a separation from the norms of the dominating variety and to the development of a language of it's own.
Abstract. At present, little is known about the structure of the mental lexicons of children who use cochlear implants. In this paper, we report preliminary analyses of errors from the Lexical Neighborhood Test, an open-set monosyllabic word recognition test, by children with profound hearing loss who use cochlear implants. Frequency and lexical neighborhood characteristics of errors were compared to those of target words using lexical databases derived from both child and adult word-frequency counts. Results demonstrate that pediatric cochlear implant users recognize words in the context of other, phonologically similar words. The structure of the mental lexicons of these children is discussed.
We describe a technique to minimize weighted tree automata (WTA), a powerful formalisms that subsumes probabilistic context-free grammars (PCFGs) and latent-variable PCFGs. Our method relies on a singular value decomposition of the underlying Hankel matrix defined by the WTA. Our main theoretical result is an efficient algorithm for computing the SVD of an infinite Hankel matrix implicitly represented as a WTA. We provide an analysis of the approximation error induced by the minimization, and we evaluate our method on real-world data originating in newswire treebank. We show that the model achieves lower perplexity than previous methods for PCFG minimization, and also is much more stable due to the absence of local optima.
This paper explores the problem of parsing Chinese long sentences. Inspired by human sentence processing, a second-stage parsing method, referred as main structure parsing in this paper, are proposed to improve the pars-ing performance as well as maintaining its high accuracy and efficiency on Chinese long sentences. Three different methods have at-tempted in this paper and the result shows that the best performance comes from the method using Chinese comma as the bounda-ry of the sub- sentence. According to our ex-periment about testing on the Chinese de-pendency Treebank 1.0 data, it improves long dependency accuracy by around 6.0 % than the baseline parser and 3.2 % than the previ-ous best model. 1
The article deals with the problem of linguistic and extra-linguistic description of sport report texts as a variety of sports discourse. Texts of sports report, presenting an implementation of oral public speech, unlike informal texts, are prepared, so their content can be reduced to logical-semantic structure set by a specific theme. Logical-semantic structure of the text is shown in the selection of language means. Compositional and stylistic features of a sports report are based on the rules of its construction. Violation of these rules leads to a variety of errors in the speech of sports journalists. The article provides examples of various violations of linguistic norms of the literary language and the requirements to the utterance. The speech of sports journalists is considered in two forms: monologue and dialogue. The results of the described sports report texts can be used in practice in studying of the subject “Russian language. Culture of speech” and also in teaching Russian as a foreign language (teaching listening skills).
The article discusses the issues of teachers’ speech culture which is an important component of pedagogical culture. Mastering the linguistic norms is one of the topical questions of increasing the teachers’ pedagogical culture which are considered at the professional development courses “Russian as State Language” realized by department of language and literary education of the Chelyabinsk Institute of Retraining and Improvement of Professional Skill of Educators the framework of implementing the Federal targeted program “Russian language”. These courses are designed for various categories of educators: to specialists of educational governances of the Russian Federation; to specialists of educational governances of municipal education; to heads and deputy heads of educational institutions; to teachers of Russian language and literature; to teachers of other subjects, specialists of the preschool education system. The work aimed at developing the educators’ culture of speech should be considered as an important means contributing to the functioning of the Russian language as a state language of the Russian Federation.
<p><b>A.</b> Physical properties. Pixel values are approximate. Differences between male and female faces used here can be noted. Male faces enlarge more with smiling, and have smaller eyes but bigger teeth. Both male and female faces get wider with smaller eyes when smiling. <b>B.</b> Emotional properties from a validation study in adults [<a href="http://www.plosone.org/article/info:doi/10.1371/journal.pone.0129812#pone.0129812.ref030" target="_blank">30</a>]. All stimuli adequately conveyed the desired emotion. Differences between sets were greater than differences between male and female faces within each set. Hit rates, intensity, and arousal ratings are typical of neutral and smiling faces [<a href="http://www.plosone.org/article/info:doi/10.1371/journal.pone.0129812#pone.0129812.ref030" target="_blank">30</a>].</p> <p>(XLSX)</p>
In the article the author makes an attempt to classify radiodiscourse in according to the existing classifications of discourses. Being a type of oral discourse of mass media it possesses its certain characteristics such as unprepared nature of official speech, tendencies to deterioration of a linguistic norm and an intimate manner of speaking. Moreover, clearly defined roles, specified place and context of communication, chronotope, certain typical discursive formulae allow to define radiodiscourse as an institutional discourse. However, extralinguistic conditions in which radiocommunication takes place bring about such peculiarities of radiodiscourse as distance and mediation, mass nature of audience, omnitude, high information rate, direct connection to time, in particular irreversibility, linearity, continuity. Though radiodiscourse has universal characteristics of both oral mediaand institutional discourses, the named peculiarities make it possible to distinguish it in discourse diversity, where it has a certain quite independent position.
For a resource poor language like Hindi, it becomes very difficult to bracket a noun sequence using approaches which are only based on corpus or lexical database. For semantic knowledge, power of both type of resources is needed to be combined. Therefore, affinity in between two nouns is preferred to be measured using backoff association which is the combination of lexical and conceptual association. Also, syntax is important for this task. But syntactic rules do not work for the compound nouns which is a special case of noun sequences and it may also occur as the sub-sequence. Using hybrid approach, accuracy of 86.33% has been obtained.
Many algorithms for natural language processing rely on manual feature engineering. However, manually finding effective features is a labor-intensive task. Moreover, whenever these algorithms are applied on new types of content, they do not perform that well anymore and new features need to be engineered. For example, current algorithms developed for Part-of-Speech (PoS) tagging of news articles with Penn Treebank tags perform poorly on microposts posted on social media. As an example, the state-of-the-art Stanford tagger trained on news article data reaches an accuracy of 73% when PoS tagging microposts. When the Stanford tagger is retrained on micropost data and new micropost-specific features are added, an accuracy of 88.7% can be obtained. We show that we can achieve state-of-the-art performance for PoS tagging of Twitter microposts by solely relying on automatically inferred distributed word representations as features and a neural network. To automatically infer the distributed word representations, we make use of 400 million Twitter microposts. Next, we feed a context window of distributed word representations around the word we want to tag to a neural network to predict the corresponding PoS tag. To initialize the weights of the neural network, we pre-train it with large amounts of automatically high-confidence labeled Twitter microposts. Using a data-driven approach, we finally achieve a state-of-the-art accuracy of 88.9% when tagging Twitter microposts with Penn Treebank tags.
The aim of the article is to examine the usefulness of etiquette books or manuals as sources for reconstructing the pragmatic history of Spanish. Etiquette books recorded the codes of behaviour, both verbal and non verbal, of bourgeois society and so served to perpetuate the established linguistic norms. They also testified to certain changes that appeared in the period being studied, that is to say, the second half of the 19th century. I evaluate this kind of source by contrasting data pertaining to the speech act of requesting recorded in etiquette books with that taken from other text types from the period, accessible via CORDE. I conclude that the usage suggested in etiquette books is adhered to but only by certain social groups. Such books provide valuable data about courtesy and speech acts from previous times, but need to be combined with other research methods.
Following recent trends on hybridization of machine translation architectures, this paper presents an experiment on the integration of a phrase-based system with syntactically-motivated bilingual pairs, namely the so-called catenae, extracted from a dependency-based parallel treebank. The experiment consisted in combining in different ways a phrasebased translation model, as typically conceived in Phrase-Based Statistical Machine Translation, with a small set of bilingual pairs of such catenae. The main goal is to study, though still in a preliminary fashion, how such units can be of any use in improving automatic translation quality.
Rich and comprehensive academic oeuvre of Professor Zaręba is a result of his research carried not only in the fi elds of dialectology, onomastics, lexicography and lexicology but also sociolinguistics, the last being the subject of the present paper. Alfred Zaręba’s deep interest in sociolinguistics was a consequence of his careful observations of the changes taking place in the contemporary language, analysed by him not merely from the perspective of the system itself but, above all, in the wide context of social and cultural factors. One of these issues was the linguistic norm, and precisely, its considerable loosening. He was also interested in the problem of so called literary “micro-languages”, both in Poland and abroad. His academic intuition turned out to be right as these phenomena belong to the most crucial research fi elds at present.
Previous studies indicate that emotion regulation may occur unconsciously, without the cost of cognitive efforts; and that conscious acceptance effectively reduces the emotional consequences of negative events. However, it has yet to be determined how conscious and unconscious acceptance strategies differ in behavioral and physiological consequences of emotion regulation. As unconscious regulation occurs with little cost of cognitive resources, the current study hypothesizes that unconscious acceptance regulates the emotional consequence of negative events more effectively compared to conscious acceptance. Subjects were randomly assigned to conscious acceptance, unconscious acceptance and control conditions. A frustrating arithmetic task was used to induce negative emotion. Emotional experiences were assessed by the positive affect and negative affect scale(PANAS) while emotion-related physiological activation was assessed by the heart-rate reactivity. The results showed that unconscious acceptance produced less reductions of positive affect ratings compared to conscious acceptance during frustration. In addition, both conscious and unconscious acceptance strategies significantly decreased emotion-related heart-rate activity(to a similar extent) in comparison with the control condition. Moreover, heart-rate reactivity showed a trend of positive correlation with negative affect rating and a trend of negative correlation with positive affect rating during frustration compared to baseline phases. Thus, unconscious acceptance is not only able to decrease emotion-related physiological activity, but also able to produce better emotional experiences compared to conscious acceptance. This suggests that it is practically important to consider unconscious acceptance for emotion regulation in real-life settings.
The Analects of Confucius records a large number of examples of character criticism.Character criticism in The Analects of Confucius is the efficient way that Confucius expounds their views,which cultivate and educate people.Character criticism is consistent with principle of education evaluation that Confucius makes no social distinctions in teaching and teaches students in accordance with their aptitude.According to each student's personality,he analyzes and evaluates the full range of the students character to inspire them.At the same time,the evaluation method of The Analects of Confucius is flexible and effective.The main methods of character criticism in The Analects of Confucius are personal evaluation,article evaluation and ranking,metaphor,category and word rating.
У статті розглянуто різні випадки використання неадаптованих англіцизмів у назвах організацій. Проаналізовано специфіку їх функціонування як частини складних багатокомпонентних ергонімів. Наголошено на необхідності узгодженості та семантичної точності у використанні таких одиниць. (The article deals with different cases of non-adapted Anglicism usage in organization names. Their functioning specificity as a part of complex many-component ergonyms is analyzed. In such situations graphical adaptation facilitates perception of the name as a whole. Consistency and semantic correctness necessity is stressed in usage of units like these. Usage of words borrowed from other languages in their English phonetic form is considered to be inadvisable. They are difficult to understand and violate the linguistic norms.)
This paper is interested in the assisted interrogation of lexical databases designed according to the LMF standard (Lexical Markup Framework) ISO-24613. The proposed solution is based on a requirement-based lexical web service generation approach that makes easier the task of engineers when developing NLP (Natural Language Processing) systems. Using this approach, the developer will not deal with the database content or its structure. Also, he will not use any language query. For each lexical requirement, we generate through a lexical web service for interrogating LMF databases. Also we generate with it its enhanced WSDL description file that we have enriched by semantic data describing data categories used in the requirement specification.
In a multi-lingual country like India where every language has its own phonological system, there is a need for language specific articulation test. Although in recent times there has been increasing awareness among parents for early intervention in children with articulation problems within the regional areas of the nation, the availability of articulation tests in the regional languages is very limited. The present study makes a preliminary attempt at developing an assessment tool to assess the articulatory skills of Tulu speaking children, who form a significant population in South India. Word list was developed based on familiarity rating and was administered on 50 children, aged 3-8 years. The target speech sounds were embedded in words which were presented in picture form to elicit responses from the participants. The responses were analysed qualitatively and in terms of production accuracy across age groups.
Parsing models for all Universal Depenencies 1.2 Treebanks, created solely using UD 1.2 data (http://hdl.handle.net/11234/1-1548). To use these models, you need Parsito binary, which you can download from http://hdl.handle.net/11234/1-1584.
The article introduces two internet sources designated to the study of Older Czech language (13th to 18th centuries); both have been designed and run by The Department of Language Development at The Institute of the Czech Language at the Academy of Sciences of the Czech Republic. The first source, Vokabulář webový [Web Vocabulary] (http://vokabular.ujc.cas.cz), makes texts, images and audio materials available to the study of Older Czech language. The accessible materials are, primarily, both modern and historical dictionaries, amongst which the most salient is the, gradually growing, Elektronický slovník staré češtiny [Electronic Old-Czech Vocabulary] that treats Old-Czech lexicon from the dawn of Czech language to the end of the 15th century. Furthermore, Vokabulář includes electronic editions of the works originating in the period from the 13th century to the beginning of the 19th century, presented both as continuous texts and in the corpus version; digitalized copies of Older-Czech grammar books; basic scientific literature; audiobooks of Older-Czech texts; and software tools utilized for the work with historical texts. The second source is Lexikální databáze hu-manistické a barokní češtiny [Lexical Database of Humanistic and Baroque Czech] (http://madla.ujc.cas.cz). It records the Czech vocabulary of the 16th to 18th centuries based on the excerption of the authentic contemporary texts (both old prints and manuscripts): Lexical database illustrates the Czech vocabulary with direct quotations, including stating the source. Thus, Lexical Database partly substitutes the missing Czech vocabulary of the mentioned period.
The article introduces two internet sources designated to the study of Older Czech language (13th to 18th centuries); both have been designed and run by The Department of Language Development at The Institute of the Czech Language at the Academy of Sciences of the Czech Republic. The first source, Vokabulář webový [Web Vocabulary] (http://vokabular.ujc.cas.cz), makes texts, images and audio materials available to the study of Older Czech language. The accessible materials are, primarily, both modern and historical dictionaries, amongst which the most salient is the, gradually growing, Elektronický slovník staré češtiny [Electronic Old-Czech Vocabulary] that treats Old-Czech lexicon from the dawn of Czech language to the end of the 15th century. Furthermore, Vokabulář includes electronic editions of the works originating in the period from the 13th century to the beginning of the 19th century, presented both as continuous texts and in the corpus version; digitalized copies of Older-Czech grammar books; basic scientific literature; audiobooks of Older-Czech texts; and software tools utilized for the work with historical texts. The second source is Lexikální databáze hu-manistické a barokní češtiny [Lexical Database of Humanistic and Baroque Czech] (http://madla.ujc.cas.cz). It records the Czech vocabulary of the 16th to 18th centuries based on the excerption of the authentic contemporary texts (both old prints and manuscripts): Lexical database illustrates the Czech vocabulary with direct quotations, including stating the source. Thus, Lexical Database partly substitutes the missing Czech vocabulary of the mentioned period.
The various views to the sovietness and the opposite tendencies of desovietization and resovietization are observed in different spheres of modern Russian society, including in language. This article is devoted to the consideration of the (anti-) sovietness in modern Russian discourse and meta-discourse about modern Russian language. In the modern Russian language there is a clear tendency to deviate from the previous linguistic norms, and the substandard units widely intruded in the literary language. This article analyzes the desovietization shifts in the system of functional styles, language play and the parodic use of sovietisms as a precedent text. As the desovietization of modern Russian language and the violation of norms consistently increases, the protection of language purity is called for to a great extent. Under puristic approach toward Russian language inviolability of linguistic norms and ideality of language of the past are emphasized, and any language innovations are rejected. Especially the aspiration to the ideal language and non-recognition of language variation remind of the Soviet puristic ideology. Soviet puristic ideology was aimed at compliance of spontaneous oral speech with standardized written language and thereby at the unification of the mass language use.
Children entering institutional education can activate the linguistic and non-linguistic norms they bring from home. These norms, however, are very diverse in their nature and children from the same age-group remain at different levels in terms of language-acquisition. In our paper we seek to identify where and what sort of problems may arise in acquiring the mother-tongue, what factors may hinder the acquisition of the first language, what are the symptoms of backwardness in the field of linguistic–communication and which areas measuring tests tend to focus upon. Finally by presenting an indication system we would like to show what opportunities observation may have in purposeful development.Keywords: native language acquisition, linguistic competence, measurementof communication skills, observation of children’s language performance.
This work describes a system that performs morphological analysis and generation of Pali words. The system works with regular inflectional paradigms and a lexical database. The generator is used to build a collection of inflected and derived words, which in turn is used by the analyzer. Generating and storing morphological forms along with the corresponding morphological information allows for efficient and simple look up by the analyzer. Indeed, by looking up a word and extracting the attached morphological information, the analyzer does not have to compute this information. As we must, however, assume the lexical database to be incomplete, the system can also work without the dictionary component, using a rule-based approach.
The article discusses a number of relatively little-known aspects of Marcus Tullius Cicero’s linguistic views. Special attention is paid to his ideas about language learning, issues of purity of language, as well as linguistic norms and anomalies. These conclusions are of interest to researchers in the field of ancient linguistic thought.
This paper examines L2 learners’ familiarity and correct use of formulaic sequences in English. In particular, the extent to which L2 proficiency level and collocational frequency affected L2 learners’ knowledge and use was investigated. Thirty L2 learners of English were tested on their correct use of 32 formulaic sequences in English, made of up ‘Verb + out’ and asked to rate their familiarity with these formulaic sequences. The results showed that the familiarity ratings given by the low L2 proficiency participants showed a significant positive correlation with their accuracy scores, but no such correlation was found for the high proficiency group. Familiarity ratings correlated positively with both word frequency and formulaic sequence frequency, but neither word frequency nor formulaic sequence frequency showed a significant correlation with accuracy of use, although formulaic sequence frequency showed a trend toward a positive correlation. Effects of frequency and decomposability on various types of formulaic sequences are discussed, along with the disparity in L2 learners’ exposure to formulaic sequences and their ability to use them accurately.
Wide-coverage resources for lexicalized grammars have been obtained by converting the existing treebanks into collections of derivations. Additional annotations to the source treebank can be used to improve these derivations. A treebank annotation called the NTT treebank was used for this paper to improve a CCGbank for Japanese. The source treebank of the CCGbank itself is created by automatically converting chunk-dependencies, but the CCGbank contains errors caused by noisier phrase structures and a lack of linguistic information, which is difficult to represent in chunk-dependency. The NTT treebank provides cleaner trees and functional and semantic information, e.g., coordinations and predicate-argument structures. The effect of the improvement process is empirically evaluated in terms of the changes in the dependency relations extracted from the resulting derivations.
One avenue for supporting the continued use and revitalization of endangered languages in the current, pervasively computerized world is the creation of computational models of the often rich and complex morphology of these languages. Such computational models can be used as a basis for creating a suite of reader’s and writer’s tools, including e.g. (1) an intelligent electronic dictionary that combines the computational model and a lexical database allowing for linking any inflected form with the appropriate dictionary entry, as well as the generation of word paradigms, (2) an intelligent computer-aided language learning application (ICALL) that allows for the dynamic generation of large numbers of exercises combining the entire core vocabulary (up to several thousand of the most common words) and a substantially smaller set of exercise templates, and (3) a spell-checker that supports adherence with one or more existing orthographical conventions, and thus the production of good-quality texts. Importantly, these tools can be made publicly available over the Internet and integrated as part of general software applications such as web browsers and word processors, to be used with little or no cost by any speakers or language-learners in the respective communities as well as any researchers, anywhere – instead of remaining on an individual researcher’s computer drive or on a library bookshelf. Based on our recent experiences on trying out various practical approaches in developing computational morphological models for Plains Cree and Northern Haida, using Finite-State Transducer (FST) technology (Beesley & Karttunen, 2003), once one gains access both to (a) a comprehensive set of full word paradigms, for every possible paradigm type, and (b) an accompanying extensive electronic lexical resource with coding indicating the relevant paradigm type, we have been able to create surprisingly rapidly, potentially within only several months, initial but already full-fledged FST models that can be readily adapted into the aforementioned software tools (1-3). Nevertheless, these first versions will certainly benefit from further work, where one cannot do without the active participation of the language community. However, we will demonstrate how, when a researcher or community linguist pays careful attention in their lexical documentation work on the systematic and detailed coding of the morphological characteristics of the vocabulary in some structured electronic format (e.g. when using software such as ToolBox), they will at the same time facilitate the rapid initial development of computational tools which will make benefits of their work available to the entire community. References Beesley, Kenneth R. and Lauri Karttunen. 2003. Finite State Morphology. Stanford (CA): CSLI Publications.
While cognitive models of the design process have long dominated, many design innovation approaches advocate the importance of exploring affective concepts such as emotion, meaning and lived experiences in the creation of innovations. We suggest the capacity to think abstractly – to question, make connections and broaden understanding based on affect and meaning – is a fundamental skill for the abductive problem solving characteristic of expert designers. There are, however, few tools to promote questioning and reflection based on affect within the design innovation process. We see a need for such tools in design innovation workshops, particularly for non-designers who are less experienced with this type of thinking. We prototype a novel creativity tool for exploring affect within design innovation processes. It utilizes Affect Control Theory's dictionaries of affective meanings for social events to explore affective space. The dictionaries contain standardized affective ratings for a range of concepts. These ratings allow the linking of concepts that have similar affective properties. The initial creativity tool prototype is illustrated within Dorst's (2015) Frame Creation design innovation method. We envisage the tool being one tool among a range used for the analysis of themes and the development of frames within design innovation processes. affect; design innovation; affective meaning; creativity tool In this paper we propose a novel affective creativity tool to assist in the design innovation process. Several cognitive heuristic tools exist to act as prompts in the design ideation phases (see, Daly, Yilmaz, Christian, Seifert, and Gonzalez (2012), for a review). Here, we stress the importance of understanding the role of affect in the design process, and present a tool designed to facilitate the exploration affect within design innovation workshops. This tool is in the early stages of development, and while we are encouraged by a number of the simulated examples shown in this paper, we recognize the real value of the tool will only be established through its use and evaluation in range of design innovation contexts.
We address some theoretical and practical issues relating to generation, processing, and management of Translation Corpus (TC) in Indian languages, which is developed in a consortiummode project (ILCIII) 1 under the DeitY, Govt. of India. Issues are discussed here for the first time keeping in mind the ready application of TC in various domains of computational and applied linguistics. We first define what is a TC; describe the process of its construction; identify its features; exemplify the processes of text alignment in TC; discuss methods of text analysis; propose for restructuring of translational units; define the process of extraction of translational equivalents; propose for generating bilingual lexical database and TermBank from a structured TC; and finally identify areas where a TC and information extracted from it may be utilized. Since construction of TC in Indian languages is full of hurdles, we try to construct a roadmap with a focus on techniques and methodologies that may be applied for achieving the task. The issues are brought under focus to justify the work that generated TC for some Indian languages for future reference and application. 1. What is a Translation Corpus? Theoretically, a Translation Corpus (TC) suggests that it contains texts and their translation. It is entitled to include bilingual (and multilingual) texts as well as texts that may fit under translation. A TC, by virtue of its character and composition, is made of two parts: a text from a source language (SL) and its translation from a target language (TL) (15) (24), (39). Although, a TC is normally bilingual and bidirectional (28), it can be multilingual and multi� directional as well (37), as it actually happens in case of the ILCII and ILCIII projects for the Indian language s. In these two projects a new strategy is adopted where Hindi is treated as the only SL and several other Indian languages are treated as the TL (Fig. 1). The issue of multidirectionality can be understood if all the target languages can establish linguistic links with each other as they are linked up with SL. Since the ILCII TC has not tried to venture into this direction, it makes sense to keep the present discussion confined within a scheme of bilingualism and bidirectionality, with, for examp le, Hindi