Papers reviewed and determined not to be word norm studies. Use the flag icon to report errors or suggest re-inclusion.
16504 papers
M. Yamazaki, A. W. Ellis, C. M. Morrison, and M. A. Lambon Ralph in 1997 (see record [rid]1997-05966-008[/rid]) demonstrated that written and spoken age-of-acquisitions had a stronger effect on the naming latency of single Kanji words than any other variable including familiarity. The present study was designed to reanalyze M. Yamazaki, et al.'s data, using the ratings of written and spoken age-of-acquisitions and visual and auditory familiarities taken from the Nippon Telephone and Telegram Corporation lexical database. This analysis showed that visual familiarity exerted a stronger independent effect on naming latency than two types of age-of-acquisitions. (PsycINFO Database Record (c) 2016 APA, all rights reserved)
This thesis investigates the perceptual categories associated with contrasting pitch accents in Tokyo Japanese. Various aspects of pitch movements throughout the words are systematically varied in order to determine what aspects of the pitch contour affect categorization of words on the basis of accent. In addition, this thesis also investigates the effect on categorization of loss of pitch due to lack of voicing in various parts of the word. The model determined by these studies reveals two more-or-less orthogonal perceptual dimensions; pitch alignment with speech segments determines accent location and the amount of pitch drop determines accent presence. This study also investigates how to quantify the distinctive function of pitch accent in a way which incorporates the frequency of the contrasting items, as well as the peculiar category structure of accents. This model was applied in the analysis of a large-scale lexical database, revealing many irregularities in the distribution and use of accents. Comparing this quantification of the lexical use of accent with the perceptual experiments shows that accent-location detection is functionally more fundamental than accent-presence detection in short, 2-mora words. (PsycINFO Database Record (c) 2016 APA, all rights reserved)
This paper is a study of outside factors that may have a negative impact on a parser’s accuracy. The discussed phenomena can be divided into two classes, those related to treebank design, and those related to morphological tagging. Although the scope of the paper is limited to the Prague Dependency Treebank, two particular taggers and one particular parser, we believe that our observations may be of interest to future treebank designers. Anyway, since training parsers is usually not the only purpose of building treebanks, other good reasons may prevail and push aside the parser’s preferences. For that case, we suggest how the parser may use nasty input. More fine-grained elaboration of these workarounds and their evaluation is a matter of future work.
This paper describes a comparative application of Grammar Learning by Partition Search to four different learning tasks: deep parsing, NP identification, flat phrase chunking and NP chunking. In the experiments, base grammars were extracted from a treebank corpus. From this starting point, new grammars optimised for the different parsing tasks were learnt by Partition Search. No lexical information was used. In half of the experiments, local structural context in the form of parent phrase category information was incorporated into the grammars. Results show that grammars which contain this information outperform grammars which do not by large margins in all tests for all parsing tasks. It makes the biggest difference for deep parsing, typically corresponding to an improvement of around 5%. Overall, Partition Search with parent phrase category information is shown to be a successful method for learning grammars optimised for a given parsing task, and for minimising grammar size. The biggest margin of improvement over a base grammar was a 5.4% increase in the F-Score for deep parsing. The biggest size reductions were 93.5% fewer nonterminals (for NP identification), and 31.3% fewer rules (for XP chunking)
This paper describes a national cooperational project STO, which has the aim of developing a large-scale Danish lexical database for computational use. We discuss some organisational aspects of the project and present the current project structure and main activities. Further, we discuss in more detail some of the linguistic issues that have required thorough consideration before a large-scale encoding could be initiated, encompassing topics such as the morphological encoding of compounds and proper names, as well as the syntactic encoding of there-constructions, phrasal verbs and reflexive verbs.
The redundancy is a method which strengthens the transmitted message in language. And the redundancy has two ways: expanding syllable and explanatory repetition. Well, the repeated faulty wording is a wordiness of the sentence structure, which should be deleted. So, retaining the redundancy is a seeking for superfluity, and it attributes the success to the sufficient principle of transmitting message. Leaving the redundancy out is a seeking for simplicity, and it attributes the success to the economical principle of expression. So it is a problem of how to hold the right simplicity. Nowadays there is a tendency that the demand is too strict for the linguistic norm. Now it may begin with the distinction between the redundancy and the repeated faulty wording in two ways: one is linguistic sensation and another is how to express. But it sometimes is difficult to differentiate redundancy from repetition.
XML(eXtensible Markup Language)is a standard which was issued by W3C(wolld wide web consortium) in February,1998.It defines the data structure by means of an open self-description.Also a group of rules of XML can be used to set up Markup Language in accordance with specific applied fields.Furthermore,based on XML,cnXML is a linguistic norm of electronic business affairs,which is keeping with the commercial habits,tradition and circuit in the continent of China.cnXML supplies a set of unified,flexible,open and extensible exchange pattern of data,which makes all trade partners including such commercial organizations as buyers,sellers,runners and intermediary etc.carry out commercial activities conveniently through Internet.So cnXML is not only a standardized norm but also a premise of electronic business development in China.
We present a procedure for building the lexicon of a transfer-based machine translation system for NPs and PPs from English to Basque. The agglutinative nature of Basque implies the need for morphosyntactic information, so the lexicon was created by automatically extracting this necessary information from a bilingual dictionary and a wide coverage lexical database. The system translates with 83% precision, 18% better than the first approach which used the raw bilingual dictionary.
Note: This is a personal view of the experience gathered, and does not necessarily reflect the opinion of the Floresta team The Floresta Sinta(c)tica http://acdc.linguateca.pt/treebank/ • A collaboration project between VISL (Southern Denmark University) and Linguateca (SINTEF); project leaders: Eckhard Bick & Diana Santos • Bosque: 1,427 syntactically analysed and revised trees (1,405 distinct sentences, 36,408 tokens, ca. 34,256 words), automatically created • Started October 2000, stopped December 2001, some research still being done as of today See Afonso et al. (2002) at LREC'2002
Medical records have been evolving from the traditional paper-based records to digital ones, from the method of dictating reports and transcription to voice recognition systems. The transition to digital operations will not be complete until we have the ability to combine voice recognition with automated indexing of texts. This paper introduces the methods we used to evaluate existing voice recognition software programs and presents NOMINDEX, a system that turns a medical text into MeSH codes, using the French ADM lexical database. Those systems were applied to 28 patient discharge summaries in French, produced after a coronarography, and extracted from the MENELAS corpus of texts. Using the best configuration for voice recognition, the rate of accurate recognition exceeds 98 percent. Among the indexing concepts assigned by NOMINDEX, 25 percent were not pertinent and 12 percent of the relevant concepts were missing. Most errors were related to confusion between common language and medical language, and to the coverage of the ADM lexical database. Best results would be expected with a more comprehensive lexical resource In addition, only 3 percent of the errors generated by inadequate voice recognition that remained in the configuration that performed better, impacted on automatic indexing by NOMINDEX.
Reviewed by: The lexicon-encyclopedia interface ed. by Bert Peeters Adam Głaz The lexicon-encyclopedia interface. Ed. by Bert Peeters. (Current research in the semantics/pragmatics interface 5.) Amsterdam: Elsevier, 2000. Pp. viii, 499. $105.00. A three-decade-old debate on the distinction between lexical and encyclopedic knowledge is reviewed by Bert Peeters (1–52) in an introductory article to the volume. The fourteen papers that follow (of which eleven are novel contributions and three are reprints with newly added afterwords) present to the reader a number of views on the issue, grouped into five parts. In Part 1, entitled ‘Assessments’, Anne Reboul (55–95) discusses James Pustejovsky’s model of the generative lexicon, concluding that it is more valuable for artificial than for natural cognition. Then, Carlos Inchaurralde (97–114) claims that connectionist models corroborate the idea of an integrated ‘lexicopedia’. In a similar vein, John R. Taylor (115–41) considers Ronald W. Langacker’s network model, which does not seek a distinction between the lexicon and encyclopedia, as more promising than Manfred Bierwisch’s two-level model. Part 2, ‘Understanding understanding’, begins with Pierre Larivée’s (145–67) analysis of utterance interpretation which leads the author to conclude that experience-based knowledge and linguistic meaning constitute distinct but interacting levels of representation. Next is Keith Allan’s (169–217) study of quantity implicatures, which are said to reside in the lexicon, encyclopedia being perhaps the site of nonlexical implicatures. Finally, William Croft (219–56) argues for an encyclopedic approach to semantics in interpreting metaphors and metonymies. In Part 3, ‘Words, words, words’, Richard Hudson and Jasper Holmes (259–90), adopting the framework of word grammar, offer a unified view of language as a knowledge network in which the lexicon and encyclopedia cannot be clearly distinguished. In contrast, Eva Born-Rauchenecker (291–316), in her study of Russian transitive action verbs, attempts to formulate a principled distinction between the two, necessary in lexicographic endeavors. This distinction is upheld, albeit in a different way, by M. Lynne Murphy (317–48), who investigates what speakers know about words as language items. Next, Heidi Harley and Rolf Noyer (349–74) show that distributed morphology has an advantage over government and binding models in that it replaces the lexicon with encyclopedia, allowing for a fuller analysis of nominalizations. The framework of distributed morphology is also adopted by Rob Pensalfini (393–431) in his inquiry into a number of phenomena found in Jingulu, a language of Northern Australia. His contribution is preceded by Joseph Hilferty’s (377–92) suggestion that an interactionist view of the relationship between grammar, the lexicon, and encyclopedia is more adequate than modular conceptions of language. These two papers constitute Part 4, entitled ‘Grammar’. Part 5, ‘Further afield’, starts with Susanne Feigenbaum’s (435–61) examination of the strategies adopted by students of a reading course, in whose performance world and linguistic knowledge interact. In the last contribution to the volume, Victor Raskin, Salvatore Attardo, and Donalee H. Attardo (463–86) discuss the nature of lexical knowledge in a lexical database called SMEARR. A part from its other assets (author, subject, and language indices; rich bibliographical references; careful editorship; and elegant typesetting), the volume is an extremely valuable publication due to the wide range of issues addressed and (often conflicting) solutions proposed. To the interested reader, it is an up-to-date presentation of the state of the art in the lexicon-encyclopedia interface. Many of the points it raises may also become thought-provoking stimuli for further research. Adam Głaz Maria Curie-Skłodowska University, Lublin Copyright © 2002 Linguistic Society of America
A complex procedure of syntactic annotation of a large text corpus may be helpful in checking a rich descriptive framework (the Praguian Functional Generative Description) that makes it possible to distinguish between the core of natural language, structured in a relatively simple way, and its large periphery with indistinct borderlines. Such a procedure underlies the Prague Dependency Treebank, within which about 20 000 Czech sentences from running texts have been analyzed in their underlying structure; for 2000 sentences also their Topic-Focus structures have been specified. We illustrate the wide range of the phenomena handled, i.e. the syntactic relations proper (arguments and adjuncts), coordination, topic-focus articulation, word order, deletion, positions of focusing particles, morphological categories such as number, tense, modality, their morphemic and analytical means of expression, and so on.
Many recent statistical parsers rely on a preprocessing step which uses hand-written, corpus-specific rules to augment the training data with extra information. For example, head-finding rules are used to augment node labels with lexical heads. In this paper, we provide machinery to reduce the amount of human effort needed to adapt existing models to new corpora: first, we propose a flexible notation for specifying these rules that would allow them to be shared by different models; second, we report on an experiment to see whether we can use Expectation-Maximization to automatically fine-tune a set of hand-written rules to a particular corpus.
Abstract There seems to be an increase in the use of foreignisms in Finnish business translation: a new linguistic norm, which incorporates foreignisms appears to be developing alongside traditional Finnish usage. Yet how justified is the use of foreignisms in business translations? Here is an analysis of a Nokia document: the author explores the status of this text as translation and offers criteria for justified and unjustified foreignisms in Finnish business translation.
Ambiguity resolution in the parsing of natural language requires a vast repository of knowledge to guide disambiguation. An effective approach to this problem is to use machine learning algorithms to acquire the needed knowledge and to extract generalizations about disambiguation decisions. Such parsing methods require a corpus-based approach with a collection of correct parses compiled by human experts. Current statistical parsing models suffer from sparse data problems, and experiments have indicated that more labeled data will improve performance. In this dissertation, we explore methods that attempt to combine human supervision with machine learning algorithms to try and extend accuracy beyond what is possible with the use of limited amounts of labeled data. In each case we do this by exposing a machine learning algorithm to unlabeled data in addition to the existing labeled data. Most recent research in parsing has shown the advantage of having a lexicalized model, where the word relationships mediate knowledge about disambiguation decisions. We use Lexicalized Tree Adjoining Grammars (TAGs) as the basis of our machine learning algorithm since they arise naturally from the lexicalization of Context Free Grammars (CFGs). We show in this dissertation that probability measures applied to TAGs retain the simplicity of probabilistic CFGs along with its elegant formal properties and that while PCFGs need additional independence assumptions to be useful in statistical parsing, no such changes need to be made to probabilistic TAGs. The main results presented in this dissertation are: (1) We extend the Co-Training algorithm (Yarowsky 1995; Blum and Mitchell 1998), a machine learning technique for combining labeled and unlabeled data previously used with classifiers with 2/3 labels to the more complex problem of statistical parsing. Using empirical results based on parsing the Wall Street Journal corpus we show that training a statistical parser on the combined labeled and unlabeled data strongly outperforms training only on the labeled data. (2) We present a machine learning algorithm that can be used to discover previously unknown subcategorization frames. The algorithm can then be used to label dependents of a verb in a treebank as either arguments or adjuncts. We use this algorithm to augment the Czech Dependency Treebank with argument/adjunct information. (3) We extend a supervised classifier for automatically identifying verb alternation classes for a set of verbs so that it can be used on minimally annotated data. Previous work (Merlo and Stevenson 2001) provided a classifier for this task that used automatically parsed text. With the use of learning of subcategorization frames we construct the same type of classifier which now requires text annotated with part-of-speech tags and phrasal chunks. In each of these results we use some existing linguistic resource that has been annotated by humans and add some further significant linguistic annotation by applying statistical machine learning algorithms.
Automatic Text Categorization (ATC) is an important task in the field of Information Access. The prevailing approach to ATC is making use of a a collection of prelabeled texts for the induction of a document classifier through learning methods. With the increasing availability of lexical resources in electronic form (including Lexical Databases (LDBs), Machine Readable Dictionaries, etc.), there is an interesting opportunity for the integration of them in learning-based ATC. In this paper, we present an approach to the integration of lexical knowledge extracted from the LDB WordNet in learning-based ATC, based on Stacked Generalization (SG). The method we suggest is based on combining the lexical knowledge extracted from the LDB interpreted as a classifier with a learning-based classifier, through SG. We have performed experiments which results show that the ideas we describe are promising and deserve further investigation.
The Papillon project aims at building a multilingual lexical database for extracting dictionaries. This paper describes the Papillon monolingual lexie structure with an example and propose some changes for spotted problems.
In this paper, we observe various syntactic information for Korean parsing and propose a method to learn constraints and improve the efficiency of a parsing model by using the constraints. The proposed method has the following three characteristics. First, it improves the parsing efficiency since we use constraints that can prevent the parser from generating unsuitable candidates. Second, it is robust on a given Korean sentence because the attributes for the constraints are selected based on the syntactic and lexical idiosyncrasy of Korean. Third, it is easy to acquire constraints automatically from a treebank by using a decision tree learning algorithm. The experimental results show that the parser using acquired constraints can reduce the number of overgenerated candidates up to 1/2~1/3 of candidates and it runs 2~3 times faster than the one without any constraints.
In this paper we analyze the problems set up in border lands, especially when a confluence of linguistic norms has taken place; an example is what happened in the Kingdom of Murcia along the Low Middle Ages, where settlers of different origins and also of different religion or race, Christians (Castilians and Catalans), Mussulmans or Jews lived together during some periods and followed one another in other time, leaving their traces on the onomastics and the toponymy. Some times the settlers' mark remained in the way of naming, but other times they reduced themselves to translate the names given by preceding settlers. With regard to onomastics, the traditions of each people remained evident and so have transmitted along the centuries; in the XIIIth. century it is very important the Catalan influence, a reflex of the repopulations; in the XIVth. century instead, because of the predominance of the Castilian model, a graphic adaptation of the family names received from preceding stages took place. Analyzing the documentation of that age permits us verify that the life together of peoples and languages enriched the toponomastic stock.
This mainly technological paper first provides a description of the web site called PapiLex. This first part shows how a file containing XML-structured lexical entries can be managed as a lexical database by using the Document Object Model (DOM) API. PapiLex offers the three essential management functions: creation, modification and deletion of a lexical entry. In a second part, two tools for entering Unicode-formatted text are presented: one for browsers having HTML 4 and JavaScript 1.2 capability and one for Microsoft Word. Such tools can be necessary for the minority languages which have no virtual keyboard embedded in the operating systems. 1 Starting point for building a lexical base Inside the Papillon project, the construction of a lexical base for a new language may take several different ways depending on where the author has to start. The following situations may occur regarding the availability of lexical resources1,2: • no dictionary exists, • a paper dictionary exists, • an electronic form of a dictionary exists, • a lexical database exists. In the last two cases, the question is to re-work the existing data so they meet the Papillon format and to fill the remaining fields. Tools are available for recycling electronic dictionary, (e.g. Nguyen 1998). Here, we will suppose that there is no preexisting dictionary or that its existence is limited to a paper dictionary. In such cases, the lexical entries have to be typed entirely. Among the 1: In addition to the existence of lexical resource, the script used for the language has also to be in Unicode and a font has to exist for it. Actually, the scripts of a number of minority languages are not in Unicode at the moment (e.g. Shan, Tai Dam, Mon). For some of them, fonts that really work are still missing as it is the case for Khmer. 2: In case there are existing data, property rights have to be looked at to say the resource is available. different ways in which this question can be handled, we chose a particular approach that consists in creating directly the Papillon formatted base by using generic and multiplatform Internet browsers. Section 2 will present how a standard browser can be used for this task3 (PapiLex mockup) and section 3 will show that a simple JavaScript program can provide a virtual keyboard that produces Unicode text. In section 4, another issue, less directly related to Papillon, will also be presented as it provides a very practical alternative for creating Unicodeencoded entries. It addresses a Windowsspecific tool for typing Unicode text in Microsoft Word when no standard keyboard is existing yet. The software was developed for the Lao language but can be applied to others. 2 The PapiLex mockup
The increasing prestige of medicine as a science, accompanied by the social rise of the doctor, in eighteenth-century France is well documented. What I would like to argue here, however, is that there exists a correspondence between the establishment of medicine as an independent field of study in eighteenth-century France and the increasing use and influence of an autonomous form of medical discourse, namely, the aphorism, in this period. This is not so much a question of the language used by the more renowned doctors of the day but of a form of discourse deeply imbued and associated with medical practice. (It is nonetheless true that certain famous physicians combined medical and literary roles. For instance, Theophile Bordeu intervenes significantly in Diderot's Le Reve d'Alembert, and Vicq d'Azyr, Marie-Antoinette's doctor, was elected to the Académie Française in 1788 in a sort of social consecration or medical discourse, implicitly incorporating his medical figure and figures into the socio-linguistic norms of 'le bon usage’ promoted by the Académie itself.) Yet what interests me particularly here is the insinuation of the medical aphorism itself into other fields of late eighteenth-century discourse, notably those of literature and politics, the traditional domains of the maxim.
The purpose of this research is to investigate the difference on clothing image evaluation in the ratings between men and women. For this study, pilot test was conducted to 50 clothing majored university students to explore the stimulus of `cute`, `casual`, `sexy`, elegant`, intelligent`, `formal`, `romantic`, `individual`, `refined` for the 9 each image styles from the 32 spring wears in fashion magazine 『FARBE』(March, 2000). On the basis of the preliminary survey, the question items explored the 15 pairs of polar adjectives as seven-point Likert Scale. The main survey was preceded 94 female and 111 man of university students from March 13 to 24 in 2000, twice for 7-days interval. There were significant differences between the two sexes each style image ratings, it was found that the female was recorded more ordinary, stable, refined, superior, plain, like than the male for intelligent style. Meanwhile, the intelligent style was evaluated well on in years by female, but male young. The female tended that elegance style was more stable, warm and less young than the male. The cute style was evaluated more light, tender, feminine, young by the female than the male, and the female looked warm while the male cool. The formal style was more stable, unrefined, solid, unfamiliar, dislike, old by the female than the male. The casual style was revealed plain and warm by the female while splender and cool by the male, the female more active, tender, familiar than the male and individual, attractive and poor quality than the female. The sexy style was evaluated more active, good appearance, young than the female, tender than the male and the female dislike a bit while the male like. The female evaluated the refined style for more stable, refined, superior, good appearance and nature than the male. The romantic style was evaluated more like, refined, superior, good appearance nature and familiar by the male, but the female a bit unfamiliar. The individual style was revealed that the female evaluated cool and a bit dislike while the male warm and like, and the male more refined, feminine, young than female.
The performance of cursive script recognition systems may be improved by applying higher level knowledge in the form of syntax or semantics. A fundamental part of such an approach is the creation of a lexical database containing the relevant information. However, to create a semantic lexicon by hand for a large vocabulary is a considerable task, which is a major reason why so many semantic theories fail to scale up from the small, artificial domains in which they were developed. An alternative approach is to use existing sources of semantic information, such as machine-readable dictionaries (which contain definitions and domain information) and text corpora (from which collocations and domain information may be derived). The development of techniques for acquiring semantic knowledge from such resources and applying it to large vocabulary cursive script recognition is described.< <ETX xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">></ETX>
Medical records have been evolving from the traditional paper-based records to digital ones, from the method of dictating reports and transcription to voice recognition systems. The transition to digital operations will not be complete until we have the ability to combine voice recognition with automated indexing of texts. This paper introduces the methods we used to evaluate existing voice recognition software programs and presents NOMINDEX, a system that turns a medical text into MeSH codes, using the French ADM lexical database. Those systems were applied to 28 patient discharge summaries in French, produced after a coronarography, and extracted from the MENELAS corpus of texts. Using the best configuration for voice recognition, the rate of accurate recognition exceeds 98 percent. Among the indexing concepts assigned by NOMINDEX, 25 percent were not pertinent and 12 percent of the relevant concepts were missing. Most errors were related to confusion between common language and medical language, and to the coverage of the ADM lexical database. Best results would be expected with a more comprehensive lexical resource In addition, only 3 percent of the errors generated by inadequate voice recognition that remained in the configuration that performed better, impacted on automatic indexing by NOMINDEX.
This paper presents work which extends previous corpus-based work on training Machine Learning Algorithms to perform Prepositional Phrase attachment. Besides recreating others&apos; experiments to see how algorithms&apos; performance changes with the number of training examples and using n-fold cross-validation to produce more accurate error rates, we implemented our own vanilla Machine Learning Algorithms as a comparison. We also had people perform exactly the same task as the Machine Learning Algorithms to indicate whether the way forward lies in improving Machine Learning Algorithms or in improving the data sets used to train Machine Learning Algorithms. The results from all these experiments feed into our other work transforming the Penn TreeBank into a more useful resource for training Machine Learning Algorithms to do Prepositional Phrase attachment.
This paper describes new default unification, lenient default unification. It works efficiently, and gives more informative results because it maximizes the amount of information in the result, while other default unification maximizes it in the default. We also describe robust processing within the framework of HPSG. We extract grammar rules from the results of robust parsing using lenient default unification. The results of a series of experiments show that parsing with the extracted rules works robustly, and the coverage of a manually-developed HPSG grammar for Penn Treebank was greatly increased with a little overgeneration.
SLIM is a prototype interactive multimedia self-learning linguistic software for foreign language students at beginner-false beginner level. It allows students to work both in an autonomous self-directed mode or in a way of programmed learning in which the process of self-instruction is pre-programmed and monitored. In this latter mode it incorporates assessment and evaluation tools in order to behave as an automatic tutor. It is organized into three basic components: audiovisual materials; a linguistic database recording all language material in text format; the supervisor. Audiovisual materials are partially taken from commercially available courses; the linguistic database is a highly sophisticated classification of all words and utterances of the course, both in written and spoken form, from all possible linguistic aspects. The supervisor is both an attractive, enjoyable and strongly pedagogically based software that allows the user to work on language materials. The most outstanding feature of SLIM is the use of speech analysis and recognition which is a fundamental aspect of all second language learning programmes. We also assume that a learning model can be represented by a finite state automaton made up by a fixed number of possible states – corresponding to the macro and microlevels at which the student's competence may be modelled – each one being internally constituted by the actual linguistic objects of knowledge of the language that make it up.
This paper deals with up-translation - a process of lexical data transformation from any source format to the XML document. Relevant aspects of the XML format and many related technologies are surveyed first. Then, information content enhancement of existing lexical resources is discussed. The last part brings information about up-translation ofthe Dictionary of Literary Czech Language and the way of efficient storage and retrieval of data.
530 SEER, 8o, 3, 2002 I999). Zubova is well aware of the metatextual qualities of Russian postmodernism, and points to the intertextualgames of varioustexts. In addition to the more familiarnames, Zubova introducesher audience to lesser known authors, such as Vladimir Strochkov, Ian Satunovskii, and Vladimir Erl'.Unfortunately, Zubova's studydoes not contain any biographical detailsof the authorsshe quotes. Such an appendixwould be a usefultool for assessingthe spreadof linguisticdeviations, from the point of view of age groups, regional variations and the aesthetic preferences of the poets. It is difficultto assess whether some of the deviations from established linguistic norms were intentional, or derive from the contemporary sloppy usage of Russianlanguage that isparticularlynoticeable in post-Soviet Russianmedia. Krivulinand Shvarts,for example, are philologistsby training,and therefore are more inclined to have playful appropriationof some idioms, or absurd examplesof Soviet newspeak. As Zubova's study demonstrates, numerous poetic experiments reflect on the fluid state of the Russian language itself. In this respect, Zubova's discussion of the satirical elements relating to the concept of gender in contemporary Russian poetry is particularlyrewarding. Zubova's examples from Russian poetry reveal, for example, the uncertaintiesrelatingto gender of animals. Thus, some poets use the feminine form of the noun koshka (cat) with the additional note that it is used in their poem as a noun of masculine gender. Such examples are both amusing and obscure. Zubova suggeststhat contemporary Russian poets are struggling to revive the use of the neuter gender that otherwise has been steadily disappearing from the standard language (p. 301). Many poems quoted by Zubova use Church Slavonic constructions such as esi, bekh,izhe, byst', and sut', to name just a few (pp. 208-39). Another archaicelement that occurs in contemporarypoetry is the Double Nominative Case discussedon pages 368-70. Zubova also refers to the influence of the English language on the contemporary Russian language in relationto the use of nouns as describingwords (pp. 362-68): for example, Rus'-zemlia, shved-koroleva. Zubova'sbook mightbe seen as an attemptto mergelinguisticanalysiswith cultural anthropology (especially Durkheim's theory), since it claims that various linguistic experiments express an archetypal collective conscience (p. 7). The extensive bibliographywill be highly appreciatedby readersof the book, as well as Zubova's reassuring message that the poetic experiments reveal the dynamic process of language evolution and expose the mnemonic abilities of semiotic signs that enable the user to restore forgotten forms and contexts. Department ofFrench andRussian ALEXANDRA SMITH University ofCanterbugy, NewZealand Menzel,Birgit.Biirgerkrieg um Worte. Die russische Literaturkritik derPerestroika. Bohlau Verlag, Cologne, Weimar, Vienna, 200I. xi + 420 pp. Illustrations. Notes. Bibliography.Tables. Index. DM 89.80:?4s.9 I. How do you investigate a type of text like literary criticism? Not literary criticism in the general sense, but in a sense which native English speakers REVIEWS 53I normally do not use, and which the Russians, among others, do? Literary criticism in this particular sense means topical, applied writing that greatly influencesthe generalpublic and explicitlyevaluatesthe worksdiscussed. Two possibilities present themselves. We can analyse the contents for consistencyand substanceof the profferedargumentsand evaluations.Or we can focus on pragmatic aspects of literary criticism (for example, by underscoring, collating, and evaluating statistics illustrating the effects of criticismon the purchasingbehaviourof readers). In her Rostock habilitation thesis, Birgit Menzel chooses neither of the above possibilities.And with good reason, for she does not focus merely on the work of a single literary critic, or on the treatment of a single work by various schools or camps of literarycriticism. Rather, she assignsherselfthe more comprehensive task of presenting an 'overview of Russian literary criticism between I986 and I993' (p. I). The reader hoping for a detailed criticism of criticism in Menzel's book will thereforebe disappointed. What she does offer is a survey of groupings and argument-patternsalong with traditions and developments. In other words, Menzel seeks to grasp the changing structuresof literary criticism as an important form of literary communication underthe conditions of the disintegratingSoviet empire. The enticing path of marketingresearchremainsunfeasiblefor the simple reason that salesrecordsfor the period of the Soviet planned economy and the years which immediatelyfollowed do not at all provide an accuratepictureof what readersreallywanted. In her statisticalassessments,Menzel thereforeconfines herselfto extremely illuminatinginformationon developments regardingthe circulationnumbersfor the...
ABSTRACT: What links the formation of the state, the formation of linguistic norms, and the development of a grammatical tradition of the state language? In France governmental institutions not only help to spread the state language, but also inspire the first analyses of that language. The model for the grammatisation of French, and the context to promote such a process, is the formation of a national judicial system, through the demand that customary law be written down, and then approved by central authorities. Establishing local texts written in the national language required the analysis of that language, and based that analysis on concepts drawn from legal language. Establishing a fixed national norm depended on the creation of a judicial system founded on writing, and on the royal bureaucracy that resulted from that change.
Why is Canadian unity important to democratic pluralism worldwide? Democratic pluralism is the ability of different cultural and language communities to find representation under a single set of democratic institutions, however configured. Although traditional liberal arguments at best ignored culture, in practise, out of a long struggle to eliminate gargantuan prejudices, errors and wrongs, the liberal tradition has created in democratic pluralism a dialectic of culture and liberal politics that resolves the theoretical conundrums so dear to both. Canadian democracy is a monument to success in its capacity to provide dignity, freedom, opportunity, and prosperity to its citizens throughout the polity. Secession, if it takes place in Quebec, puts these achievements at risk, raising the spectre that cultural-linguistic norms, not a mature liberal democracy, will fashion the kind of state that future generations will inherit. Charles Doran examines why Canadian unity is important, what drives Quebec separatism in the American view, the concern that after Quebec succession the rest of Canada could unravel, and the nature of the historical era that has shaped and conditioned secessionist impulse.
In this paper, we present an approach to primarily achieve the semantic interpretation and the region retrieval for an attentive region in a color image. The main components of the system include image feature extraction, indexing process, as well as linguistic inference rules construction and semantic description. Based on these features, each of attentive regions in an image can be described by a global linguistic meaning. The main procedure consists of two parts: forward and recall processes. The forward process primarily performs the linguistic meaning description of objects for an image, and the recall process reconstructs the region image which is the rough mental image of human memory retrieval. Experiments confirm that our approach is reasonable and feasible. bership function) to construct a knowledge base such as human knowledge and experiences [I, 21. A useful data information may be presented by the fuzzy number, and the data presentation of a fuzzy number can be parameterized to simplify the fizzy computation. In accordance with the advantages of indexing process and fuzzy set theory, in this paper, we present an approach to perform a human-based image interpretation. This approach combines image features and linguistic database to describe semantic meaning for image region.
Language deviation is a form that deviates from linguistic norms. Language deviation is very common in language communication it is a creative use of language and is important in rhetoric. This paper shows the formation of language deviation and its humorous effect.
This paper describes Grammar Learning by Partition Search, a general method for automatically constructing grammars for a range of parsing tasks. Given a base grammar, a training corpus, and a parsing task, Partition Search constructs an optimised probabilistic context-free grammar by searching a space of nonterminal set partitions, looking for a partition that maximises parsing performance and minimises grammar size. The method can be used to optimise grammars in terms of size and performance, or to adapt existing grammars to new parsing tasks and new domains. This paper reports an example application to optimising a base grammar extracted from the Wall Street Journal Corpus. Partition Search improves parsing performance by up to 5.29%, and reduces grammar size by up to 16.89%. Parsing results are better than in existing treebank grammar research, and compared to other grammar compression methods, Partition Search has the advantage of achieving compression without loss of grammar coverage.
With the economic development,advertising has become an essential part of our life and its colorful language bears great charm. To make the language new and unconventional, advertisement designers try boldly the device of language deviation to violate linguistic norms. The paper deals with this device in three aspects: vocabulary, syntax and varieties.
Pragmatic ecological Civilization is a new subject in the study of linguistic ecological.The new linguistic norm includes two parts. One is pragmatic ecologist,which means a combination of diversification and overall norm. The other is ecologist of pragmatic action, which puts forward the linguistic ethics with harmonious space--time and green consume. The thought over pragmatic ecological civilization can help us have a new view of linguistic norm dialectically.
This study is one of the consequtive stages of a macro project designed to identify the relation between language skills and functional literacy in Turkish adults with varying educational backgrounds. Within this framework, earlier stages investigated how elementary and high school graduates in the Turkish sočiety could functionally use their language skills. The consideration is that one’s level of functional literacy observed in his linguistic practices involves how successfully he expresses himself through written and oral language, as well as how successfully he comprehends and interprets written and oral discourse in line with the dynamics of the idealized social and linguistic norms in the information society (Baynham, 1995, Barton, 1987). Although schooling is assumed to be providing individuals with the basics of such linguistic skills, the development or maintenance of these skills to meet the demands of the modern society may present individual- or groupspecific characteristics. It is thought that the quality of social interactions, either during or after school, can create a potential for the development of functional language skills. Whatever is gained at school may or may not be long-lasting once the formal education is over. Depending on the level of education and the types of social networks, individuals are assumed to be displaying varying degrees of functional literacy performance. In individuals with limited education, sometimes horizons are broadened through positive social networks, and functional literacy performance, especially in the language use, appears to be more succesful than expected. The results of the previous stages of the project presented supporting evidence for this consideration. This time focusing on a more educated group, the presented study aims to investigate the relations between the written and oral language skills and functional literacy in university graduates who have been working either at the state or private institutions. The subject group in this study consists of 240 university graduate Turkish males and females over 25 years of age. The subjects, grouped according to (a) their sexes, (b) their years of experience at work, (c) the type of institution they presently work at, and finally (d) their majors at university (independent variables), have been given a personal information questionnaire and have been tested on four language skills, i.e., reading, writing, speaking and listening through specially prepared materials. The oral performance of the subjects have been evaluated in terms of (a) standard pronunciation, (b) morphological, syntactic and semantic well-formedness of utterances, (c) use of discourse markers and idea organization, and (d) communicative competence. Their written performance has also been tested in terms of (a) use of Standard language, (b) idea organization and development, (c) syntactic and semantic well-formedness, (d) use of simple and complex sentence structures, and (e) punctuation and spelling. The data have been analyzed with the SPSS programme and the results have been evaluated in the light of linguistic and sociological theories to find out if the level of functional language skills is bound to the nature of independent variables under consideration. It was hypothesized that the level of functional literacy in university graduates would tend to be somewhat higher than that of the lower-level graduates. Results present supporting evidence; however, findings also reveal statistically significant differences among university-graduate subjects, particularly when independent variables are matched with the results of language skills tests.
In this paper we analyze the problems set up in border lands, especially when a confluence of linguistic norms has taken place; an example is what happened in the Kingdom of Murcia along the Low Middle Ages, where settlers of different origins and also of different religion or race, Christians (Castilians and Catalans), Mussulmans or Jews lived together during some periods and followed one another in other time, leaving their traces on the onomastics and the toponymy. Some times the settlers’ mark remained in the way of naming, but other times they reduced themselves to translate the names given by preceding settlers. With regard to onomastics, the traditions of each people remained evident and so have transmitted along the centuries; in the XIIIth. century it is very important the Catalan influence, a reflex of the repopulations; in the XIVth. century instead, because of the predominance of the Castilian model, a graphic adaptation of the family names received from preceding stages took place. Analyzing the documentation of that age permits us verify that the life together of peoples and languages enriched the toponomastic stock.
This talk provides an overview of current work in my research group on the syntactic annotation of the Tubingen corpus of spoken German and of the German Reference Corpus (Deutsches Referenzkorpus: DEREKO) of written texts. Morpho-syntactic and syntactic annotation as well as annotation of function-argument structure for these corpora is performed automatically by a hybrid architecture that combines robust symbolic parsing with finite-state methods (&amp;quot;chunk parsing &amp;quot; in the sense Abney) with memory-based parsing (in the sense of Daelemans). The resulting robust annotations can be used by theoretical linguists, who are interested in large-scale, empirical data, and by computational linguists, who are in need of training material for a wide range of language technology applications. To aid retrieval of annotated trees from the treebank, a query tool VIQTORYA with a graphical user interface and a logic-based query language has been developed. VIQTORYA allows users to query the treebanks for linguistic structures at the word level, at the level of
We present a flexible approach for extracting hierarchical classifications from linguistic data. To this end, the framework of observational logic is introduced, which extends the logic that underlies standard Formal Concept Analysis by allowing disjunctive rules and exclusions. We give a rigorous mathematical characterization of how the chosen rule type affects the structure of the induced hierarchy. The framework is applied to the induction of hierarchical classifications from linguistic databases. The pros and cons of several types of hierarchies are discussed in detail with respect to criteria such as compactness of representation, suitability for inference tasks, and intelligibility for the human user.
This paper presents a method for designing and organizing a multi‐purpose morpheme‐based lexical database for Modern Greek. The authors are in favour of multi‐purpose lexical databases, to avoid a repetition of effort from one application to another, and of morpheme‐based lexica, to achieve flexibility, reusability, expandability, and compact representation of data for future developments. The suggested method for modelling the lexical database in the word‐processing function is the Entity/Relationship model, according to the linguistic theory of Generative Lexical Morphology. In the framework of this model, which depicts rich linguistic information, we can introduce new data structures for storing the morphemes. These new data structures are matrix encoding schemes; one type, called the Cartesian Lexicon, has been designed as a part of our research. The matrix data structures combine the advantages of hash‐tables and tries, which are very popular data structures in supporting machine readable dictionaries. Our system was tested on the Modern Greek language, and demonstrated a satisfactory overall performance in word‐processing. These methods could also be applicable to other languages having morphological systems similar to Modern Greek.
BACKGROUND: The age-related decline of dehydroepiandrosterone (DHEA) has prompted research on its experimental replacement in women. Although no relationship to sexual functioning in healthy women has been shown to date, DHEA replacement has potential for affecting sexual response. METHODS: To investigate DHEA effects, 16 sexually functional postmenopausal women participated in a randomized, double-blind, crossover protocol in which oral administration of DHEA (300 mg) or placebo occurred 60 minutes before the presentation of an erotic video segment. Blood DHEA sulfate (DHEAS) changes, subjective and physiological sexual responses, as well as affective responses were measured in response to videotaped neutral and erotic video segments. RESULTS: The concentration of DHEAS increased 2-5-fold following DHEA administration in all 16 women. Subjective ratings across DHEA and placebo conditions showed significantly greater mental (p < 0.016) and physical (p < 0.036) sexual arousal to the erotic video with DHEA vs. placebo. Positive affect also increased during the erotic video across drug conditions. Vaginal pulse amplitude (VPA) and vaginal blood volume (VBV) demonstrated a significant increase (p < 0.001) between neutral and erotic film segments within both conditions (DHEA and placebo) but did not differentiate drug conditions. CONCLUSION: In sum, increases in mental and physical sexual arousal ratings significantly increased in response to an acute dose of DHEA in postmenopausal women.