Papers reviewed and determined not to be word norm studies. Use the flag icon to report errors or suggest re-inclusion.
18265 papers
Research on insight—the phenomenon of suddenly solving an apparently intransigent problem—has been hampered because stimulus problems have been few, ad hoc, heterogeneous, and difficult to solve. Responding to the need for a larger pool of problems of a similar type and of varying level of difficulty, we report an experiment testing the validity of rebuses as insight problems. A rebus combines verbal and visual clues to a common phrase, such as PAINS (“growing pains”). Solving a rebus requires breaking implicit assumptions of normal reading, similar to the restructuring required in insight. We hypothesized that, the more implicit assumptions are involved, the more difficult the solution. The results of a two-part experiment supported the hypothesis, with participants solving more problems involving one assumption than they did problems involving two or more. Also, rebus performance correlated significantly with self-rated insight and with scores on remote associates, but not with general verbal ability. The findings suggest that rebus puzzles may be a useful source of theoretically grounded insight problems.
Traditional Active Learning (AL) techniques assume that the annotation of each datum costs the same. This is not the case when annotating sequences; some sequences will take longer than others. We show that the AL technique which performs best depends on how cost is measured. Applying an hourly cost model based on the results of an annotation user study, we approximate the amount of time necessary to annotate a given sentence. This model allows us to evaluate the effectiveness of AL sampling methods in terms of time spent in annotation. We acheive a 77% reduction in hours from a random baseline to achieve 96.5% tag accuracy on the Penn Treebank. More significantly, we make the case for measuring cost in assessing AL methods.
Data Oriented Parsing is a natural language processing model that analyses new input based on past experience. The underlying idea is to extract a set of fragment-probability pairs from a given treebank and use these concrete experiences to construct new utterance analyses. Initially, probabilities were based on the fragments' relative frequency of occurrence. This estimator, however, was soon shown to be biased towards large corpus trees [8] and inconsistent [10]. To alleviate the effects of bias on performance a set of heuristic constraints was put in force. Other estimators addressing these issues have since then been proposed. This paper seeks to show that the most commonly used DOP estimators are in fact susceptible to strong size-sensitive bias effects and to present a new estimation algorithm that greatly reduces these effects of bias on performance without complicating the estimation process.
Since norms for vocabulary acquisition in Maltese children do not yet exist, documentation of productive vocabulary acquisition may contribute to establishing a baseline of lexical development. Clinical implications may thus be derived. The current study is a small-scale investigation of the proportions of Maltese and English lexemes in the vocabularies of ten normally-developing Maltese children aged between 12 and 30 months. The participants were primarily exposed to Maltese within their immediate environments, while receiving indirect exposure to English. Outcomes of parental report and language sampling were analysed for evidence of a bilingual dimension in these children's productive vocabularies. Translation equivalents were reported on by parents, but negligible evidence of equivalents emerged in conversational language use. In contrast, lexical borrowings were both reported and sampled. A substantial proportion of English lexemes were reported by the parents in the absence of Maltese equivalents.
Research on the second language acquisition (SLA) of Spanish has identified grammatical structures for which an analysis of errors for second-language (L2) learners is inappropriate (Geeslin, 2003; Geeslin & Guijarro-Fuentes, 2006; Gudmestad, 2006). This is because the norms of use for such structures are changing, and prescriptive grammars do not coincide with actual language use. Thus, in order to examine such sociolinguistically-variable grammatical features in learner language, researchers have shifted to an analysis of the predictors of use of a given variant, rather than an assessment of accuracy (Geeslin, 2000). Investigations following this approach on structures such as copula choice and mood choice have been largely based on written contextualized tasks (WCT), where use is contextualized and participants indicate a preference for one of the two possible variants. The advantage of this type of task is twofold. First, in comparison with grammaticality judgment tasks, participants are not forced to select one (presumably the only) grammatical sentence from the options provided. Consequently, the WCT is more in line with the idea that variation is indeed an acceptable, and even irrefutable, part of native-like speech. Secondly, in comparison with tasks that elicit less directed production, the WCT assures that each participant will respond to the same tokens (both lexically and in terms of the contextual features that predict selection of a given variant) and that each
The Arabic Treebank team at the Linguistic Data Consortium has significantly revised and enhanced its annotation guidelines and procedure over the past year. Improvements were made to both the morphological and syntactic annotation guidelines, and annotators were trained in the new guidelines, focusing on areas of low inter-annotator agreement. The revised guidelines are now being applied in annotation production, and the combination of the revised guidelines and a period of intensive annotator training has raised inter-annotator agreement f-measure scores already and has also improved parsing results.
THE VOICE TEACHER IS REGULARLY BESET WITH CHALLENGES in the studio regarding consonant clusters in sung German, as is the singer who approaches any vocal work in the German language. The reputation of the German language as being consonant rather than vowel oriented is commonly appreciated and justifiable. Statistical studies show that the burden of text intelligibility is carried principally by the consonants, to a greater extent than most languages. A language that can produce lexical items such as entsturzt [ent'∫tYrtst] and kraftstrotzend ['kraft∫trctsent] adopts a strongly marked position among the world's languages with respect to the involvement of consonants in its sound system. These words contain ten and fourteen phonemes respectively, of which only two or three are vowels. The remaining clusters of consonants are samples of the subject of this article. The consonant clusters normally encountered in German phonology will be inventoried and contrasted with English. The material is likely to be familiar to many readers, albeit presented in a different, perhaps more systematic perspective than is normally encountered. The subject of German consonant clusters is best dealt with in terms of phonetic, not orthographic consonants. A firm grasp of the relationship between spelling and pronunciation is naturally also essential. Two or three successive letters may represent a single phoneme, as in [arrow right] /c/ or /x/ [arrow right] /∫/ [arrow right] /k/ [arrow right] /t/ Conversely, a single written consonant may serve to indicate more than one phoneme, as in [arrow right] /ts/ This situation is familiar because it is even more pronounced in English. The word scythe contains two consonant digraphs and two letter-vowels, but phonetically only one diphthong and no clusters at all. Since the greatest challenge in consonant clusters is visual (i.e., orthographic), thinking in terms of phonetic consonants should serve to simplify the matter for a student. German, more than most other languages, has absorbed lexical items from other languages into its own vocabulary, particularly from English, French, and Italian. Thus Duden, the principal lexicographic publisher in modern Germany, devotes an entire book to Fremdworter in its series of dictionaries. The process of lexical transfer is a complex aspect of German linguistics, particularly regarding pronunciation norms. Some words, such as Situation, have been subsumed into the phonological patterning of German, while others have retained the pronunciation of the word in the language from whence it came, or have struck a middle ground, such as Orange and Weekend. This diversity of phonetic transfer gives modern spoken German a particular flavor, and reflects the country's central geographic position in Europe. There are similar examples in English, such as cul-de-sac (where the French pronunciation has been distorted) and naive (which retains the original, although English idiosyncratically employs only the feminine form). This article will confine itself to the consonant clusters that occur regularly in the standard lexis, referring to combinations resulting from foreign influences only when appropriate. It is useful to consider consonant clusters in two quite distinct groups: syllable-interior, and across syllable or word boundaries. Part I of the article will concern itself with the former; Part II (to appear in the March/April 2008 issue), with the latter. Recognition of which group an example belongs to is the first step toward establishing correct pronunciation, and in some cases is necessary to discriminate between two potentially correct pronunciations. Before outlining in tabular form the cluster environments of German and English, it will be useful to consider all the consonantal combinations that are admissible in each language. A detailed theoretical account of the phonotactic rules and constraints of each language will not be necessary for our purposes. …
This study investigated cued odor identification performance with a set of 64 natural common odors (half of edible and half of nonedible stimuli) in three groups of participants: one group of 30 young adults (mean age 25.3 years, range 18–30, SD 3.1) and two groups of older adults—20 young-old (mean age 64.4 years, range 60–69, SD 2.8) and 21 old-old (mean age 74.6 years, range 70–79, SD 2.5). The results showed that 49 of the 64 odors were correctly identified by over 70% of the participants in all groups. The odor identification performance of the young-old adults did not differ from that of the young adults. However, the oldest group showed a significant loss of performance in the task. Women in the young-old group performed better than men, whereas no gender differences were found in the other two age groups. The data obtained in this study will be useful for further perceptual and memory studies conducted in the olfactory modality with young as well as with older participants.
Problems in training behavioral observers to a high degree of interindividual accuracy and intraindividual stability are fundamental concerns in descriptive research, as well as in provisions of behavioral intervention services. This article presents design characteristics of and results from three formative evaluations of an adaptive computerized expert system that shapes observation and recording skills and maximizes both individual coding accuracy and stability. The system, called Train-to-Code, allows instructors or trainers to import their own video source files and to code those videos using any appropriate descriptive behavioral-coding scheme. This generates customized expert reference data that automate subsequent training on the basis of an operant response-shaping instructional design model. Successful training relies on transitions through alternative levels of prompting and feedback designed to optimize ongoing performance until stable expert-equivalent levels of interobserver accuracy are maintained without prompting or feedback.
Previous research found that the duration of segments decreases as children grow older. The development of suprasegmental duration, however, has not been explored. The present study investigated developmental changes in duration of the four Mandarin tones. 5-, 8-, and 12-year-old monolingual Mandarin-speaking children and young adults participated in the study. Tone durations were measured in participants’ production of monosyllabic target words elicited by picture identification tasks. The results were as follows (1) For each tone category, tone duration and variability decreased with age: 5- and 8-year-old children showed significantly longer durations than adults. Tone durations in 12-year-old children approximated adult values. (2) Despite longer durations, adultlike duration patterns across tone categories existed in all children: dipping tones were the longest, followed by rising and level tones, with falling tones being the shortest. (3) Duration differences between the rising and dipping tones became larger as children grew older. The results may be indicative of the general maturation of laryngeal control over age. Although 5- and 8-year-old children have already established lexical contrasts of tone, adultlike phonetic norms are still in the process of development. The developmental data also provide support for a hybrid account of speech production from a suprasegmental perspective.
Statistical parsing of noun phrase (NP) structure has been hampered by a lack of goldstandard data. This is a significant problem for CCGbank, where binary branching NP derivations are often incorrect, a result of the automatic conversion from the Penn Treebank. We correct these errors in CCGbank using a gold-standard corpus of NP structure, resulting in a much more accurate corpus. We also implement novel NER features that generalise the lexical information needed to parse NPs and provide important semantic information. Finally, evaluating against DepBank demonstrates the effectiveness of our modified corpus and novel features, with an increase in parser performance of 1.51%. 1
Reviewed by: Lexicalization and language change Jesús Fernández-Domínguez Laurel J. BrintonElizabeth Closs Traugott. 2005. Lexicalization and language change. In the series Research Surveys in Linguistics. Cambridge: Cambridge University Press. Pp. xii + 207. US $34.99 (softcover). Lexicalization has been customarily defined as “a gradual historical process, involving graphemic, phonological and semantic changes and the loss of motivation” (Lipka 2005:40), and can affect a word in its phonology, morphology, semantics, or syntax. Because it can affect the makeup of virtually any item, it stands as a central phenomenon in language change, and as such it has gathered the attention of scholars for decades. The aim of [End Page 104] Brinton and Traugott’s work is to provide a wide coverage for what has been traditionally considered under lexicalization, as well as to discuss related concepts necessary for its understanding. Lexicalization and language change develops along six chapters and progressively introduces the various conceptualizations given to the processes of language modification. Chapter 1 (pp. 1–31) sets the theoretical context of the book and introduces some basic notions, and Chapter 2 (pp. 32–61) provides a background in terms of definitions and viewpoints for lexicalization. The authors discuss next the relationship between lexicalization and grammaticalization, first in a general fashion in Chapter 3 (pp. 62–88) and then in further detail in Chapter 4 (pp. 89–110). The most relevant contents of the work are exemplified in Chapter 5 (pp. 111–140), and some conclusions and research questions are offered in Chapter 6 (pp. 141–160). Among the concepts introduced in Chapter 1, the notion of lexicon bears a special significance, as there exist various senses to it which must be clarified before attempting a definition of lexicalization (see Aronoff 1989, not mentioned by the authors). To this end, Brinton and Traugott devote several pages to outline holistic vs. componential approaches to the lexicon, to the categories of the lexicon, and to the lexicon viewed as a continuum of productivity, thus laying the conceptual background required for a proper comprehension of the book. This overview is a suitable introduction to the subject also because it is contrasted with concepts like grammar, language change, or productivity, all of which have a bearing on lexicalization and are seen by the authors as a matter of gradation. Brinton and Traugott also offer a summary of the remainder of contents, and set a number of assumptions for a study of language change “from a historical, functionalist perspective” (p. 31). Chapter 2 immerses into lexicalization proper. After a brief introduction, a central section is “Ordinary processes of word formation” (pp. 33–45), a summary of the major devices of contemporary English: compounding, derivation, conversion, back-formation, initialism, etc. Here, Brinton and Traugott rightly note that lexicalization is to be distinguished from word-formation as far as only the latter has the capacity to produce new items in a regular and predictable manner, a discussion picked up later in Chapter 4. Their review proves valuable because it offers the reader the general features of present-day word-formation in a concise and satisfactory manner, even if one can hardly agree with the inclusion of loan translation, root creation, or coinage under word-formation (see Štekauer 2005:214). A subsequent logical step is the indispensable though brief explanation of institutionalization, that is, “the spread of a usage to a community and its establishment as the norm” (p. 45), usually taken as a stage following word-formation and preceding lexicalization (see Bauer 1983:45–48; Hohenhaus 2005). A number of opinions are explained and illustrated here before turning to the core of the chapter: lexicalization as fusion (pp. 47–57) and as increase in autonomy (pp. 57–60). The authors complain of the very little attention that lexicalization as fusion has received from a historical point of view, and define it as “the development of a form from a more complex to a simpler sequence” (p. 47). The present chapter truly represents a deep and up-to-date review of the typology of the phenomenon, given that it covers lexicalization as affecting phrasal and syntactic constructions (p. 48–50), word-formation (p. 50–52), phonological...
Combining Statistical and Rule-Based Approaches to Morphological Tagging of Czech Texts This article is an extract of the PhD thesis (Spoustová, 2007) and it extends the article (Spoustová et al., 2007). Several hybrid disambiguation methods are described which combine the strength of hand-written disambiguation rules and statistical taggers. Three different statistical taggers (HMM, Maximum-Entropy and Averaged Perceptron) and a large set of hand-written rules are used in a tagging experiment using Prague Dependency Treebank. The results of the hybrid system are better than any other method tried for Czech tagging so far.
The article focuses on the issue of territorial marker in the database treating of the Czech lexicon.
Three sides existed whose connection is solved in this thesis. First, it was the Prague Dependency Treebank 2.0, one of the most advanced treebanks in the linguistic world. Second, there existed a very limited but extremely intuitive search tool - Netgraph 1.0. Third, there were users longing for such a simple and intuitive tool that would be powerful enough to search in the Prague Dependency Treebank. In the thesis, we study the annotation of the Prague Dependency Treebank 2.0, especially on the tectogrammatical layer, which is by far the most complex layer of the treebank, and assemble a list of requirements on a query language that would allow searching for and studying all linguistic phenomena annotated in the treebank. We propose an extension to the query language of the existing search tool Netgraph 1.0 and show that the extended query language satisfies the list of requirements. We also show how all principal linguistic phenomena annotated in the treebank can be searched for with the query language. The proposed query language has also been implemented - we present the search tool as well and talk about the data format for the tool. An attached CD-ROM contains the installation of the tool.
Migration to economically more prosperous areas has been an attractive choice for many Appalachians. This paper traces the effects of migration on language variation within one Appalachian family. Through qualitative and quantitative analysis of phonological, morphological, and lexical variables, we draw distinctions between family members who remained in West Virginia and those who migrated to Ohio and Michigan. The data come from interviews with nine members of one southern West Virginia family. Aside from migration status, education is the most influential factor in language variation patterns for migrant and non-migrant speakers. Our findings indicate that Appalachian migrants negotiate their sociolinguistic identities by drawing on the norms both of their family members and of their adopted homes. This phenomenon is not isolated to one family; economic conditions have fostered the introduction of external sociolinguistic norms into Appalachian communities for at least seventy years.
Stanley Kubrick's 2001: A Space Odyssey (1968) has invited an army of commentators and probably encouraged the publication of Arthur C. Clarke's more discursive of the same name shortly after the movie release in April 1968. However, the existing critical discourse on 2001 rarely foregrounds the importance of Clarke's as an independent work with inherent differences from the movie. In fact major science fiction film scholars such as Vivian Sobchak (in Screening Space), Scott Bukatman (in Terminal Identity), and J. P Telotte (in Replications) do not even mention Clarke's in their discussions about the film. Both the movie and the originated in Clarke's short story Sentinel (1948), but Sentinel merely foreshadows the complex conceptual scopes of the works that developed from it. The general critical stance regarding 2001 is a somewhat linear one--from Sentinel to Kubrick's film and then to the novelization of the film by Clarke. Early commentators such as Jeremy Bernstein, Stanley Kauffmann, and Jerome Agel even regarded Clarke's as an explanation of the film, a view which is echoed to some extent by critics like Robert Kolker even in 2006. (1) Again, commentators like David Patterson and Zoe Sofia seem to acknowledge the difference between the and the film and yet end up appropriating the to explain the film. (2) However, a close comparative examination of the and the film clearly shows that Clarke's is neither an explanation nor a novelization of the film but a work existing independently. While Clarke's is rooted directly in the tradition of hardcore science fiction, Kubrick's film subverts all the norms of traditional films to create something unique. On the one hand, Clarke exploits the conventional device of science fictional discourse to contemplate the theme of the existence of higher forms of intelligence in the universe. On the other hand, Kubrick employs a method similar to the transcendental style to bring about an ineffable quality that gives the film a quasi-religious air of mystery. This article contends that though they deal with the same theme, the film and the are the products of two completely different media and should be seen as such. Unlike the common screen adaptations or novelizations, the film and the were created simultaneously; they both function independently of one another, each with its own unique structures, themes, and significance. In Novels into Films (1957), George Bluestone observes that novel and film are both organic--in the sense that aesthetic judgments are based on total ensembles which include both formal and thematic conventions (137). But he also points out that and film are two completely different media, with limits and advantages peculiar to their forms. The aim of a successful film adaptation should not be merely to turn a into a moving version of words on the pages; rather, this is precisely the thing that can never be done. A film and a literary text operate in two distinctly different ways. Suparno Banerjee spells out this concern very clearly: A film is a photo text, a text containing both the visual image and the sound, while the latter [literature] works with words [and ideas] which arouse images but not in the cinematic sense. John Hartley [...] uses the term photopoetry to describe the cinema. Through the use of light--photopoetry (light writing)--it creates images. Media like television (far sight), video (I see), and cinema (movement) [...] share the same property of sight or visual quality [...]. Compared to these audio-visual experiences, literature is, in a sense, an extra-sensory experience [...]. (154) The eye here acts as an instrument which sends the optical impressions of the lexical symbols on the paper to the brain where the real experience takes place. The letters which combine into words, which then form paragraphs and so on, are really sequences of significations of objects which form a succession of images denoting actions, characters, and so on and raise connotative significances in the brain that then work out their meanings and suggestions. …
We have constructed a large scale and detailed database of lexical types in Japanese from a treebank that includes detailed linguistic information. The database helps treebank annotators and grammar developers to share precise knowledge about the grammatical status of words that constitute the treebank, allowing for consistent large-scale treebanking and grammar development. In addition, it clarifies what lexical types are needed for precise Japanese NLP on the basis of the treebank. In this paper, we report on the motivation and methodology of the database construction.
This research shows a new approach and development of a design methodology, based on the perspective of meanings. In this study the design process is explored as a development of the structure of meanings. The processes of search and evaluation of meanings form the foundations of developing this structure. In order to facilitate the use and operation of the meanings, the WordNet lexical database and an existing visualization of WordNet — Visuwords — is used for the process of meaning search. The basic tool used for evaluation process is the WordNet::Similarity software, measuring the relatedness of meanings in the database. In this way it is measuring the degree of interconnections between different meanings. This kind of search and evaluation techniques are later on incorporated into our methodology of the structure of meanings to support the design process. The measures of relatedness of meanings are developed as convergence criteria for application in the processes of evaluation. Further on, the methodology for the structure of meanings developed here is used to construct meanings in a verification of product design. The steps of the design methodology, including the search and evaluation processes involved in developing the structure of the meanings, are elucidated. The choices, made by the designer in terms of meanings are supported by consequent searches and evaluations of meanings to be implemented in the designed product. In conclusion, the paper presents directions for developing and further extensions of the proposed design methodology.
The Arabic Treebank (ATB), released by the Linguistic Data Consortium, contains multiple annotation files for each source file, due in part to the role of diacritic inclusion in the annotation process. The data is made available in both ”vocalized ” and ”unvocalized ” forms, with and without the diacritic marks, respectively. Much parsing work with the ATB has used the unvocalized form, on the basis that it more closely represents the ”real-world ” situation. We point out some problems with this usage of the unvocalized data and explain why the unvocalized form does not in fact represent ”real-world ” data. This is due to some aspects of the treebank annotation that to our knowledge have never before been published. 1.
This paper provides a frame-based account of the inclusion of non-prototypical members in lexical categorization and semantic classification. It proposes that sense extensions based on metaphorical mappings can be viewed as frame-to-frame transfer, constituting an essential part of our lexical knowledge. Adopting the perspective of frame semantics (Fillmore and Atkins 1991), the study helps delimit and anchors the broad notion of 'domain,' a key concept in defining metaphors, into lexically-attested 'semantic frames' in a principled and systematic manner. Most metaphorical or extended meanings, such as the use of mo 'to touch' in wo mo bu qing ta de yong yi 'I don't understand his intension,' are often neglected by the existing databases. Literally, as a verb of touching, mo is classified as a contact verb, but the above usage of mo clearly indicates its affiliation with cognition verbs. By exploring a number of non-prototypical cognition verbs, kan 'to see', xiu 'to smell', mo, and chi 'to eat', this paper aims to propose a mechanism that allows an effective categorization of non-core members into the appropriate verb class by incorporating and redefining metaphorical extensions in a frame-based approach within the framework of frame semantics. In previous studies of lexical semantics, extended meanings are often excluded from the lexical database. For example, both Levin (1993) and Fillmore, Wooters & Baker (2001) fail to grasp the cognition-related meaning of buy 'to believe' as in I don't buy your story and that of catch 'to understand' in I just can't catch the idea of the book. However, the frequent association of the cognition sense with the two verbs still requires an explanation. In an attempt to account for the usages of various non-prototypical Mandarin cognition verbs, we propose that the usages can be viewed as metaphorical in nature and semantically, the conceptual transfer (cf. Lakoff and Johnson 1980) can be redefined as a partial transfer of frame elements from a source domain/frame to a target domain/frame. For instance, the meaning of kan in [wo/Cognizer] kan kan [yao bu yao bang ta /Issue] 'I am considering whether to help him or not,' belongs to the Cogitating Frame (Target Domain) through an extension from the meaning of kan in [wo/Perceiver] kan kan [zhe zhang zi tiao /Phenomenon] 'I take a look at this note,' in the Perception_active Frame (Source Domain). Through a convergence of frame elements, the key frame element in the Cogitating frame, the Issue, can be viewed as the corresponding frame element in the Perception frame, the Phenomenon. Syntactically, the non-prototypical cognition verb (e.g. kan) displays a similar range of lexical and grammatical collocations in its source domain (as a perception verb) as well as in the target domain (as a cognition verb). Based on lexical aspectual properties, it is shown that Mandarin cognition verbs may encode three different event types, viz. activity achievement and state (cf. Hu 2007). Examples are sorted by the three event types and operated in this paper. With clear operational mechanisms under the framework of frame semantics, metaphorically extended meanings of non-prototypical members can be readily included and well represented in lexical databases and verbal classification. The study ultimately provides a unified account of extended verb members in a grammatically-relevant and semantically-motivated way.
Abstract Food is significant beyond its nutritive value and its dietary customs are culturally contextualised. Folklore, the unwritten cultural evidence of a people, presents a stable platform for cultural analysis of oral food cultures. Using a biocultural approach, this study traces folkloristic influences on African indigenous leafy vegetables preference and dietary habits. Folkloristic products with a semiotic dimension are of particular interest. Norms, acts and events that dictate their use are analysed from a sociolinguistic perspective. These studies show that the folklore of the agropastoral Luo abound with useful reference to vegetables; indigenous leafy vegetables are more than just food. Gender, taste, textural preferences, recipe constructs and olfactory attributes of vegetable foods and sectarian taboos are discussed. The argument is that in general vegetable consumption reflects cultural backgrounds and experiences. Sixteen recorded sayings, proverbs, illustrative metaphors, mantras, lexical phrases, tropes and folktales depicting both wrong and right meanings suggest that vegetable foods are a less preferred food. Cultural factors forcefully determine semiotic workings that underlie food consumption and are more imposing largely determining what is palatable and what is not.
In this paper, a new approach for linguistic database summarization is proposed. It is naturally designed to provide to the user synthetic views of groups of tuples over the database. Summaries are represented as concepts which organized into a hierarchy, defining different levels of granularity to intelligently parse the database and refine user queries.
Despite the importance of spoken vocabulary use to improve spoken skills, little has been conducted to describe spoken features of Korean learners. The purpose of this study is to investigate spoken vocabulary use of Korean learners and find out how far they deviate from native speaker norms. For this purpose, 40 Korean college students` spoken interaction data are transcribed and analysed. Using Wordsmith tool with 17,436 words, frequent single words category (modal items, delexical verbs, interactive words, and discourse markers) and frequent multi-words clusters (discourse markers, vagueness & approximation, politeness & face, and hedgeing) are compared with the spoken British National Corpus (BNC). Frequency analysis revealed that among 56 lexical items investigated, 34 items were underused and 17 items are not represented in KLC. It is evident that Korean learners used limited variety and range in those spoken vocabulary use. Based on this result, some suggestions are made to improve Korean learners` spoken vocabulary teaching and learning.
\n Dans cette proposition, nous plaidons pour une meilleure mutualisation des résultats de recherche sur le lexique à travers loutil informatique que représente le Web. Après avoir analysé les conditions de réussite dune telle mutualisation et limportance des normes et standards en ce domaine, nous montrons quelques exemples de réussite dune telle mutualisation tant en lexicographie contemporaine, à travers le Trésor de la langue française informatisé, quen lexicographie historique. Ainsi à travers le DMF (Dictionnaire du Moyen Français), nous explicitons le concept nouveau de lexicographie évolutive et montrons quelques exemples de résultats de recherche qui nauraient pas vu le jour sans sappuyer sur la richesse dexploitation inégalée, rendue possible grâce à son informatisation: \n - en lexicologie, par exemple sur la datation dapparition de sens nouveaux dun lexème dans la langue,\n - en pragmatique, à travers lexemple dune anté-datation de près de deux siècles de lusage de enfin énumératif,\n - ou en morphologie constructionnelle, à travers létude des formations en inr- qui pour certaines furent ensuite abandonnées au profit de formation en irr- (tel inrégulier versus irrégulier).\nToujours dans le domaine de la lexicographie historique, nous montrons lintérêt de mutualiser nos connaissances sur létymologie, tel quil se pratique dans le projet TLF-Etym, ou sur les « mots fantômes », pseudo lexèmes disposant à tort dun statut lexicographique (« ces mots qui nexistent pas »), et les lemmatisations erronées qui se trouvent encore trop souvent dans les dictionnaires historiques et étymologiques français de référence.\nNous terminons enfin par la présentation dun exemple dintégration et de valorisation de données lexicographiques et lexicales au sein du portail lexical du Centre National de Ressources Textuelles et Lexicales (CNRTL, www.cnrtl.fr) qui à travers les quelque 300 000 requêtes quil sert par jour est aujourdhui une magnifique vitrine des résultats de recherche en lexicographie, morpho-syntaxe, étymologie, synonymie, antonymie.\n\n
In this paper we present a corpus representation format which unifies the representation of a wide range of dependency treebanks within a single model. This approach provides interoperability and reusability of annotated syntactic data which in turn extends its applicability within various research contexts. We demonstrate our approach by means of dependency treebanks of 11 languages. Further, we perform a comparative quantitative analysis of these treebanks in order to demonstrate the interoperability of our approach.
It is generally acknowledged nowadays that and are inseparable and there exists a direct connection between a and the used by its members. Culture occupies a prominent position on the foreign teaching agenda for the time being and the role of cultural learning has become one of the essential issues in foreign teaching theory today. There are a lot of definitions of culture suggested by different authors from various perspectives. For us as teachers of English as a foreign one of the most useful approaches in this context might be the definition provided by G. Hofstede who sees as collective programming of the mind which distinguishes the members of one group or category of people from another. M. Seidl proposes to consider a concept of that links it to a oriented analysis that in turn defines in terms of the norms and values shared by the members of a social group. The author states that language proficiency, … is a matter of familiarity with commonly held norms and values which constitute hidden meaning encoded in discourse structures. She believes that when someone learns a foreign and wants to understand another it is not enough to come to terms with another lexical or grammatical code. One has to view the world from a different perspective since speaking another means adopting another point of view.
Abstract This paper offers a model to explain the general observation that lexical items are more often borrowed from a higher status language into a lower status one, than visa versa. Material from Lahore, Pakistan, shows that in casual speech among plurilinguals codeswitching is the norm. In formal contexts, in which there is attention to proper language, educated speakers filter out features which are not part of the standard language. Constraints on language and education in the hierarchical social structure withhold from most speakers of the lower status languages the knowledge necessary to evaluate their own speech in this way, thus allowing features of other languages to become established in their language.
This paper presents a methodology for automatic learning of ontologies from Thai text corpora, by extraction of terms and relations. A shallow parser is used to chunk texts on which we identify taxonomic relations with the help of cues: lexico-syntactic patterns and item lists. The main advantage of the approach is that it simplify the task of concept and relation labeling since cues help for identifying the ontological concept and hinting their relation. However, these techniques pose certain problems, i.e. cue word ambiguity, item list identification, and numerous candidate terms. We also propose the methodology to solve these problems by using lexicon and co-occurrence features and weighting them with information gain. The precision, recall and F-measure of the system are 0.74, 0.78 and 0.76, respectively.
This paper presents recent advances in an established treebank annotation framework comprising of an abstract XML-based data format, fully customizable editor of tree-based annotations, a toolkit for all kinds of automated data processing with support for cluster computing, and a work-in-progress database-driven search engine with a graphical user interface built into the tree editor.
ABSTRACT Using lexical items from Martin Durrell's classification of register variation as a sample, the study investigates how the current advanced monolingual learners' dictionaries of German as an additional language treat such variation and indicate to their users what they consider to be standard usage: how do they set the standard? Abbreviated usage labels as conventionally found in dictionaries for first‐language users are the primary indications, and the dictionaries seldom go beyond such labels. German Standard German is the norm. At its core are unmarked or unlabelled items, while its range extends to include less formal items from everyday use, especially spoken, which are typically labelled umg. or gespr., and more formal items, more particularly found in written usage, which are labelled geh. or geschr. Non‐standard items, if entered as headwords, may be labelled derb or vulgär, veraltet or lit. No one dictionary stands out from the others as setting the standard in terms of treating register variation, and it must be questioned whether learners of German as an additional language would not be better served by more detailed, discursive information on different contexts of use and stylistic levels.
Over the past 15 years, there has been increasing use of linguistically annotated sentence collections, such as the Penn Treebank (PTB), for constructing statistically based parsers.While these parsers have generally been built for engineering purposes, more recently such approaches have been advanced as potentially cognitively relevant, e.g., for addressing the problem of human language acquisition.Here we examine this possibility critically: we assess how well these Treebank parsers actually approach human/child language competence.We find that such systems fail to replicate many, perhaps most, empirically attested grammaticality judgments; seem overly sensitive, rather than robust, to training data idiosyncrasies; and easily acquire "unnatural" syntactic constructions, those never attested in any human language.Overall, we conclude that existing statistically based treebank parsers fail to incorporate much "knowledge of language" in these three senses.
Processing speed (Gs) and working memory (WM) tasks have received considerable interest as correlates of more complex cognitive performance measures. Gs and WM tasks are often repetitive and are often rigidly presented, however. The effects of Gs and WM may, therefore, be confounded with those of motivation and anxiety. In an effort to address this problem, we assessed the concurrent and predictive validity of computer-game-like tests of Gs (Space Code) and WM (Space Matrix) across two experiments. In Experiment 1, within a university sample (N =70), Space Matrix exhibited concurrent validity as a WM measure, whereas Space Code appeared to be a mixed-ability measure. In Experiment 2, Space Matrix exhibited concurrent validity as well as predictive validity (as a predictor of school grades) within a school-aged sample (N=94), but the results for Space Code were less encouraging. Relationships between computer-game-like tests and gender, handedness, and computergame experience are also discussed.
WordNet is a lexical database describing English words and their senses. We propose a method for automatically producing similar resources for new languages by taking advantage of the original WordNet in conjunction with translation dictionaries. A small set of training mappings is used to learn a model for predicting associations between terms and senses. The associations are represented using a variety of scores that take into account structural properties as well as semantic relatedness and corpus frequency information. For evaluation, we created a German-language wordnet, and the data indicate a significantly better coverage and higher precision than previous heuristics. The resulting resources provide not only valuable information for monolingual NLP tasks but also enable a high degree of cross-lingual interoperability. 1
There are many expressive and structural differences between product names and general named entities such as person names, location names and organization names. To date, there has been little research on product named entity recognition (NER), which is crucial and valuable for information extraction in the field of market intelligence. This paper focuses on product NER (PRO NER) in Chinese text. First, we describe our efforts on data annotation, including well-defined specifications, data analysis and development of a corpus with annotated product named entities. Second, a hierarchical hidden Markov model-based approach to PRO NER is proposed and evaluated. Extensive experiments show that the proposed method outperforms the cascaded maximum entropy model and obtains promising results on the data sets of two different electronic product domains (digital and cell phone).
This paper describes the building of a valency lexicon of Arabic verbs using a morphologically and syntactically annotated corpus, the Prague Arabic Dependency Treebank, as its primary source. We present the theoretical account on valency developed within the Functional Generative Description theory. We apply the framework to Arabic and discuss various valency-related phenomena with respect to examples from the corpus. We then outline the methodology and the linguistic and technical resources used in the building of the lexicon. Valency lexicons can find application in automatic parsing as well as in language generation. 1.
What’s the best way to assess the performance of a semantic component in an NLP system? Tradition in NLP evaluation tells us that comparing output against a gold standard is a good idea. To define a gold standard, one first needs to decide on the representation language, and in many cases a first-order language seems a good compromise between expressive power and efficiency. Secondly, one needs to decide how to represent the various semantic phenomena, in particular the depth of analysis of quantification, plurals, eventualities, thematic roles, scope, anaphora, presupposition, ellipsis, comparatives, superlatives, tense, aspect, and time-expressions. Hence it will be hard to come up with an annotation scheme unless one permits different level of semantic granularity. The alternative is a theory-neutral black-box type evaluation where we just look at how systems react on various inputs. For this approach, we can consider the well-known task of recognising textual entailment, or the lesser-known task of textual model checking. The disadvantage of black-box methods is that it is difficult to come up with natural data that cover specific semantic phenomena. 1. Evaluating Meaning Formal methods for the analysis of the meaning of natural language expressions have long been restricted to the ivory tower built by semanticists, logicians, and philosophers of language. It was only in exceptional cases that they made their way directly into open domain NLP tools. Recently, this situation has changed. Thanks to the development of treebanks (large collections of texts annotated with syntactic structures), robust statistical parsers trained on such treebanks, and the development of large-scale semantic lexica, we now have at our disposal systems that are able to produce formal semantic representations achieving
The problem of significance of the knowledge of cultural norms and standards of national communication is discussed. Studying a foreign language implies not only the knowledge of its lexical units, grammar and word combinations but also the knowledge of behavior stereotypes, etiquette norms and rules in different situations of communication. This thesis is analyzed on the comparison of address forms both in Russian and American cultures.
L’objectif premier de ce travail etait de caracteriser les images (et mots correspondants) proposees par Bonin, Peereman, Malardier, Meot et Chalard (2003) en termes de frequence cumulee et trajectoire frequentielle, selon le modele utilise par Bonin, Barry, Meot et Chalard (2004) pour caracteriser les images de Snodgrass et Vanderwart (1980). Correlations et regressions multiples revelent que la frequence cumulee et la trajectoire frequentielle sont les principaux predicteurs de l’âge d’acquisition (AoA) estime ou objectif. La comparaison des deux jeux d’images renforce l’idee selon laquelle ces facteurs, construits objectifs puisque simplement deduits des frequences lexicales de l’adulte (LEXIQUE) et de l’enfant (MANULEX), sont des candidats valides pour remplacer la frequence lexicale et l’AoA. Leur utilisation permet alors de d’echapper aux biais des etudes comportementales denonces par Zevin et Seidenberg (2002) portant sur le choix de la frequence lexicale et l’utilisation de l’AoA, variable de performance. En particulier, frequence cumulee et trajectoire frequentielle ne sont pas correlees. L’implication theorique et methodologique de ces constatations est discutee dans cet article.
Native speakers of languages perceive differences in the acceptability of phrases even when those phrases are both grammatical and novel (previously unseen). We suggest that smoothing, a statistical technique used by natural language processing engineers, provides several candidate mechanisms for investigating this phenomenon. We describe the creation of a large data set of predictions from several smoothing algorithms about the acceptability of unseen grammatical phrases and a novel experimental method for the pairwise comparison of these models. We use this method to compare three smoothing methods and consider the results in light of the differences among the models. We argue that the data support the idea that similarity in this domain is best thought of as a form of asymmetric representational distortion and that the informational basis over which such estimates are made is broad, rather than narrow, as has been previously suggested.