Papers reviewed and determined not to be word norm studies. Use the flag icon to report errors or suggest re-inclusion.
18265 papers
This paper presents a current status of Thai resources and tools for CG development. We also proposed a Thai categorial dependency grammar (CDG), an extended version of CG which includes dependency analysis into CG notation. Beside, an idea of how to group a word that has the same functions are presented to gain a certain type of category per word. We also discuss about a difficulty of building treebank and mention a toolkit for assisting on a Thai CGs tree building and a tree format representations. In this paper, we also give a summary of applications related to Thai CGs. 1
Two of the main corpora available for training discourse relation classifiers are the RST Discourse Treebank (RST-DT) and the Penn Discourse Treebank (PDTB), which are both based on the Wall Street Journal corpus. Most recent work using discourse relation classifiers have employed fully-supervised methods on these corpora. However, certain discourse relations have little labeled data, causing low classification performance for their associated classes. In this paper, we attempt to tackle this problem by employing a semi-supervised method for discourse relation classification. The proposed method is based on the analysis of feature cooccurrences in unlabeled data. This information is then used as a basis to extend the feature vectors during training. The proposed method is evaluated on both RST-DT and PDTB, where it significantly outperformed baseline classifiers. We believe that the proposed method is a first step towards improving classification performance, particularly for discourse relations lacking annotated data.
For centuries, scholars have explored the deep links among human languages. In this paper, we present a class of probabilistic models that use these links as a form of naturally occurring supervision. These models allow us to substantially improve performance for core text processing tasks, such as morphological segmentation, part-of-speech tagging, and syntactic parsing. Besides these traditional NLP tasks, we also present a multilingual model for the computational decipherment of lost languages. 1. Overview Electronic text is currently being produced at a vast and unprecedented scale across the languages of the world. Natural Language Processing (NLP) holds out the promise of automatically analyzing this growing body of text. However, over the last several decades, NLP research efforts have focused on the English language, often neglecting the thousands of other languages of the world (Bender, 2009). Most of these languages are currently beyond the reach of NLP technology due to several factors. One of these is simply the lack of the kinds of hand-annotated linguistic resources that have helped propel the performance of English language systems. For complex tasks of linguistic analysis, hand-annotated corpora can be prohibitively time-consuming and expensive to produce. For example, the most widely used annotated corpus in the English language, the Penn Treebank (Marcus et al., 1994), took years for a team of professional linguists to produce. It is unrealistic to expect such resources to ever exist for the majority of the world’s languages.
We describe a process for converting the Penn Arabic Treebank into the CCG formalism. Previous efforts have yielded CCGbanks in English, German, and Turkish, thus opening these languages to the sophisticated computational tools developed for CCG and enabling further cross-linguistic development. Conversion from a context free grammar treebank to a CCGbank is a four stage process: head finding, argument classification, binarization, and category conversion. In the process of implementing a basic CCGbank conversion algorithm, we reveal properties of Arabic grammar that interfere with conversion, such as subject topicalization, genitive constructions, relative clauses, and optional pronominal subjects. All of these problematic phenomena can be resolved in a variety of ways- we discuss advantages and disadvantages of each in their respective sections. We detail these and describe our categorial analysis of each of these Arabic grammatical phenomena in depth, as well as technical details on their integration into the conversion algorithm. 1.
The paper introduces principles of rescript and lemmatization of lexemes with German origin for Lexical database of Baroque and Humanist Czech made in the Czech Language Institute of the Academy of Sciences of the Czech Republic in Prague.
In this paper we present an experimental toolbox for automatic tree-to-tree alignment based on local classification and alignment inference. The aligner implements a recurrent architecture for structural prediction using history features and a sequential classification procedure. The discriminative base classifier uses a log-linear model which enables simple integration of various features extracted from the data. The Lingua-Align toolbox provides a flexible framework for feature extraction including contextual properties and implements several alignment inference procedures. Various settings and constraints can be controlled via a simple frontend or called from external scripts. Lingua-Align supports different treebank formats and includes additional tools for conversion and evaluation. In our experiments we can show that our tree aligner produces results with high quality and outperforms unsupervised techniques proposed otherwise. It also integrates well with another existing tool for manual tree alignment which makes it possible to quickly integrate additional training material and to run semi-automatic alignment strategies. 1.
The creation of language resources for less-resourced languages like the historical ones benefits from the exploitation of language-independent tools and methods developed over the years by many projects for modern languages. Along these lines, a number of treebanks for historical languages started recently to arise, including treebanks for Latin. Among the Latin treebanks, the Index Thomisticus Treebank is a 68,000 token dependency treebank based on the Index Thomisticus by Roberto Busa SJ, which contains the opera omnia of Thomas Aquinas (118 texts) as well as 61 texts by other authors related to Thomas, for a total of approximately 11 million tokens. In this paper, we describe a number of modifications that we applied to the dependency parser DeSR, in order to improve the parsing accuracy rates on the Index Thomisticus Treebank. First, we adapted the parser to the specific processing of Medieval Latin, defining an ad-hoc configuration of its features. Then, in order to improve the accuracy rates provided by DeSR, we applied a revision parsing method and we combined the outputs produced by different algorithms. This allowed us to improve accuracy rates substantially, reaching results that are well beyond the state of the art of parsing for Latin. 1.
Raven’s Progressive Matrices is a widely used test for assessing intelligence and reasoning ability (Raven, Court, & Raven, 1998). Since the test is nonverbal, it can be applied to many different populations and has been used all over the world (Court & Raven, 1995). However, relatively few matrices are in the sets developed by Raven, which limits their use in experiments requiring large numbers of stimuli. For the present study, we analyzed the types of relations that appear in Raven’s original Standard Progressive Matrices (SPMs) and created a software tool that can combine the same types of relations according to parameters chosen by the experimenter, to produce very large numbers of matrix problems with specific properties. We then conducted a norming study in which the matrices we generated were compared with the actual SPMs. This study showed that the generated matrices both covered and expanded on the range of problem difficulties provided by the SPMs.
Variation is a ubiquitous feature of speech. Listeners must take into account context-induced variation to recover the interlocutor's intended message. When listeners fail to normalize for context-induced variation properly, deviant percepts become seeds for new perceptual and production norms. In question is how deviant percepts accumulate in a systematic fashion to give rise to sound change (i.e., new pronunciation norms) within a given speech community. The present study investigated subjects' classification of /s/ and /∫ / before /a/ or /u/ spoken by a male or a female voice. Building on modern cognitive theories of autism-spectrum condition, which see variation in autism-spectrum condition in terms of individual differences in cognitive processing style, we established a significant correlation between individuals' normalization for phonetic context (i.e., whether the following vowel is /a/ or /u/) and talker voice variation (i.e., whether the talker is male or female) in speech )
In this paper, we present a system that automatically extracts lexicalized tree adjoining grammars (LTAG) from treebanks. We first discuss in detail extraction algorithms and compare them to previous works. We then report the first LTAG extraction result for Vietnamese, using a recently released Vietnamese treebank. The implementation of an open source and language independent system for automatic extraction of LTAG grammars is also discussed. 1
Despite the prevalence of infidelity, there is relatively little research regarding the long-term effects on the children. Combining the views of transgenerational theory with the existing literature on infidelity as a family stressor and infidelity as a trauma, this project examines the potential lasting effects of parental infidelity on the adult child. Study participants included a small sample of adults over the age of 18 who were aware of parental infidelity in their family of origin. Through a series of regression analyses using a moderator model, the researchers found that higher negative self schemata and affect ratings related to the perception of the infidelity were associated with more conservative attitudes toward sex. Suggestions for future utilization of the proposed model and the clinical implications of the findings are also discussed.
This paper investigates whether high-quality annotations for tasks involving semantic disambiguation can be obtained without a major investment in time or expense. We examine the use of untrained human volunteers from Amazon’s Mechanical Turk in disambiguating prepositional phrase (PP) attachment over sentences drawn from the Wall Street Journal corpus. Our goal is to compare the performance of these crowdsourced judgments to the annotations supplied by trained linguists for the Penn Treebank project in order to indicate the viability of this approach for annotation projects that involve contextual disambiguation. The results of our experiments show that invoking majority agreement between multiple human workers can yield PP attachments with fairly high precision, confirming that this crowdsourcing approach to syntactic annotation holds promise for the generation of training corpora in new domains and genres.
Cytotoxic T cell (CTL) covers several subtypes, which are CD8+, CD4 and CD4-CD8-. CTL derives from T cell repertoire in lymphoid hematopoietic stem cells. It matures in thymus and is activated in peripheral lymphoid tissues. Effector CTL kills the target cells by 2 ways. One is apoptotic effect mediated by FasL-Fas pathway and the other one is cytolytic effect mediated by granzymes. CTL has aroused great attention due to its significance in anti-tumor and anti-virus.
We propose an effective approach to automatically identify predicate heads in Chinese sentences based on statistical pre-processing and rule-based post-processing. In the preprocessing stage, the maximal noun phrases in a sentence are recognized and replaced by “NP ” labels to simplify the sentence structure. Then a CRF model is trained to recognize the predicate heads of this simplified sentence. In the post-processing stage, a rule base is built according to the grammatical features of predicate heads. It is then utilized to correct the preliminary recognition results. Experimental results show that our approach is feasible and effective, and its accuracy achieves 89.14 % on Tsinghua Chinese Treebank. 1
This paper proposes a method of correcting annotation errors in a treebank. By using a synchronous grammar, the method transforms parse trees containing annotation errors into the ones whose errors are corrected. The synchronous grammar is automatically induced from the treebank. We report an experimental result of applying our method to the Penn Treebank. The result demonstrates that our method corrects syntactic annotation errors with high precision. 1
The Prague Dependency Treebank (henceforth PDT) is a large collection of texts in Czech. It contains several layers of rich annotation, ranging from morphology to deep syntax. It is unique in its size and theoretical background, especially for a language like Czech, which can be, with regard to the number of its speakers, considered a small language. In this article, we use PDT 2.0 to demonstrate that within real NLP systems, complex annotations may cut both ways. We present several issues that might pose problems when extracting data from PDT, and complex structures in general, and hint on possible solutions.
Proceedings of the Ninth International Workshop \non Treebanks and Linguistic Theories. \nEditors: Markus Dickinson, Kaili Müürisep and Marco Passarotti. \nNEALT Proceedings Series, Vol. 9 (2010), 211-222. \n© 2010 The editors and contributors. \nPublished by \nNorthern European Association for Language \nTechnology (NEALT) \nhttp://omilia.uio.no/nealt. \nElectronically published at \nTartu University Library (Estonia) \nhttp://hdl.handle.net/10062/15891.
This paper presents a new perspective on the origin and development of the Mary-merry-marry merger, the conditioned merger, or neutralization, of mid and low front vowels before /r/ in dialects of North American English. The city of Montreal, Quebec represents one of very few regions in which this merger has not taken hold, despite the fact that a near-complete merger is found in the nearby rural region of Quebec’s Eastern Townships. This paper attempts to shed light on this puzzling geographic distribution using data from archival interviews conducted with Eastern Townshippers born between 1895 and 1915. An acoustic analysis of the vowels before /r/ is presented and compared with data from recent studies of Montreal English. Acoustic analysis of the mean values of the first and second vowel formants shows a great deal of variation in these speakers’ productions of the historically low front vowel before /r/. In some tokens it is clearly merged with the mid vowel, while in others the two phonemes remain clearly distinct. Further, this variation is found both between speakers and in the speech of individuals themselves. Although not entirely homogenous, the speech community does appear to share general norms with regard to which words are or are not merged. These results demonstrate that the merger was not a lexically abrupt sound change. Rather, the results are consistent with a theory of sound change via lexical diffusion, which implies a much longer timeline for this change than previously assumed, suggesting its origins may go back many more generations. As such, it is suggested that the current geolinguistic pattern of the merger may be traced to the different settlement histories of Montreal and the Eastern Townships. This working paper is available in University of Pennsylvania Working Papers in Linguistics: http://repository.upenn.edu/pwpl/ vol16/iss2/3 U. Penn Working Papers in Linguistics, Volume 16.2, 2010 Lexical Diffusion in the Early Stages of the Merry-Marry Merger
This paper presents a method for the automatic detection and correction of malapropism errors found in documents using the WordNet lexical database, a search engine (Google) and a paronyms dictionary. The malapropisms detection is based on the evaluation of the cohesion of the local context using the search engine, while the correction is done using the whole text cohesion evaluated in terms of lexical chains built using the linguistic ontology. The correction candidates, which are taken from the paronyms dictionary, are evaluated versus the local and the whole text cohesion in order to find the best candidate that is chosen for replacement. The testing methods of the application are presented, along with the obtained results.
The article presents the Russian-Czech lexical database which is being compiled in the Institute of Slavonic Studies AS CR, defines its content, structure and functioning.
Background: Matrix-assisted laser desorption ionisation time of flight mass spectrometry (MALDI TOF-MS) allows the identification of most bacteria and an increasing number of fungi. The potential for the highest clinical benefit of such methods would be in severe acute infections that require prompt treatment adapted to the infecting species. Our objective was to determine whether yeasts could be identified directly from a positive blood culture, avoiding the 1-3 days subculture step currently required before any therapeutic adjustments can be made. Methodology/Principal Findings: Using human blood spiked with Candida albicans to simulate blood cultures, we optimized protocols to obtain MALDI TOF-MS fingerprints where signals from blood proteins are reduced. Simulated cultures elaborated using a set of 12 strains belonging to 6 different species were then tested. Quantifiable spectral differences in the 5000-7400 Da mass range allowed to discriminate between these species and to build)
The insula has been implicated as a component of central networks subserving evaluative and affective processes. This study examined evaluative valence and arousal ratings in response to picture stimuli in patients with lesions of the insula and two contrast groups: a control-lesion group (the primary contrast group) and an amygdala-lesion group. Patients rated the positivity and negativity of picture stimuli (from very unpleasant to very pleasant) and how emotionally arousing they found the pictures to be. Compared with patients in the control-lesion group, patients with insular lesions reported reduced arousal in response to both unpleasant and pleasant stimuli, as well as marked attenuation of valence ratings. In contrast, the arousal ratings of patients with amygdala lesions were selectively attenuated for unpleasant stimuli, and these patients' positive and negative valence ratings did not differ from those of the control-lesion group. Results support the view that the insular cortex may play a broad role in integrating affective and cognitive processes, whereas the amygdala may have a more selective role in affective arousal, especially for negative stimuli.
Proceedings of the Workshop on Annotation and \nExploitation of Parallel Corpora AEPC 2010. \nEditors: Lars Ahrenberg, Jörg Tiedemann and Martin Volk. \nNEALT Proceedings Series, Vol. 10 (2010), 1-13. \n© 2010 The editors and contributors. \nPublished by \nNorthern European Association for Language \nTechnology (NEALT) \nhttp://omilia.uio.no/nealt. \nElectronically published at \nTartu University Library (Estonia) \nhttp://hdl.handle.net/10062/15893.
Several software programs exist to assist researchers in setting up online questionnaires. Existing tools are of little help for delivering online rating studies, for which it is often desirable to collect data from participants for only a subset of a stimulus set. OR-Vis enables researchers to quickly set up online rating studies by supplying the set of items to be rated, the number of stimuli an individual participant responds to, the number of participants an item is shown to, and the rating questions. The software then generates and delivers unique questionnaires for each participant, while managing the data collection process. The present article describes OR-Vis, its installation process, and how to use it to gather data. OR-Vis is open-source software and can be downloaded from www.orvis.uni-muenster.de.
In this paper, we present several ways to measure and evaluate the annotation and annotators, proposed and used during the building of the Czech part of the Prague Czech-English Dependency Treebank. At first, the basic principles of the treebank annotation project are introduced (division to three layers: morphological, analytical and tectogrammatical). The main part of the paper describes in detail one of the important phases of the annotation process: three ways of evaluation of the annotators- inter-annotator agreement, error rate and performance. The measuring of the inter-annotator agreement is complicated by the fact that the data contain added and deleted nodes, making the alignment between annotations non-trivial. The error rate is measured by a set of automatic checking procedures that guard the validity of some invariants in the data. The performance of the annotators is measured by a booking web application. All three measures are later compared and related to each other. 1.
Latent variable grammars take an observed (coarse) treebank and induce more fine-grained grammar categories, that are better suited for modeling the syntax of natural languages. Estimation can be done in a generative or a discriminative framework, and results in the best published parsing accuracies over a wide range of syntactically divergent languages and domains. In this paper we highlight the commonalities and the differences between the two learning paradigms. 1
In the architecture of a natural language processing system based on linguistic knowledge, two types of component are important: the knowledge databases and the processing modules. One of the knowledge databases is the lexical database, which is responsible for providing the lexical unities and its properties to the processing modules. The systems that process two or more languages require bilingual and/or multilingual lexical databases. These databases can be constructed by aligning distinct monolingual databases. In this paper, we present the interlingua and the strategy of aligning the two monolingual databases in REBECA, which only stores concepts from the “wheeled vehicle” domain.
We describe an effective constituent projection strategy, where constituent projection is performed on the basis of dependency projection. Especially, a novel measurement is proposed to evaluate the candidate projected constituents for a target language sentence, and a PCFG-style parsing procedure is then used to search for the most probable projected constituent tree. Experiments show that, the parser trained on the projected treebank can significantly boost a state-of-the-art supervised parser. When integrated into a tree-based machine translation system, the projected parser leads to translation performance comparable with using a supervised parser trained on thousands of annotated trees. 1
Discriminative parse reranking has been shown to be an effective technique to im-prove the generative parsing models. In this paper, we present a series of exper-iments on parsing the Tsinghua Chinese Treebank with hierarchically split-merge grammars and reranked with a perceptron-based discriminative model. In addition to the homogeneous annotation on TCT, we also incorporate the PCTB-based parsing result as heterogeneous annotation into the reranking feature model. The rerank-ing model achieved 1.12 % absolute im-provement on F1 over the Berkeley parser on a development set. The head labels in Task 2.1 are annotated with a sequence labeling model. The system achieved
Studies of discourse relations have not, in the past, attempted to characterize what serves as evidence for them, beyond lists of frozen expressions, or markers, drawn from a few well-defined syntactic classes. In this paper, we describe how the lexicalized discourse relation annotations of the Penn Discourse Treebank (PDTB) led to the discovery of a wide range of additional expressions, annotated as AltLex (alternative lexicalizations) in the PDTB 2.0. Further analysis of AltLex annotation suggests that the set of markers is open-ended, and drawn from a wider variety of syntactic types than currently assumed. As a first attempt towards automatically identifying discourse relation markers, we propose the use of syntactic paraphrase methods.
Internet dating is now ranked third as the way people meet, behind meeting at work or school, and through a friend or family member. This study researches the use of social and linguistic norms in online dating advertisements. Previous research has posed that social groups create unique identities and group members will selectively present themselves in ways consistent with these identities. Using Craigslist to assess the similarities and differences between genders and sexualities in online personal postings, an online quiz-like survey was created. This research reports on people's abilities to predict the sexual orientation and gender of the writer based on linguistic cues.
This paper proposes a method of correcting annotation errors in a treebank. By using a synchronous grammar, the method transforms parse trees containing annotation errors into the ones whose errors are corrected. The synchronous grammar is automatically induced from the treebank. We report an experimental result of applying our method to the Penn Treebank. The result demonstrates that our method corrects syntactic annotation errors with high precision.
Real world audio clips contain numerous acoustic sources. The rich acoustic information they carry cannot be fully described with single or even multiple terms about the acoustic sources alone. For instance, the label birds assigned to a birds singing clip that includes sounds of trees and a small river does not properly capture the experience it creates in the person listening to it. In this paper we introduce a novel scheme where the subjective experience of listening to sound clips containing mixture of sources is captured using affective measures. Furthermore, in contrast to the conventional approach of simple label-based methods, the affective ratings are then used to evaluate the performance of an example-based audio retrieval system. We argue that audio retrieval systems can benefit from using affective measures which are well established in experimental psychology, especially when dealing with real world audio clips. We present experimental results of our pilot study to support this motivation where the latent indexing framework has been employed for example-based retrieval on a collection of clips from the BBC sound effects library. The result of our study indicates that using the scheme of affective measures for representation and evaluation is indeed a promising direction to explore.
Discrete phonological phenomena form our conscious experience of language: continuous changes in pitch appear as distinct tones to the speakers of tone languages, whereas the speakers of quantity languages experience duration categorically. The categorical nature of our linguistic experience is directly reflected in the traditionally clear-cut linguistic classification of languages into tonal or non-tonal. However, some evidence suggests that duration and pitch are fundamentally interconnected and co-vary in signaling word meaning in non-tonal languages as well. We show that pitch information affects real-time language processing in a (non-tonal) quantity language. The results suggest that there is no unidirectional causal link from a genetically-based perceptual sensitivity towards pitch information to the appearance of a tone language. They further suggest that the contrastive categories tone and quantity may be based on simultaneously covarying properties of the speech signal and t)
DeSR is a statistical transition-based dependency parser which learns from annotated corpora which actions to perform for building parse trees while scanning a sentence. We describe the experiments performed for the ICON 2010 Tools Contest on Indian Dependency Parsing. DesR was configured to exploit specific features from the Indian treebanks. The submitted run used a stacked combination of four configurations of the DeSR parser and achieved the best unlabeled accuracy scores in all languages. The contribution to the result of various choices is analyzed.
Reviewed by: Laboratory phonology 8 Jaye Padgett Laboratory phonology 8. Ed. by Louis Goldstein, D. H. Whalen, and Catherine T. Best. (Phonology and phonetics 4-2.) Berlin: Mouton de Gruyter, 2006. Pp. xvi, 675. ISBN 9783110176780. $192 (Hb). This collection of papers is the end-product of the eighth Conference on Laboratory Phonology (LabPhon), held in New Haven, Connecticut, in June 2002, and hosted by Yale University and Haskins Laboratories. The volume is dedicated to the memory of Catherine P. Browman. If the LabPhon conferences and volumes were a bit renegade when they began, they are now more of an institution. It was still unusual in the 1980s to combine phonological theorizing with experimental methods and with theories drawn from phonetics and psycholinguistics, but to do so now seems more the norm. Out of thirteen papers published in the 1989 issue of Phonology, three incorporate experimental methodologies. (I construe ‘experimental’ broadly to include, for example, gestural or neural network modeling and formal learning theory.) In 2009 it was nine out of thirteen. (Six of these were from a special issue called ‘Phonological models and experimental data’; the reader can decide whether this strengthens or weakens the point.) The early ‘labphon’ movement can take credit for much of this change. This year we should see the inaugural publication of a laboratory phonology journal to replace the published volumes. In my view, this shift to a regular, peer-reviewed, and more accessible forum is very welcome. The book contains twenty-six contributions (including four commentary pieces) and an introduction. It is divided into three sections (two of them further subdivided): ‘Qualitative and variable faces of phonological competence’, ‘Sources of variation and their role in the acquisition of phonological competence’, and ‘Knowledge of language-specific organization of speech gestures’. I found these groupings to be nebulous; what comes through much more clearly is a second theme, on sign languages and comparisons between spoken and sign language. The deployment of laboratory methods and ‘philosophy’ in exploring sign languages is an exciting development. Other leitmotifs in the book draw on gestural phonology, exemplar modeling, acquisition, and the roles of abstract and categorical vs. concrete and gradient notions in representation and usage. Some of the papers are probably longer and less clearly written than they might be, but this is a minor complaint about a very interesting collection of works. Given space limitations here, I could not do justice to all twenty-six contributions; instead I focus on highlighting a few of them. Mirjam Ernestus and Harald Baayen, in ‘The functionality of incomplete neutralization in Dutch: The case of past-tense formation’ (27–49), replicate, for one speaker, the finding in Warner et al. 2004, 2006 of incomplete neutralization (IN) of final devoicing in Dutch, based on a reading task involving nonce verb forms. (Unlike in Warner et al., the forms were not presented as minimal pairs.) Particularly interesting are the results of their perception experiments using the speaker’s productions as stimuli. Ernestus and Baayen show that subjects not only detected IN but also used it to choose the appropriate past-tense ending (-te or -de) for the nonce verb stimulus forms, a task that requires the listener to infer the underlying voicing of the stem-final obstruent. The authors argue that IN, as well as their perception results, are due to the storage of lexical paradigms, among other things. Consider for example the form [vεrvεit] ‘widen’ and its infinitival form [vεrvεid n]. Even if the former is stored in its surface form (contrary to the assumption of [End Page 957] most generative phonologists), both the production and the perception of its final consonant will be influenced by activation of the associated form [vεrvεid n] (see also Bybee 2001). IN has posed a serious problem for the traditional understanding of the phonology-phonetics relation, in which discrete phonology is transduced into continuous phonetics, because if /vεrvεid/ is categorically devoiced to [vεrvεit] by phonology, then phonetic implementation has no means of recovering underlying voicing in order to produce IN. The storage and use of entire paradigms circumvents this problem. Since not all...
The paper introduces principles of rescript and lemmatization of lexemes with German origin for Lexical database of Baroque and Humanist Czech made in the Czech Language Institute of the Academy of Sciences of the Czech Republic in Prague.
The verb google is intriguing for the study of morphology, loanwords, assimilation, language contrast and neologisms. We present data for it for nineteen languages from nine language families.
In this paper, we introduce our recent work on re-annotating the deep information, which includes both the grammatical functional tags and the traces, in a Chinese scientific tree-bank. The issues with regard to re-annotation and its corresponding solutions are discussed. Furthermore, the process of the re-annotation work is described.
What kind of mental objects are letters? Research on letter perception has mainly focussed on the visual properties of letters, showing that orthographic representations are abstract and size/shape invariant. But given that letters are, by definition, mappings between symbols and sounds, what is the role of sound in orthographic representation? We present two experiments suggesting that letters are fundamentally sound-based representations. To examine the role of sound in orthographic representation, we took advantage of the multiple scripts of Japanese. We show two types of evidence that if a Japanese word is presented in a script it never appears in, this presentation immediately activates the ("actual") visual word form of that lexical item. First, equal amounts of masked repetition priming are observed for full repetition and when the prime appears in an atypical script. Second, visual word form frequency affects neuromagnetic measures already at 100-130 ms whether the word is pre)
Background: Absolute pitch (AP) is the ability to identify or produce isolated musical tones. It is evident primarily among individuals who started music lessons in early childhood. Because AP requires memory for specific pitches as well as learned associations with verbal labels (i.e., note names), it represents a unique opportunity to study interactions in memory between linguistic and nonlinguistic information. One untested hypothesis is that the pitch of voices may be difficult for AP possessors to identify. A musician's first instrument may also affect performance and extend the sensitive period for acquiring accurate AP. Methods/Principal Findings: A large sample of AP possessors was recruited on-line. Participants were required to identity test tones presented in four different timbres: piano, pure tone, natural (sung) voice, and synthesized voice. Note-naming accuracy was better for non-vocal (piano and pure tones) than for vocal (natural and synthesized voices) test tones. Th)
This paper reports on a treebanking project where eight different modern Chinese translations of the Bible are syntactically analyzed. The trees are created through dynamic treebanking which uses a parser to produce the trees. The trees have been going through manual checking, but correc-tions are made not by editing the tree files but by re-generating the trees with an updated grammar and dictionary. The accuracy of the treebank is high due to the fact that the grammar and dictionary are optimized for this specif-ic domain. The tree structures essen-tially follow the guidelines of the Penn Chinese Treebank. The total number of characters covered by the treebank is 7,872,420 characters. The data has been used in Bible translation and Bi-ble search. It should also prove useful in the computational study of the Chi-nese language in general. 1
We investigate parsing accuracy on the Korean Treebank 2.0 with a number of different grammars. Comparisons among these grammars and to their English counterparts suggest different aspects of Korean that contribute to parsing difficulty. Our results indicate that the coarseness of the Treebank’s nonterminal set is a even greater problem than in the English Treebank. We also find that Korean’s relatively free word order does not impact parsing results as much as one might expect, but in fact the prevalence of zero pronouns accounts for a large portion of the difference between Korean and English parsing scores. 1
A variety of query systems have been developed for interrogating parsed corpora, or treebanks. With the arrival of efficient, widecoverage parsers, it is feasible to create very large databases of trees. However, existing approaches that use in-memory search, or relational or XML database technologies, do not scale up. We describe a method for storage, indexing, and query of treebanks that uses an information retrieval engine. Several experiments with a large treebank demonstrate excellent scaling characteristics for a wide range of query types. This work facilitates the curation of much larger treebanks, and enables them to be used effectively in a variety of scientific and engineering tasks. 1