Papers reviewed and determined not to be word norm studies. Use the flag icon to report errors or suggest re-inclusion.
16504 papers
In the present article, we outline the architecture of a computer program for simulating the process by which humans comprehend texts. The program is based on psycholinguistic theories about human memory and text comprehension processes, such as the construction-integration model (Kintsch, 1998), the latent semantic analysis theory of knowledge representation (Landauer & Dumais, 1997), and the predication algorithms (Kintsch, 2001; Lemaire & Bianco, 2003), and it is intended to help psycholinguists investigate the way humans comprehend texts.
This paper explores techniques to take advantage of the fundamental difference in structure between hidden Markov models (HMM) and hierarchical hidden Markov models (HHMM). The HHMM structure allows repeated parts of the model to be merged together. A merged model takes advantage of the recurring patterns within the hierarchy, and the clusters that exist in some sequences of observations, in order to increase the extraction accuracy. This paper also presents a new technique for reconstructing grammar rules automatically. This work builds on the idea of combining a phrase extraction method with HHMM to expose patterns within English text. The reconstruction is then used to simplify the complex structure of an HHMM The models discussed here are evaluated by applying them to natural language tasks based on CoNLL-2004 1 and a sub-corpus of the Lancaster Treebank 2.
Abstract Attempting to automatically learn to identify verb complements from natural language corpora without the help of sophisticated linguistic resources like grammars, parsers or treebanks leads to a significant amount of noise in the data. In machine learning terms, where learning from examples is performed using class-labelled feature-value vectors, noise leads to an imbalanced set of vectors: assuming that the class label takes two values (in this work complement/non-complement), one class (complements) is heavily underrepresented in the data in comparison to the other. To overcome the drop in accuracy when predicting instances of the rare class due to this disproportion, we balance the learning data by applying one-sided sampling to the training corpus and thus by reducing the number of non-complement instances. This approach has been used in the past in several domains (image processing, medicine, etc) but not in natural language processing. For identifying the examples that are safe to remove, we use the value difference metric, which proves to be more suitable for nominal attributes like the ones this work deals with, unlike the Euclidean distance, which has been used traditionally in one-sided sampling. We experiment with different learning algorithms which have been widely used and their performance is well known to the machine learning community: Bayesian learners, instance-based learners and decision trees. Additionally we present and test a variation of Bayesian belief networks, the COr-BBN (Class-oriented Bayesian belief network). The performance improves up to 22% after balancing the dataset, reaching 73.7% f-measure for the complement class, having made use only a phrase chunker and basic morphological information for preprocessing.
DepAnn is an interactive annotation tool for dependency treebanks, providing both graphical and text-based annotation interfaces. The tool is aimed for semi-automatic creation of treebanks. It aids the manual inspection and correction of automatically created parses, making the annotation process faster and less error-prone. A novel feature of the tool is that it enables the user to view outputs from several parsers as the basis for creating the final tree to be saved to the treebank. DepAnn uses TIGER-XML, an XML-based general encoding format for both, representing the parser outputs and saving the annotated treebank. The tool includes an automatic consistency checker for sentence structures. In addition, the tool enables users to build structures manually, add comments on the annotations, modify the tagsets, and mark sentences for further revision.
Previous articleNext article FreeCurrent ApplicationsLinguistic AnthropologyV.ChandV.Chand Search for more articles by this author PDFPDF PLUSFull Text Add to favoritesDownload CitationTrack CitationsPermissionsReprints Share onFacebookTwitterLinked InRedditEmailQR Code SectionsMoreImmigration Practices in Belgium: African AsylumSeeker DiscourseJan Blommaert, a Belgian anthropologist currently at the Institute of Education, University of London, is known in Belgium as a public campaigner on immigration issues. He became involved in African asylumseekers rights in 1998, when the death of Semira Adamu during her forced repatriation provoked a public outcry over the implementation of Belgian immigration policies, in which more than 95% of applicants for asylum are rejected. Adamu had fled Nigeria in her teens to avoid entering into a polygamous marriage with a 65yearold man. Shackled at the ankles and vigorously resisting, she was suffocated by Belgian police attempting to restrain her.Poster for a 2003 commemoration of Adamu's death.View Large ImageDownload PowerPointAdamus death was a catalyst for the formation of new action groups, and existing organizations also became involved: churches opened their doors to asylum seekers, and NGOs such as OXFAM and the League of Human Rights campaigned for asylum seekers rights. Early on, Blommaert contributed to this organic collaborative effort: I gave tons of public lectures for any audience willing to listen, wrote opeds in major newspapers, campaigned with MPs close to the government, participated in public debates on these matters, and wrote expert articles for a wider audience. Additionally, he saw that his understanding of the legal and sociolinguistic issues was directly applicable to the problem. He was motivated, he explains, by awareness of the real stakes and real people involved.African asylum seekers face circumstances not shared with those from Europe and the Gulf because they come from wartorn areas with unclear state boundaries, may be illiterate and lack documentation, have long migration paths to Belgium, and do not share a language with immigration authorities. In particular, Blommaert points out, their choices of language, background texts, and genres for storytelling affect their chances of acquiring refugee status. African asylumseeker language is stereotypically filled with language mixing, language impurities, varying degrees of literacy, and different styles of storytelling and discourse. When asylum seekers present their stories, they are unaware of how the officials are judging them on their linguistic habits. The officials note these details and find them inconsistent with Belgian expectations; they judge the asylum seekers in terms of these nave choices and reject them. In many such cases, rejection is a matter of life or death for the refugees, given the risks associated with deportation and with repatriation into the home country that originally motivated them to seek asylum.Blommaert has mobilized an informal network of Belgian academics and institutions that has produced many academically informed public statements. Additionally, by documenting and analyzing oral immigration interviews with an eye to understanding the disconnect between asylum seekers presentations and immigration officials expectations, he has improved practice in the Immigration Department. He has used his analysis to train members of the department, helping to adjust views of what can be gathered in an immigration interview, to improve interview techniques, and to promote an awareness of variability in sociocultural linguistic norms and presentation styles.He continues to work as an official expert for Belgian legal, government, and security authorities, translating and analyzing documentation and corroboratory written texts provided for immigration procedures and advising on specific issues. He has been able to influence outcomes for some individual asylum seekers. His public campaigning has raised awareness of the issues surrounding African asylum seekers, and his publications have provoked work within academia that may lead to further interventions by academics with regard to the interview process.Blommaert hopes that public campaigning by NGOs and academic scholarship will promote continued dialogue with immigration officials, eventually producing policies and procedures that take into account the politics of migration and displacement and the way in which asylum seekers frame their life stories. Previous articleNext article DetailsFiguresReferencesCited by Current Anthropology Volume 47, Number 3June 2006 Sponsored by the Wenner-Gren Foundation for Anthropological Research Article DOIhttps://doi.org/10.1086/504161 Views: 287Total views on this site Citations: 1Citations are reported from Crossref PDF download Crossref reports the following articles citing this article:Kevin D. O'Gorman, Cailein Gillespie The mythological power of hospitality leaders?, International Journal of Contemporary Hospitality Management 22, no.55 (Jul 2010): 659–680.https://doi.org/10.1108/09596111011053792
AIM: (i) To determine the ability of general practitioners (GPs) and paediatricians to correctly identify children as overweight or obese by visual cues alone; (ii) to describe the current management practices of overweight and obese children by these practitioners; and (iii) to compare these with National Health and Medical Research Council (NHMRC) Clinical Practice Guidelines. METHODS: Forty-four GPs and 29 paediatricians participated in the study. Respondents completed a questionnaire based on a series of body images, rating the size of the child as acceptable weight, overweight or obese and indicating the likelihood of carrying out a series of management options. RESULTS: There was considerable variation in ability to rate images correctly with the total number of correct responses being 72% and 68%, respectively, for GPs and paediatricians. There were statistically significant differences in management between GPs and paediatricians in terms of conducting appropriate anthropometry and screening for co-morbidities, with paediatricians performing closer to the NHMRC Clinical Practice Guidelines. CONCLUSION: GPs and paediatricians have the opportunity to screen children for overweight and obesity during their everyday practice. Accurate determination of weight status cannot be performed by visualisation alone and all children should have height and weight measured and correctly interpreted. Some areas of current GP and paediatrician management of overweight and obese children fall short of the NHMRC clinical guidelines and areas for improvement are highlighted in this paper.
Annotated corpora constitute a crucial resource to acquire or induce linguistic knowledge about how languages are used. In this sense, it is widely admitted that tagged corpora appear to be a very useful resource for computational and linguistic analysis of languages. The more explicit linguistic information they contain, the more interesting and useful they are. In this paper we present the theoretical basis for semantic annotation of two treebanks, CESS-ESP and CESS-CAT, focusing specially on the verbal semantic classes that determine the mapping between syntactic functions and semantic roles.
This document describes the information used for summarization-inspired temporal-relation extraction [Dorr and Gaasterland, 2007]. We present a set of tense/aspect extraction templates that are applied to a Penn Treebank-style analysis of the input sentence. We also present an analysis of tense-pair combinations for different temporal connectives based on a corpus analysis of complex tense structures in Treebank-3. Finally, we include analysis charts and temporal relation tables for all combinations of intervals/points for each legal BTS combinations.
Research has shown that repeated statements are rated as more credible than new statements. However, little research has examined whether such "illusions of truth" can be produced by contextual (nonmnemonic) influences, or compared to the magnitude of these illusions in younger and older adults. In two experiments, we examined how manipulations of perceptual and conceptual fluency influenced truth and familiarity ratings made by young and older adults. Stimuli were claims about companies or products varying in normative familiarity. Results showed only small effects of perceptual fluency on rated truth or familiarity. In contrast, manipulating conceptual fluency via semantic/textual context had much larger effects on rated truth and familiarity, with the effects modulated by normative company familiarity such that fluency biases were larger for lesser-known companies. In both experiments, young and older adults were equally susceptible to fluency-based biases.
Many social phenomena involve a set of dyadic relations among agents whose actions may be dependent. Although individualistic approaches have frequently been applied to analyze social processes, these are not generally concerned with dyadic relations, nor do they deal with dependency. This article describes a mathematical procedure for analyzing dyadic interactions in a social system. The proposed method consists mainly of decomposing asymmetric data into their symmetric and skew-symmetric parts. A quantification of skew symmetry for a social system can be obtained by dividing the norm of the skew-symmetric matrix by the norm of the asymmetric matrix. This calculation makes available to researchers a quantity related to the amount of dyadic reciprocity. With regard to agents, the procedure enables researchers to identify those whose behavior is asymmetric with respect to all agents. It is also possible to derive symmetric measurements among agents and to use multivariate statistical techniques.
Ambulatory accelerometry is a technique that allows objective measurement of aspects of everyday human behavior. The aim of our research has been to develop, validate, and apply this technique, which recently resulted in an upper limb activity monitor (ULAM). The ULAM consists of body-mounted acceleration sensors connected to a waist-worn data recorder and allows valid and objective assessment of activity of both upper limbsduring performance of also automatically detected mobility-related activities: lying, sitting, standing, walking, cycling, and general movement. The ULAM can be used to determine (limitations of) upper limb activity and mobility in freely moving subjects with upper limb disorders. This article provides a detailed description of its characteristics, summarizes the results of a feasibility study and four application studies in subjects having upper limb complex regional pain syndrome, discusses the most important practical, technical, and methodological issues that were encountered, and describes current and future research projects related to measuring (limitations of) upper limb activity.
This paper evaluates four of the most commonly used, freely available, state-of-the-art parsers on a standard benchmark as well as with respect to a set of data relevant for measuring text cohesion, as one example of a learning technology application that requires fast and accurate syntactic parsing. We outline advantages and disadvantages of existing technologies and make recommendations. Our performance report uses traditional measures based on a gold standard as well as novel dimensions for parsing evaluation. To our knowledge, this is the first attempt to evaluate parsers across genres and grade levels for the implementation in learning technology using both gold standard and directed evaluation methods.
We exploit the resources in the Arabic Treebank (ATB) and Arabic Gigaword (AG) to determine the best features for the novel task of automatically creating lexical semantic verb classes for Modern Standard Arabic (MSA). The verbs are classified into groups that share semantic elements of meaning as they exhibit similar syntactic behavior. The results of the clustering experiments are compared with a gold standard set of classes, which is approximated by using the noisy English translations provided in the ATB to create Levin-like classes for MSA. The quality of the clusters is found to be sensitive to the inclusion of syntactic frames, LSA vectors, morphological pattern, and subject animacy. The best set of parameters yields an Fβ=1 score of 0.456, compared to a random baseline of an Fβ=1 score of 0.205.
Since ancient linguistics, the studies of Indo-European word order work with the conception of universal natural word order (ordo naturalis) – an order of the verb-dependent constituents in the linear organization of a clause. The description of the natural word order is usually based on occasional (and in some degree random) observations of clauses in a certain language. In Czech linguistics, the idea of the natural word order was formulated in a more precise way as the hypothesis of the systemic ordering (Sgall, Hajicova and Buraňova, 1980). According to the authors, the contextually non-bound participants and adverbials are ordered as follows (o. c., page 77):
We investigate generalizations of the all-subtrees "DOP" approach to unsupervised parsing. Unsupervised DOP models assign all possible binary trees to a set of sentences and next use (a large random subset of) all subtrees from these binary trees to compute the most probable parse trees. We will test both a relative frequency estimator for unsupervised DOP and a maximum likelihood estimator which is known to be statistically consistent. We report state-of-the-art results on English (WSJ), German (NEGRA) and Chinese (CTB) data. To the best of our knowledge this is the first paper which tests a maximum likelihood estimator for DOP on the Wall Street Journal, leading to the surprising result that an unsupervised parsing model beats a widely used supervised model (a treebank PCFG).
We present the implementation of a system which extracts not only lexicalized grammars but also feature-based lexicalized grammars from Korean Sejong Treebank. We report on some practical experiments where we extract TAG grammars and tree schemata. Above all, full-scale syntactic tags and well-formed morphological analysis in Sejong Treebank allow us to extract syntactic features. In addition, we modify Treebank for extracting lexicalized grammars and convert lexicalized grammars into tree schemata to resolve limited lexical coverage problem of extracted lexicalized grammars.
Grammars can be induced from treebanks for a potential variety (e.g., parsing), but this task faces two major problems. One is the existence of annotation er-rors in the treebank which can affect the final system performance (cf. [4]). The other is that the sheer number of rules is overwhelming for most applications. As
A grammatical method of combining two kinds of speech repair cues is presented. One cue, prosodic disjuncture, is detected by a decision tree-based ensemble classifier that uses acoustic cues to identify where normal prosody seems to be interrupted (Lickley, 1996). The other cue, syntactic parallelism, codifies the expectation that repairs continue a syntactic category that was left unfinished in the reparandum (Levelt, 1983). The two cues are combined in a Treebank PCFG whose states are split using a few simple tree transformations. Parsing performance on the Switchboard and Fisher corpora suggests that these two cues help to locate speech repairs in a synergistic way.
Solar-driven interfacial evaporation technology has attracted significant attention for water purification. However, design and fabrication of solar-driven evaporator with cost-effective, excellent capability and large-scale production remains challenging. In this study, inspired by plant transpiration, a tri-layered hierarchical nanofibrous photothermal membrane (HNPM) with a unidirectional water transport effect was designed and prepared via electrospinning for efficient solar-driven interfacial evaporation. The synergistic effect of the hierarchical hydrophilic-hydrophobic structure and the self-pumping effect endowed the HNPM with unidirectional water transport properties. The HNPM could unidirectionally drive water from the hydrophobic layer to the hydrophilic layer within 2.5 s and prevent reverse water penetration. With this unique property, the HNPM was coupled with a water supply component and thermal insulator to assemble a self-floating evaporator for water desalination. Under 1 sun illumination, the water evaporation rates of the designed evaporator with HNPM in pure water and dyed wastewater reached 1.44 and 1.78 kg·m<sup>-2</sup>·h<sup>-1</sup>, respectively. The evaporator could achieve evaporation of 11.04 kg·m<sup>-2</sup> in 10 h under outdoor solar conditions. Moreover, the tri-layered HNPM exhibited outstanding flexibility and recyclability. Our bionic hydrophobic-to-hydrophilic structure endowed the solar-driven evaporator with capillary wicking and transpiration effects, which provides a rational design and optimization for efficient solar-driven applications.
Powerful lookup capabilities are included in almost every electronic dictionary. However, improvements are still conceivable in the exploration of the lexical information within the dictionary and in the integration of the dictionary into other applications. In this paper, we illustrate the most salient functions of the Base lexicale dufrançais (BLF), an online lexical database for general French. The major features of this interactive database are the integration of a dictionary (DAFLES), automatically generated exercises (ALFALEX) and corpus applications (CATS) in order to create a powerful learning environment for learners of French. The DAFLES focuses on the learner's needs by providing separate access paths according to receptive or productive needs. ALFALEX trains the language learner on global lexical knowledge as it is described in the DAFLES. The corpus applications enrich the lexical description by providing the (comparative) combinatorial profiles of words as well as examples and sentences for the exercises.
In the past, a divide could be seen between ’deep ’ parsers on the one hand, which construct a semantic representation out of their input, but usually have significant coverage problems, and more robust parsers on the other hand, which are usually based on a (statistical) model derived from a treebank and have larger coverage,
While much of the research and labor in treebanks has focused on modern languages, recent scholarship has also seen the rise of treebanks for historical languages as well, such as Middle English (Kroch and Taylor [15]), Early Modern English (Kroch et al. [16]), Old English (Taylor et al. [28]), Early New High German (Demske et al. [11]) and Medieval Portuguese (Rocio et al. [27]). Like their modern counterparts, these historical treebanks serve two distinct ends and often two different audiences: they provide crucial datasets for NLP projects such as automatic parsing and grammar induction while also providing a valuable corpus for scholars researching the state of a language and its progression across time. Historical treebanks, however, also offer one additional benefit over modern treebanks: they provide an annotated set of texts that scholars actually care about. When linguists of modern languages base theories on corpus evidence, their analysis is generally directed toward the language at large; few, if any, pore over the Wall Street Journal examining its use of an arcane literary device. If the corpus is Vergil, however, we do. The sheer volume of Latin texts available electronically1 not to mention the enormous mass still locked in print is much larger than the small community of scholars and students who can read it. This alone justifies a treebank as a resource for those attempting to learn the language, but it also highlights the need for automatic methods of parsing and machine translation. To this end a Latin treebank will well serve the NLP community, which has a long history of applying such research to modern languages.2 Classical scholars, however, largely operate on a fixed canon of texts. The value of a treebank for them is not so much in training
Each year the Conference on Computational Natural Language Learning (CoNLL) features a shared task, in which participants train and test their systems on exactly the same data sets, in order to better compare systems. The tenth CoNLL (CoNLL-X) saw a shared task on Multilingual Dependency Parsing. In this paper, we describe how treebanks for 13 languages were converted into the same dependency format and how parsing performance was measured. We also give an overview of the parsing approaches that participants took and the results that they achieved. Finally, we try to draw general conclusions about multi-lingual parsing: What makes a particular language, treebank or annotation scheme easier or harder to parse and which phenomena are challenging for any dependency parser?
This study examined whether the presence of a label and the length of viewing time affect the perception of art. Participants were 152 undergraduate students at an urban university in central New Jersey. A computer program randomly assigned participants to viewing 4 paintings under a label or no label condition and a 1-s, 5-s, 30-s, or 60-s time condition. The paintings were Cezanne's Still Life With Apples, Mondrian's Composition with Red, Yellow and Blue, Monet's Garden at Sainte-Adresse, and Davis' Report from Rockport. After viewing each painting, participants completed a rating scale of 24 adjective pairs. Multivariate analyses of variance indicated little support for the hypotheses that the label condition will lead to different ratings of the art as compared to the non-label condition and that length of viewing time will affect ratings of the art. The reliability coefficients for the rating scale were fairly strong. Results are discussed in terms of how ratings of art are dependent largely on the work of art itself.
Objective:This study aimed to examine the effects of haloperidol and amphetamine on human startle response modulated by emotionally-toned film clips. Method: Sixty participants, in two groups (one receiving haloperidol and the other receiving amphetamine) were tested using electromyography (EMG) to measure eye-blink muscle (orbicular oculi) while different emotions were induced by six 2-minute film clips. Results: An affective rating shows the negative and positive effects of the two drugs on emotional reactivity, neither amphetamine nor haloperidol had any impact on the modulation of the startle response. Conclusion: The methodological and theoretical aspects of the study and findings will be discussed.
Line drawings are commonly used in perception research. A basic strategy used in such research is to remove portions of the line drawings in order to determine what features of an object are important for recognition. However, it is important to monitor the amount of contour and type of information that are deleted when one is making partially deleted or fragmented objects. With the Image Fragmenting Program, researchers can use random or manual contour deletion strategies to create fragmented objects while controlling for the amount of contour removed from the images.
The movements of newborns have been thoroughly studied in terms of reflexes, muscle synergies, leg coordination, and target-directed arm/hand movements. Since these approaches have concentrated mainly on separate accomplishments, there has remained a clear need for more integrated investigations. Here, we report an inquiry in which we explicitly concentrated on taking such a perspective and, additionally, were guided by the methodological concept of home base behavior, which Ilan Golani developed for studies of exploratory behavior in animals. Methods from nonlinear dynamics, such as symbolic dynamics and recurrence plot analyses of kinematic data received from audiovisual newborn recordings, yielded new insights into the spatial and temporal organization of limb movements. In the framework of home base behavior, our approach uncovered a novel reference system of spontaneous newborn movements.
An HMM-based single character recovery (SCR) model is proposed in this paper to extract a large set of atomic abbreviations and their full forms from a text corpus. By an “atomic abbreviation,” it refers to an abbreviated word consisting of a single Chinese character. This task is important since Chinese abbreviations cannot be enumerated exhaustively but the abbreviation process for compound words seems to be compositional. One can often decode an abbreviated word character by character to its full form. With a large atomic abbreviation dictionary, one may be able to handle multiple character abbreviation problems more easily based on the compositional property of abbreviations.
Use of structural information and lexicalization are two of the main challenges facing syntactic analysis, and they are investigated in this paper. First, the probabilities of lexical dependencies are obtained by training a large-scale dependency treebank and used to build the lexical model. Second, the governing degree of words is introduced to utilize the structure information. The lexical method overcomes the weakness of POS dependencies in the past work; meanwhile the governing degree of words is helpful to distinguish the syntactic structures so some ill-formed structures are avoided. Finally, the paper shows a good experimental result of around 74% accuracy on the test set that consists of 4000 sentences.
cjelokupni tekst: Abstract: From the translator's point of view, collocations and idioms belong to rather demanding text units, which often require a high level of linguistic, communicative, cultural and translational competence. The translator needs to be aware, and appreciative of their semantics (the difficulty with idioms being that their meaning is not deducible from that of the individual elements), their syntax (non-free and sometimes puzzling), their pragmatics (related to a variety of linguistic and textual circumstances, more specific of which can stem from conscious breach of the frozen syntax, the metaphorical quality of their meaning or culture-specific usage patterns) on both the source and the target end of the translation process. In addition, a successful choice of an appropriate equivalent requires a well-founded translational decision as to the most relevant aspect(s) of the value of the phrase in question i.e. the aspect which should be matched in the translation, as well as an awareness of the various procedures that can be employed. The paper discusses translation equivalents of Swedish collocations and idioms manifest in Croatian translations of about 1000 book pages of Swedish fiction. In an attempt to check whether there are any patterns in the treatment of collocations and idioms in translation, the discussion focuses on the following characteristics of the established equivalents: - whether they are the same type of linguistic units as the original phrase (collocation or idiom); - which aspect of the original phrase’ s linguistic and communicative value has been given prominence in the translation (i.e. is best matched by the chosen translation equivalent); - whether they can be considered lexical or grammatical calques (if yes, of what kind). The basis for the discussion is provided by twofold procedure: - a systematic analysis of all the translation equivalents of ten Swedish lexical collocations / idioms established in the translation of seven books of fiction; - an analysis of the equivalents of some other lexical collocations and idioms established in randomly selected extracts from three other novels. The analysis is primarily expected to offer an insight into the ways in which lexical collocations and idioms are commonly perceived and treated. In addition it will feed into a more general picture of translation practices and translation norms in Croatia.
On the norm of current Chinese characters,象 像 of non-noun mainly the verb lexical meaning have not got the ideal social effect yet in the division and usage.And the present state is ambiguous.By reviewing the differences in non-noun lexical meaning and evidence of their division,it is thought that only divided their usage clearly can they be used normally.
This empirical study evaluates the factors that influence Korean students’ reading comprehension of target culture-embedded texts in the US. These texts contain US socio-cultural facts and words that represent native speakers’ beliefs, norms, and values in their daily lives. They also introduce diverse word meanings depending on each context. Through this study, I found that the most important factors influencing the Korean students’ comprehension of a target culture-embedded text are their lack of US cultural knowledge and US culture-embedded lexical knowledge. These factors were linked to the students’ poor meaning-making strategy in comprehending US culture-embedded texts appropriately. Contextual and lexical knowledge is a necessary factor for Korean students to make appropriate meaning and better comprehend US culture-embedded texts through diverse exposures in EFL reading classes.
This paper presents a Chinese dependency syntax for treebanking. The syntax contains 13 word classes and 34 dependency types. A format of treebank based on the syntax is also proposed for the applications of computational and general linguistic research. Some experiments show that the treebank based on the proposed dependency syntax can be used for training and evaluating the dependency parser and for quantitative analysis of Chinese syntax.
Reviewed by: Using corpora to explore linguistic variation ed. by Randi Reppen, Susan M. Fitzmaurice, and Douglas Biber Merja Kytö Using corpora to explore linguistic variation. Ed. by Randi Reppen, Susan M. Fitzmaurice, and Douglas Biber. (Studies in corpus linguistics 9.) Amsterdam: John Benjamins, 2002. Pp. xii, 274. ISBN 1588112837. $126 (Hb). Among the increasing number of books dedicated to the study of linguistic variation and aspects of language use, this volume stands out. It offers a well-balanced selection of corpus-based studies that cover a broad range of linguistic features, registers, and dialects (or varieties) of English, but also includes one study in which the target language is Russian. In the editors’ words, ‘adequate descriptions of variation and use must be based on empirical analyses of natural texts’ and ‘on multiple texts collected from many speakers’. Moreover, these descriptions ‘must simultaneously consider the influence of a range of contextual factors on linguistic variability’ (vii). The studies included in the volume are carried out with these methodological aims in mind. [End Page 438] The approach has been gaining momentum over the past few decades, along with the recent advances made in the compilation and exploitation of computerized corpora. The book comprises an illuminating introduction and thirteen chapters divided into three sections. In Part 1 (Chs. 1–9), the main research goal is to explore variation in the use of different linguistic features across various dialects and registers. Two of the nine studies are devoted to hedging, one of which focuses on uses of the modal would. There are two further studies on modals, and five studies devoted to various phraseological and lexical phenomena, ranging from formulaic language to lexical bundles, pseudo-titles, and issues in pattern grammar. The two chapters in Part 2 are devoted to dialect or register variation. Ch. 10 describes syntactic features in written Indian English, and Ch. 11 investigates linguistic variation in academic lectures. The two chapters in Part 3 focus on historical variation. Ch. 12 deals with patterns of negation in eighteenth-century English, and Ch. 13 looks at register variation in nineteenth-century English. Deanna Poos and Rita Simpson investigate the functions of hedging in academic English using a pilot version of the Michigan Corpus of Spoken English (MICASE). Focusing on kind of and sort of, they show convincingly that the hedging frequencies do not depend so much on the gender variable as on academic discipline, type of interaction, and speaker needs within a given academic institution. Hedging is shown to be more characteristic of discourse in the humanities than in the physical sciences; of particular interest in this context are the possible reasons the authors suggest might account for this empirical finding. The other important result presented concerns the multifunctionality of kind of and sort of in spoken interaction. In addition to appearing as indicators of tentativeness, these items also ‘serve a variety of often overlapping sociopragmatic purposes in spoken interaction’ (21). Thus the study brings together the gender and multifunctionality issues, emphasizing the importance of acknowleding the multiplex nature of speaker identity. Fiona Farr and Anne O’keeffe compare the use of would as a hedging device in spoken Irish English (sampled from phone-in conversations on radio and post-observation teacher training interaction) with the uses attested in corpora representative of spoken British and American English. The corpus evidence shows that would is used more in Irish English than in the other varieties, and more specifically, for broader pragmatic functions than merely to mitigate or tone down the force of an utterance. To account for their findings, the authors arrive at a hierarchical three-tiered model where register, setting, and sociocultural norms interact to influence language choice. The model clearly has potential for further research. By contrast, Graeme Kennedy limits his study of modals to one variety, British English. However, as the object of this study is the 100-million-word British National Corpus (BNC), we are presented with a gigantic dataset (1,457,721 tokens), extracted on the basis of the tagging accompanying the text. Kennedy’s study is a fascinating piece of work, one of the few in which...
Several recent Information Extraction (IE) systems have been restricted to the identification facts which are described within a single sentence. It is not clear what effect this has on the difficulty of the extraction task or how the performance of systems which consider only single sentences should be compared with those which consider multiple sentences. This paper compares three IE evaluation corpora, from the Message Understanding Conferences, and finds that a significant proportion of the facts mentioned therein are not described within a single sentence. Therefore systems which are evaluated only on facts described within single sentences are being tested against a limited portion of the relevant information in the text and it is difficult to compare their performance with other systems. Further analysis demonstrates that anaphora resolution and world knowledge are required to combine information described across multiple sentences. This result has implications for the development and evaluation of IE systems.
Wordnets, which are repositories of lexical semantic knowledge containing semantically linked synsets and lexically linked words, are indispensable for work on computational linguistics and natural language processing. While building wordnets for Hindi and Marathi, two major Indo-European languages, we observed that the verb hierarchy in the Princeton Wordnet was rather shallow. We set to constructing a verb knowledge base for Hindi, which arranges the Hindi verbs in a hierarchy of is-a (hypernymy) relation. We realized that there are unique Indian language phenomena that bear upon the lexicalization vs. syntactically derived choice. One such example is the occurrence of conjunct and compound verbs (called Complex Predicates) which are found in all Indian languages. This paper presents our experience in the construction of lexical knowledge bases for Indian languages with special attention to Hindi. The question of storing versus deriving complex predicates has been dealt with linguistically and computationally. We have constructed empirical tests to decide if a combination of two words, the second of which is a verb, is a complex predicate or not. Such tests provide a principled way of deciding the status of complex predicates in Indian language wordnets.
Detecting idioms in a sentence is important to sentence understanding. This paper discusses the linguistic knowledge for idiom detection. The challenges are that idioms can be ambiguous between literal and idiomatic meanings, and that they can be “transformed” when expressed in a sentence. However, there has been little research on Japanese idiom detection with its ambiguity and transformations taken into account. We propose a set of linguistic knowledge for idiom detection that is implemented in an idiom dictionary. We evaluated the linguistic knowledge by measuring the performance of an idiom detector that exploits the dictionary. As a result, more than 90% of the idioms are detected with 90% accuracy.