Papers reviewed and determined not to be word norm studies. Use the flag icon to report errors or suggest re-inclusion.
18265 papers
In data-driven approaches to natural language processing, a common problem is the lack of data for many languages. Within the project Stochastic Dependency Grammars for Natural Language Parsing at Växjö University, we (Joakim Nivre, Johan Hall and Jens Nilsson) are developing a deterministic data-driven dependency parser, which is language independent. In this project we intend to enlarge the data resources for our parser. For the moment, we have only tested our parser on small Swedish treebank converted to dependency structure, and on English using Penn Treebank converted to dependency trees. Since we do not have more Swedish dependency treebanks at hand, we want to broaden our view towards treebanks for other languages, especially the bigger ones, to investigate the influence of data size. Primarily, we are focusing on the Danish Dependency Treebank (DDT) and the Prague Dependency Treebank (PDT). These treebanks are not in a format that we can use for our parser and therefore we have to convert them to MALT-XML, a format which our parser can handle.
Discourse connectives can be analyzed as discourse level predicates which projectpredicate-argument structure on a par with verbs at the sentence level. The PennDiscourse Treebank (PDTB) reflects this view in its design providing annotation ofthe discourse connectives and their arguments. Like verbs, discourse connectiveshave multiple senses. We present a set of manual sense annotation studies for threeconnectives whose arguments have been annotated in the PDTB. Using syntacticfeatures computed from the Penn Treebank and a simple MaxEnt model, we haveachieved some success in automatically disambiguating among their senses.
Development and testing of a large range of NLP applications presuppose corpora annotated at levels more advanced than those of part–of–speech and shallow syntax. Therefore, multi–layered annotation schemes have been designed in order to provide deeper representations of intra – and inter–sentential structure and meaning.
In the first part of this technical report we describe our approach to design a new data format, based on XML (Extensible Markup Language) and aimed to provide a better and unifying alternative to various legacy data formats used in various areas of corpus linguistics and specifically in the field of structured annotation. We introduce the first version of the format, called Prague Markup Language (PML). This version has already been employed as the main data format for the upcoming Prague Dependency Treebank 2.0 (PDT). Finally we outline our ideas and proposals for further improvement of PML, based on our current experience with using and processing data in PML format in the PDT 2.0 project. The second part of the technical report contains the state-of-the-art specification of PML. Technicka zprava c. TR-2005-29 Technicka zprava projektu Integrace jazykových zdrojů za ucelem extrakce informaci z přirozených textů Projekt Informacni spolecnosti Grantove agentury Akademie věd CR Registracni cislo GA AV CR: 1ET101120503 Interni kod MFF: 207-14 / 242083
In this paper we present a number of experiments to test the portability of existing treebankinduced LFG resources. We test the LFG parsing resources of Cahill et al. (2004) on the ATIS corpus which represents a considerabley different domain to the Penn-II Treebank Wall Street Journal sections, from which the resources were induced. This testing shows an underperformance at both c- and f-structure level as a result of the domain variation. We show that in order to adapt the LFG resources of Cahill et al. (2004) to this new domain, all that is necessary is to retrain the c-structure parser on data from the new domain.
We describe a parallel annotation approach for PubMed abstracts. It includes both entity/relation annotation and a treebank containing syntactic structure, with a goal of mapping entities to constituents in the treebank. Crucial to this approach is a modification of the Penn Treebank guidelines and the characterization of entities as relation components, which allows the integration of the entity annotation with the syntactic structure while retaining the capacity to annotate and extract more complex events.
Traditional Constraint Grammar (CG) is a methodological, rather than a descriptive paradigm, designed for robust parsing, not the implementation of a specific linguistic theory. Therefore, if used for treebank generation, it is not immediately clear, which linguistic formalism would be easiest to
Great progress has been made in parsing the Wall Street Journal portion of the Penn Treebank. Now parsing languages other than English is an intensive research area. Head-driven model is one of the best English parsing models. It has been successfully applied to Czech but failed to outperform a base-line model in parsing German. This paper attempts to parse Chinese with head-driven model. Promising experimental results demonstrate that head-driven model works well for Chinese. We propose a hybrid parsing strategy, which combines head-driven model with a Chinese base phrases parsing model. The combined model not only improves the performance but also makes the parser space and time efficient. We evaluate our method in PARSEVAL measures, and the combined model performances are at 79.88% precision, 81.97% recall.
This research shows a cost-effective approach to simultaneously produce proteases, amylases, and endoglucanases from <i>Stachybotrys microspora</i> that could be considered a compatible detergent additive in the green detergent industry.
Abstract This corpus-based contrastive study examines the thematic use of the semantic field of research and researchers in the Discussion section of biomedical reports in Spanish native texts and English-Spanish translations. This semantic field was divided into integral reference (specific named researchers), general nouns for researchers, and singular and plural nouns referring to research. Themes containing these lexical items were examined with regard to their syntactic manifestations and their lexicogrammatical relations with the main finite verb. Quantitative analysis was used to establish reference values for the native texts and to reveal differences between the two subcorpora. Qualitative contextual analysis then investigated how the data might be applied to the translated texts. The quantitative study showed that the Spanish texts had more integral references and more general researcher nouns in their themes whereas the translations had more singular research nouns, especially those referring to the current study. Singular research nouns were associated with more prepositional adjuncts in the Spanish texts but with more subject themes, either as head or as modifier, in the translations. The distribution of tenses was different in all categories except for plural research nouns, with a higher percentage of present and present perfect in the Spanish texts and more past indefinite in the translations. Differences were also found in the distribution of lexical verbs related to integral references and singular research nouns. The contextual analysis revealed that awareness of these differences and strategic choices based on them could lead to thematic and discourse patterns that come closer to the target-language norms for this genre.
It is not clear a priori how well parsers trained on the Penn Treebank will parse significantly different corpora without retraining. We carried out a competitive evaluation of three leading treebank parsers on an annotated corpus from the human molecular biology domain, and on an extract from the Penn Treebank for comparison, performing a detailed analysis of the kinds of errors each parser made, along with a quantitative comparison of syntax usage between the two corpora. Our results suggest that these tools are becoming somewhat over-specialised on their training domain at the expense of portability, but also indicate that some of the errors encountered are of doubtful importance for information extraction tasks.
Computer-based studies usually produce log files as raw data. These data cannot be analyzed adequately with conventional statistical software. The Chemnitz LogAnalyzer provides tools for quick and comfortable visualization and analyses of hypertext navigation behavior by individual users and for aggregated data. In addition, it supports analogous analyses of questionnaire data and reanalysis with respect to several predefined orders of nodes of the same hypertext. As an illustration of how to use the Chemnitz LogAnalyzer, we give an account of one study on learning with hypertext. Participants either searched for specific details or read a hypertext document to familiarize themselves with its content. The tool helped identify navigation strategies affected by these two processing goals and provided comparisons, for example, of processing times and visited sites. Altogether, the Chemnitz LogAnalyzer fills the gap between log files as raw data of Web-based studies and conventional statistical software.
We present a methodology for extracting subcategorization frames based on an automatic lexical-functional grammar (LFG) f-structure annotation algorithm for the Penn-II and Penn-III Treebanks. We extract syntactic-function-based subcategorization frames (LFG semantic forms) and traditional CFG category-based subcategorization frames as well as mixed function/category-based frames, with or without preposition information for obliques and particle information for particle verbs. Our approach associates probabilities with frames conditional on the lemma, distinguishes between active and passive frames, and fully reflects the effects of long-distance dependencies in the source data structures. In contrast to many other approaches, ours does not predefine the subcategorization frame types extracted, learning them instead from the source data. Including particles and prepositions, we extract 21,005 lemma frame types for 4,362 verb lemmas, with a total of 577 frame types and an average of 4.8 frame types per verb. We present a large-scale evaluation of the complete set of forms extracted against the full COMLEX resource. To our knowledge, this is the largest and most complete evaluation of subcategorization frames acquired automatically for English.
This dissertation is a study on the works of the Swedish author Carl Jonas Love Almqvist during the final years of his exile in America. Focusing on the monumental 1438-page unpublished manuscript 'About Swedish Rhymes', the study first presents the textual material and then discusses the text from different formal and content-based aspects essential to an understanding of Almqvist's works in exile. In the manuscripts preserved from his last years of exile, i.e. the period after 1860, Almqvist refers to 'Mr Hugo's Academy, established in the year 1838' introduced in one of the volumes of The Book of the Wild Rose (1839). In comparison with his earlier fiction about academic "cabinet meetings", this fiction of such an academy, conceived in exile, is in some ways extraordinary. A close reading of the texts reveals that the aging Almqvist, contrary to previous opinions about him, maintained strict control over the activity: the extension and division of the record, as well as its references to time and space, all indicate a complete consistency and an exact mimetic order. The consideration of 'About Swedish Rhymes' starts out from exterior qualities. The observations are first considered in relation to the author's statements on the importance of the manuscript for the literary work of art. Subsequently, the genesis of the "exile" texts is re-examined. One key question here is whether the manuscript was completed in Philadelphia, or was continued in Bremen during the final year of his life. The content of the conversations in the records of the cabinet meetings is also analyzed. Although questions of metre and versification dominate, the text also deals with a variety of widely differing subjects, including discussions about the use of language and linguistic norms. The fictitious frame that the cabinet meeting provides for the purpose of discussing metre and rhyme is also considered. Here we find various improvised verses composed at the cabinet meeting and put into the mouth of the authentic versifier H.J. Seseman. One important question is whether the cabinet-meeting discussions about the metre in these verses are intended to be a serious contribution to scholarly debate, or whether they in fact have ironic undertones. Next, the narration of the "exile" texts is discussed from the point of view provided by its own fictitious perspective, together with the author’s relation to irony, satire and parody. The concluding chapter deals with verse-making in the record of rhyming. The emphasis is laid on the analysis and characterization of the various rhymed verses collected under the title Sesemana. One essential question concerns the 'rubbishy' or 'plain' character of these poems. The present analysis indicates that questions of rubbish, textual triviality and the like must bow to the broader question of the character of the poems in a deeper sense. Seseman's poetry is considered in relation to the Songes collection. Finally the question of how rhythm manifests itself as 'free verse' in a number of these poems with more serious content is also discussed.
We present a heuristic technique for converting a constituency treebank into a dependency treebank. In particular, we comment on our experience in converting the Spanish treebank Cast3LB. We extract a context-free grammar from the treebank, automatically identify the head in each rule, and use this information for constructing the dependency tree. Our heuristics have 99 % precision and 80 % recall in identifying the head in the rules, which gives 92% accuracy in identifying dependencies between words.
Linguistic politeness is intimately connected with social norms. Estoniansociety has gone through considerable change over the last ten years. It hasregained independence and, at the same time, switched from a planned toa market economy as well as from dictatorship to democracy. A decade ismost probably not long enough for linguistic norms to change drastically:as we know, the structure of a language often takes much longer to change.Politeness, however, may to some extent be subject to deliberate influence,as witnessed, for example, by the reform of Swedish du (you, sg.) where therecommendations of some left-wing organisations on the usage of mutualdu (T) have won general social acceptance. It is, thus, not unlikely that change is taking place in Estonian politeness at present....
The paper investigates the use of richer syntactic dependencies in the structured language model (SLM). We present two simple methods of enriching the dependencies in the syntactic parse trees used for initializing the SLM. We evaluate the impact of both methods on the perplexity (PPL) and word-error-rate (WER, N-best rescoring) performance of the SLM. We show that the new model achieves an improvement in PPL and WER over the baseline results reported using the SLM on the UPenn Treebank and Wall Street Journal (WSJ) corpora, respectively.
We present a strictly lexical parsing model where all the parameters are based on the words. This model does not rely on part-of-speech tags or grammatical categories. It maximizes the conditional probability of the parse tree given the sentence. This is in contrast with most previous models that compute the joint probability of the parse tree and the sentence. Although the maximization of joint and conditional probabilities are theoretically equivalent, the conditional model allows us to use distributional word similarity to generalize the observed frequency counts in the training corpus. Our experiments with the Chinese Treebank show that the accuracy of the conditional model is 13.6% higher than the joint model and that the strictly lexicalized conditional model outperforms the corresponding unlexicalized model based on part-of-speech tags.
This paper reports the corpus-oriented development of a wide-coverage Japanese HPSG parser. We first created an HPSG treebank from the EDR corpus by using heuristic conversion rules, and then extracted lexical entries from the treebank. The grammar developed using this method attained wide coverage that could hardly be obtained by conventional manual development. We also trained a statistical parser for the grammar on the treebank, and evaluated the parser in terms of the accuracy of semantic-role identification and dependency analysis.
Cette thèse aborde le problème de la structuration de bases lexicales multilingues (BDLM) en lexies et axies, à partir de ressources existantes. Ce travail est motivé par l'inadéquation des techniques existantes utilisées isolément, pour la structuration de BDLM. Pour résoudre ce problème, la stratégie proposée est de composer des techniques existantes de désambiguïsation pour structurer semi-automatiquement des bases lexicales multilingues à lexies et acceptions interlingues. De plus, cette thèse propose une catégorisation des critères d'évaluation de la qualité des BDLM, ainsi que les mesures correspondantes. Cette stratégie a été implémentée dans Jeminie, un système logiciel adaptable qui permet d'implémenter à la fois des méthodes de structuration de BDLM et des mesures de qualité, sous la forme de modules logiciels réutilisables. Des compositions arbitraires de ces modules peuvent être définies par un lexicologue dans un langage de haut niveau d'abstraction, ce qui permet d'adapter facilement la structuration et l'évaluation de qualité en fonction des objectifs du lexicologue et des ressources disponibles sans nécessiter de connaissances en programmation. L'intérêt de cette approche a été validé expérimentalement: la qualité des BDLM obtenues est meilleure par combinaison de techniques qu'avec chaque technique antérieure utilisée seule.
This paper presents the design and construction of the PolyU Treebank, a manually annotated Chinese shallow treebank. The PolyU Treebank is based on shallow annotation where only partial syntactical structures within sentences are annotated. Guided by the Phrase-Standard Grammar proposed by Peking University, the PolyU Treebank has been designed and constructed to provide a large amount of annotated data containing shallow syntactical information and limited semantic information for use in natural language processing (NLP) research. This paper describes the relevant design principles, annotation guidelines, and implementation issues, including the achievement of high quality annotation through the use of well-designed annotation workflow and effective post-annotation checking tools. Currently, the PolyU Treebank consists of a one-million-word annotated corpus and has been used in a number of NLP research projects with promising results.
The influence of massmedia on the norm of present Czech both on lexical and on communicative-pragmatic layer.
In this paper we present a quantitative and qualitative analysis of annotation in the Hinoki treebank of Japanese, and investigate a method of speeding annotation by using part-of-speech tags. The Hinoki treebank is a Redwoods-style treebank of Japanese dictionary definition sentences. 5,000 sentences are annotated by three different annotators and the agreement evaluated. An average agreement of 65.4% was found using strict agreement, and 83.5% using labeled precision. Exploiting POS tags allowed the annotators to choose the best parse with 19.5% fewer decisions.
Dismal is a spreadsheet that works within GNU Emacs, a widely available programmable editor. Dismal has three features of particular interest to those who study behavior: (1) the ability to manipulate and align sequential data, (2) an open architecture that allows users to expand it to meet their particular needs, and (3) an instrumented and accessible interface for studies of human-computer interaction (HCI). Example uses of each of these capabilities are provided, including cognitive models that have had their sequential behavior aligned with subject’s protocols, extensions useful for teaching and doing HCI design, and studies in which keystroke logs from the timing package in Dismal have been used.
To ease the interpretation of higher order factor analysis, the direct relationships between variables and higher order factors may be calculated by the Schmid-Leiman solution (SLS; Schmid & Leiman, 1957). This simple transformation of higher order factor analysis orthogonalizes first-order and higher order factors and thereby allows the interpretation of the relative impact of factor levels on variables. The Schmid-Leiman solution may also be used to facilitate theorizing and scale development. The rationale for the procedure is presented, supplemented by syntax codes for SPSS and SAS, since the transformation is not part of most statistical programs. Syntax codes may also be downloaded from www.psychonomic.org/archive/.
A visual presentation procedure is introduced that presents target words followed by a dynamic mask until recognition. This form of stimulus degradation prolongs the word recognition process. Differences in word recognition latencies—which are usually quite small—are magnified, and thus can be more easily observed. The results of two experiments on the Internet with a total of 141 participants establish the task’s ability to magnify differences in word recognition latencies stemming from word familiarity (Experiment 1) and word prototypicality (Experiment 2). Both factors interact with stimulus degradation, but at different presentation intervals; these results are discussed as evidence for comparing models of word recognition. The new procedure can be used for assessing individual differences, such as implicit motives and self-focused attention. Further applications are discussed.
The European Language Resources Association (ELRA) was founded in 1995 with the mission of providing language resources (LR) to European research institutions and companies. In this paper we describe the background, the mission and the major activities since then.
We present a new method to describe the contextual meaning of a key word in a corpus. The vocabulary of the sentences containing this word is compared to that of the entire corpus in order to highlight the words which are significantly overutilized in the neighbourhood of this key word (they are associated in the author’s mind) and the ones which are significantly underutilized (they are mutually exclusive). This method provides an interesting tool for lexicography and literary studies as is shown by applying it to the word amour (love) in the work of Pierre Corneille, the most famous French playwright of the 17th century.
The role of language resources and language technology evaluation is now recognized as being crucial for the development of written and spoken language processing systems. Given the increasing challenge of multilingualism in Europe, the development of language technologies requires a more internationally distributed effort. This paper first describes several recent and on-going activities in France aimed at the development of language resources and evaluation. We then outline a new project intended to enhance collaboration, cooperation, and resource sharing among the international language processing research community.
Although some progress has been made on the quality of Machine Translation in recent years, there is still a significant potential for quality improvement. There has also been a shift in paradigm of machine translation, from “classical” rule-based systems like METAL or LMT1 towards example-based or statistical MT.2 It seems to be time now to evaluate the progress and compare the results of these efforts, and draw conclusions for further improvements of MT quality.
There was simply linguistics at the beginning. During the years, linguistics has been accompanied by various attributes. For example corpus one. While a name corpus is relatively young in linguistics, its content related to a language - collection of texts and speeches - is nothing new at all. Speaking about corpus linguistics nowadays, we keep in mind collecting of language resources in an electronic form. There is one more attribute that computers together with mathematics bring into linguistics - computational. The progress from working with corpus towards the computational approach is determined by the fact that electronic data with the "unlimited" computer potential give opportunities to solve natural language processing issues in a fast way (with regard to the possibilities of human being) on a statistically significant amount of data.Listing the attributes, we have to stop for a while by the notion of annotated corpora. Let us build a big corpus including all Czech text data available in an electronic form and look at it as a sequence of characters with the space having dominating status -- a separator of words. It is very easy to compare two words (as strings), to calculate how many times these two words appear next to each other in a corpus, how many times they appear separately and so on. Even more, it is possible to do it for every language (more or less). This kind of calculations is language independent -- it is not restricted by the knowledge of language, its morphology, its syntax. However, if we want to solve more complex language tasks such as machine translation we cannot do it without deep knowledge of language. Thus, we have to transform language knowledge into an electronic form as well, i.e. we have to formalize it and then assign it to words (e.g., in case of morphology), or to sentences (e.g., in case of syntax). A corpus with additional information is called an annotated corpus.We are lucky. There is a real annotated corpus of Czech -- Prague Dependency Treebank (PDT). PDT belongs to the top of the world corpus linguistics and its second edition is ready to be officially published (for the first release see (Hajič et al., 2001)). PDT was born in Prague and had arisen from the tradition of the successful Prague School of Linguistics. The dependency approach to a syntactical analysis with the main role of verb has been applied. The annotations go from the morphological level to the tectogrammatical level (level of underlying syntactic structure) through the intermediate syntactical-analytical level. The data (2 mil. words) have been annotated in the same direction, i.e., from a more simple level to a more complex one. This fact corresponds to the amount of data annotated on a particular level. The largest number of words have been annotated morphologically (2 mil. words) and the lowest number of words tectogramatically (0.8 mil. words). In other words, 0.8 million words have been annotated on all three levels, 1.5 mil. words on both morphological and syntactical level and 2 mil. words on the lowest morphological level.Besides the verification of 'pre-PDT' theories and formulation of new ones, PDT serves as training data for machine learning methods. Here, we present a system Styx that is designed to be an exercise book of Czech morphology and syntax with exercises directly selected from PDT. The schoolchildren can use a computer to write, to draw, to play games, to page encyclopedia, to compose music - why they could not use it to parse a sentence, to determine gender, number, case,...? While the Styx development, two main phases have been passed:1. transformation of an academic version of PDT into a school one. 20 thousand sentences were automatically selected out of 80 thousand sentences morphologically and syntactically annotated. The complexity of selected sentences exactly corresponds to the complexity of sentences exercised in the current textbooks of Czech. A syntactically annotated sentence in PDT is represented as a tree with the same number of nodes as is the number of the words in the given sentence. It differs from the schemes used at schools (Grepl and Karlík, 1998). On the other side, the linear structure of PDT morphological annotations was taken as it is -- only morphological categories relevant to school syllabuses were preserved.2. proposal and implementation of exercises. The general computer facilities of basic and secondary schools were taken into account while choosing a potential programming language to use. The Styx is implemented in Java that meets our main requirements -- platform-independent system and system stability.At least to our knowledge, there is no such system for any language corpus that makes the schoolchildren familiar with an academic product. At the same time, our system represents a challenge and an opportunity for the academicians to popularize a field devoted to the natural language processing with promising future.A number of electronic exercises of Czech morphology and syntax were created. However, they were built manually, i.e. authors selected sentences either from their minds or randomly from books, newspapers. Then they analyzed them manually. In a given manner, there is no chance to build an exercise system that reflects a real usage of language in such amount the Styx system fully offers.
INTRODUCTORY REMARKS: HISTORICAL LINGUISTICS AND THE DATING OF HEBREW TEXTS CA. 1000–300 B.C.E.* Ziony Zevit University of Judaism In 1927, M. H. Segal’s A Grammar of Mishnaic Hebrew (Oxford University Press, 1927, reprinted in 1958 with corrections and addenda) helped launch a new sub-discipline in historical linguistics: The History of Hebrew. In order for him to establish that Mishnaic Hebrew was a well-defined linguistic stage in the history of Hebrew meriting a description on its own terms, it was necessary to demonstrate the ways in which it was unlike Biblical Hebrew. He produced impressive lists of data illustrating that the differences between Biblical Hebrew and Mishnaic Hebrew extended to style of expression, vocabulary, and grammar, that is, phonology, morphology, and syntax. His lists illustrated that of the 1350 verbs in the Biblical Hebrew lexicon, Mishnaic Hebrew lost 250 verbs while gaining about 300 new ones. Through analysis of its lexicon, Segal showed how Aramaic semantic calques on Hebrew changed the meanings of Biblical Hebrew words that continued into Mishnaic Hebrew or how Biblical Hebrew words were replaced by Aramaic words or how new Hebrew words replaced old Hebrew words. Segal’s research indicated beyond doubt that Mishnaic Hebrew was not a debased or slightly evolved form of Biblical Hebrew. The repertoire of its linguistic norms was not described in grammars of Biblical Hebrew while its lexical resources were larger and more diverse than those of Biblical Hebrew. From an historical perspective it had to be studied on its own because it was geographically discontinuous with most of Biblical Hebrew, because the linguistic environment in which it was spoken differed significantly from that of Biblical Hebrew, and because it was a few centuries younger than Biblical Hebrew but not necessarily its direct stemmatic continuation. Historical linguistics begins by noting that living languages change. Their phonology changes as do their morphology and syntax and vocabulary when new words are introduced and old ones drop out of use or when the semantic load of individual vocables shift. Linguists have observed, on the basis of two centuries of research into many languages, that change occurs more easily and hence rapidly—when and if it occurs—in phonology and lexicon than in *!These introductory remarks were delivered November 22, 2004 before presentations by a panel of scholars dealing with the question of whether or not biblical texts can be dated linguistically. Hebrew Studies 46 (2005) 322 Zevit: Introductory Remarks morphology and syntax. But change occurs, exactly the type of changes that Segal described in 1927. Since the 1920s, work on delimiting the characteristic features of Hebrew in many of its historical periods has continued unabated, primarily at institutions in Israel, but also in some located in Europe and North America. Nowadays, scholars talk about Modern Israeli Hebrew, Haskalah Hebrew, Medieval Hebrew, Mishnaic/Tannaitic Hebrew, and, of course, Biblical Hebrew. Researchers in Israel work on all periods of Hebrew, from Biblical Hebrew through the contemporary language, whereas those outside of Israel work primarily on Hebrew from both the First and Second Temple periods, including some Mishnaic Hebrew, but more often on the Hebrew of the Dead Sea Scrolls, an ill-defined type that fits chronologically somewhere between Biblical Hebrew and Mishnaic Hebrew. A bibliographically rich summary of the achievements and the state of research in Mishnaic Hebrew is available in Moshe Bar Asher, “Mishnaic Hebrew: An Introduction,” HS 40 (1999): 115– 151. Projects aimed at refining notions about Hebrew of the First Temple period were stimulated not only by the comparative data supplied by the Ugaritic after the 1930s, but also by research into Aramaic dialects from early antiquity through the modern period, and by work on Akkadian in general and the Amarna dialects in particular. In addition, such projects benefited directly from advances in semantics and dialect studies, by studies of the living linguistic and textual traditions in diasporic Jewish communities, and by the study and analysis of newly discovered Hebrew and Aramaic inscriptions. The inscriptions proved to be of major importance because they supplied archaeologically dated, uncurated texts for linguistic analysis. Many scholars contributed to the advance of knowledge in this area and I name a few whose...
The gradient descent optimization method has been a de facto standard learning algorithm in computational models of category learning. However, it can be considered as a normative (vs. descriptive) model of human learning processes. In particular, there are three concerns associated with the learning algorithm& #x2014;namely, complexity, regularity, and context independency. In response to these limitations, the present study introduces an alternative, hypothesis-testing& #x2014;like learning algorithm on the basis of a stochastic optimization method. The new learning model, termed SCODEL, provides qualitatively simple interpretations for its implied category-learning processes. Moreover, SCODEL is the first modeling attempt to depict individually unique and context-dependent learning processes. Four simulation studies were conducted and showed that the present model has the competence to operate as several different types of learners in various plausibly real-life situations.
We describe briefly the redevelopment of Space Fortress (SF), a research tool widely used to study training of complex tasks involving both cognitive and motor skills, to be executed on currentgeneration systems with significantly extended capabilities, and then compare the performance of human participants on an original PC version of Space Fortress (SF) with the revised Space Fortress (RSF). Participants trained on SF or RSF for 10 sets of eight 3-min practice trials and two 3-min test trials. They then took tests involving retention, resistance to secondary task interference, and transfer to a different control system. They then switched from SF to RSF or from RSF to SF for 2 sets of final tests and completed rating scales comparing RSF and SF. Slight differences were predicted on the basis of a scoring error in the original version of SF used and on slightly more precise joystick control in RSF. The predictions were supported. The SF group started better but did worse when they transferred to RSF. Despite the disadvantage of having to be cautious in generalizing from RSF to SF, we conclude that RSF has many advantages, which include accommodating new PC hardware and new training techniques. A monograph that presents the methodology used in creating RSF, details on its performance and validation, and directions on how to download free copies of the system may be downloaded from www .psychonomic.org/archive/.
We examine methods for measuring performance in signal-detection-like tasks when each participant provides only a few observations. Monte Carlo simulations demonstrate that standard statistical techniques applied to ad’ analysis can lead to large numbers of Type I errors (incorrectly rejecting a hypothesis of no difference). Various statistical methods were compared in terms of their Type I and Type II error (incorrectly accepting a hypothesis of no difference) rates. Our conclusions are the same whether these two types of errors are weighted equally or Type I errors are weighted more heavily. The most promising method is to combine an aggregated’ measure with a percentile bootstrap confidence interval, a computerintensive nonparametric method of statistical inference. Researchers who prefer statistical techniques more commonly used in psychology, such as a repeated measurest test, should useγ (Goodman & Kruskal, 1954), since it performs slightly better than or nearly as well asd’. In general, when repeated measurest tests are used,γ is more conservative thand’: It makes more Type II errors, but its Type I error rate tends to be much closer to that of the traditional .05 α level. It is somewhat surprising thatγ performs as well as it does, given that the simulations that generated the hypothetical data conformed completely to thed’ model. Analyses in which H—FA was used had the highest Type I error rates. Detailed simulation results can be downloaded fromwww.psychonomic.org/archive/Schooler-BRM-2004.zip.
This study compared four common methods for scoring a popular working memory span task, Daneman and Carpenter’s (1980) reading span test. More continuous measures, such as the total number of words recalled or the proportion of words per set averaged across all sets, were more normally distributed, had higher reliability, and had higher correlations with criterion measures (reading comprehension and Verbal SAT) than did traditional span scores that quantified the highest set size completed or the number of words in correct sets. Furthermore, creation of arbitrary groups (e.g., high-span and low-span groups) led to poor reliability and greatly reduced predictive power. It is recommended that researchers score span tasks with continuous measures and avoid post hoc dichotomization of working memory span groups.
Contrasting linguistic and nonlinguistic processing has been of interest to many researchers with different scientific, theoretical, or clinical questions. However, previous work on this type of comparative analysis and experimentation has been limited. In particular, little is known about the differences and similarities between the perceptual, cognitive, and neural processing of nonverbal environmental sounds and that of speech sounds. With the aim of contrasting verbal and nonverbal processing in the auditory modality, we developed a new on-line measure that can be administered to subjects from different clinical, neurological, or sociocultural groups. This is an on-line task of sound to picture matching, in which the sounds are either environmental sounds or their linguistic equivalents and which is controlled for potential task and item confounds across the two sound types. Here, we describe the design and development of our measure and report norming data for healthy subjects from two different adult age groups: younger adults (18–24 years of age) and older adults (54–78 years of age). We also outline other populations to which the test has been or is being administered. In addition to the results reported here, the test can be useful to other researchers who are interested in systematically contrasting verbal and nonverbal auditory processing in other populations.
There is a strong relationship between evaluation and methods for automatically training language processing systems, where generally the same resource and metrics are used both to train system components and to evaluate them. To date, in dialogue systems research, this general methodology is not typically applied to the dialogue manager and spoken language generator. However, any metric for evaluating system performance can be used as a feedback function for automatically training the system. This approach is motivated with examples of the application of reinforcement learning to dialogue manager optimization, and the use of boosting to train the spoken language generator.
Ontologies are recognised as important tools, not only for effective and efficient information sharing, but also for information extraction and text mining. In the biomedical domain, the need for a common ontology for information sharing has long been recognised, and several ontologies are now widely used. However, there is confusion among researchers concerning the type of ontology that is needed for text mining , and how it can be used for effective knowledge management, sharing, and integration in biomedicine. We argue that there are several different ways to define an ontology and that, while the logical view is popular for some applications, it may be neither possible nor necessary for text mining. We propose a text-centered approach for knowledge sharing, as an alternative to formal ontologies. We argue that a thesaurus (i.e. an organised collection of terms enriched with relations) is more useful for text mining applications than formal ontologies.