Papers reviewed and determined not to be word norm studies. Use the flag icon to report errors or suggest re-inclusion.
16504 papers
The article discusses the specific features of the English used in The Gambia by looking at the phonetic and lexical markers that distinguish Gambian English from the other national varieties of West African English. The study shows that Gambian English has a number of established and exclusive features owing to the formation of a national norm and the influence of certain indigenous languages, yielding a national quasi-standard easy to identify.
The article is devoted to the problem of regulation and unification of law terminology in modern Ukrainian \nlegislation. Analyzing with this purpose the text of the Criminal code of Ukraine, author pays attention to cases \nof non-compliance with lexical, grammar, stylistic norms, replication of Russian syntactical constructions in \nseveral items of Code. Author also gives recommendations of the correct usage of terms and words.
The aim of this study is to show how clusteranalysis can shed light on very complexvariation in a transitional dialect zone ineastern Finland. In the course of history thisarea has been on the border between Sweden andRussia and the population has clearly been oftwo kinds: the Savo people and the Karelians.It is a well-known fact that there is variationamong these dialects, but the spread and extentof the variation has not been demonstrated previously.The idiolects of the area were studied in thelight of ten phonological and morphologicalfeatures. The material consisted of recordingsof 198 idiolects, totalling around 195 hoursand representing 19 parishes. The variation wasanalysed using hierarchical cluster analysis.While the analysis showed the extent of thevariation between idiolects and parishes, italso demonstrated how the effects of the oldparishes, borders and settlements are stillvisible in the dialects. On the parish level,the data formed clear clusters that correspondwith the main dialects in the area and itssurroundings. On the idiolect level, however,the speakers from the surrounding areas formedfairly homogenous clusters but the idiolectsfrom the Savonlinna area were spread acrossalmost all clusters.
This paper presents a study that attributes verb serialization in the interlanguage of Vietnamese-speaking ESL learners to language transfer and, furthermore, puts forward the view that such transfer bears a resemblance to substrate influence in creoles with serial verb constructions (SVCs). In a task that elicited English causatives through pictures representing the causation of events, a subset of the Vietnamese-speaking participants in this study produced a number of serial-type constructions that reflected lexicosemantic aspects of causative SVCs in Vietnamese. Speakers of Hindi-Urdu, a nonserializing language used for comparative purposes, did not produce any equivalents. Additionally, serial-type constructions with second verbs (V2s) representing a result (e.g., cook butter melt ) predominated at lower levels of lexical proficiency, whereas serials with make and a result (e.g., make broken ) were more evenly distributed across proficiency levels. One inference based on the results is that certain serials are eliminated early in the acquisition process through positive evidence obtained via English input, whereas others continue to appear beyond the elementary level because of misleadingly similar constructions in the input. A comparison of the proficiency-based transfer of “ cook butter melt ” serials in this study and the inferred transfer of SVCs in creolization suggests that, whereas transfer processes in the two contexts are congruent in certain ways (often resulting from the exigencies of communication, limited access to the TL, and linguistic convergence), the processes diverge because of differences in target norms and input conditions. The latter two factors provide one explanation for why SVC-related transfer effects were limited to a subgroup of Vietnamese-speaking participants in this study.
MLRy 98.1, 2003 195 Valdman and Thomas A. Klinger examine the phonology and grammar of Louisiana Creole and, in so doing, document how tenuous is the line of demarcation between LC and Cajun French. They also return to the issue of possible African influence on the Creole but conclude that the parallels with vernacular and regional French forms suggest a process of convergence between the varieties spoken by the white settlers and the languages of the slave population in the genesis of the Creole. Margaret M. Marshall outlines the historical emergence of LC and argues for the early existence in the colony of a linguistic continuum and a level of social fluidity in the early years that allowed the Creole to spread to the white community. The fuzziness of the boundaries separating LC and CF is highlighted by Klinger, Michael D. Picone, and Valdman in their description of the lexicon: they argue for a fundamental unity among the French-related varieties, justifying the combination of the lexicons under the single label of 'Louisiana French'. Their study underlines the decline in lexical productivity of LC and CF, where internal processes are yielding to borrowings from English. Jacques Henry and Becky Brown, respectively, chart the socio-political and historical background and the shift in language attitudes that led to the effortsto revitalize French in Louisiana through the creation of the CODOFIL (Council for the Development of French in Louisiana), the ensuing debate on Cajun identity,and the emergence of a more informal revival movement focusing on issues of culture rather than language. Brown considers the issue of developing a Louisiana French norm and underlines the need for language planners to take into account speakers' desire to identifywith a local variety rather than an external norm. Finally, Barry Jean Ancelet documents the present state of research on Louisiana's folklore and folklife. The scope of this volume in reality goes beyond that suggested by the title, since four remaining chapters deal with related varieties outside Louisiana. Julianne Maher provides a description of Saint Barth patois and Saint Barth Creole. Karin Flikeid and Raymond Mougeon deal with Acadian and Ontarian French respectively, and Pierre Rezeau examines lexical links between Louisiana French and varieties within France, and proposes a methodology for comparative lexicographical research. The broad sweep of this volume, the wealth of useful and comprehensive information it contains, and the range of issues addressed have much to offer scholars interested in French language and culture outside France as well as Creolists and linguists interested in language loss and maintenance. University of Leeds Marie-Anne Hintze Jehan etBlonde, Poems and Songs. By Philippe de Remi. Ed. by Barbara N. SargentBaur. (FauxTitre, 201) Amsterdam and Atlanta, GA: Rodopi. 2001. ii +586 pp. $127.50. ISBN 90-420-1504-7. This very substantial volume completes Barbara Sargent-Baur's edition of the com? plete works of Philippe de Remi, begun in her 1999 edition of his Roman de la Manekine, prepared in collaboration with Alison Stones and Roger Middleton and also published by Rodopi. The two volumes together provide a modern edition of all of the works attributed to this poet, including not only the narrative works and shorter pieces found in Bibl. Nat. fr.1588 but also the songs transmitted in the chansonnier Bibl. Nat. fr. 24406 and the Resveries found in Bibl. Nat. fr.837. As in the edition of La Manekine, Sargent-Baur has provided an English translation of Jehan et Blonde, but she has not translated any of the other pieces. All ofthe texts, however, are equipped with notes commenting primarily on the establishment of the text and on linguistic peculiarities, though occasionally on parallels with other texts. Jehan et Blonde tells the story of a young man, son of an impoverished knight, who succeeds through his wits and his chivalric abilities in marrying the daughter 196 Reviews of the earl of Oxford. The narrative is fairly conventional in its representation of court life and amorous intrigue, lacking the lurid elements of La Manekine with its tale of treachery, incest, infanticide, and self-mutilation, and its eventual miraculous resolution. None the less, Philippe tells a lively story with plenty of entertaining moments...
Article Sprachwissen im Konflikt. Sprachliche Zweifelsfälle zwischen Linguistik und Sprachnorm. Arbeitsgruppe auf der Jahrestagung der Deutschen Gesellshaft für Sprachwissenschaft 2003 mit dem Rahmenthema Sprache. Wissen. Sprachwissenschaft, München 26.28. Februar 2003 [Languageknowledge in conflict. Borderline cases between linguistics and linguistic norms. Working group at the annual conference of the German Linguistic Society 2003 with the topic Language. Knowledge. Linguistics.] was published on March 25, 2004 in the journal Zeitschrift für germanistische Linguistik (volume 31, issue 2).
In this paper, we present a modular incremental statistical model for English full parsing. Unlike other full parsing approaches in which the analysis of the sentence is a uniform process, our model separates the full parsing into shallow parsing and sentence skeleton parsing. In shallow parsing, we finish POS tagging, Base NP identification, prepositional phrase attachment and subordinate clause identification. In skeleton parsing, we use a layered feature-oriented statistical method. Modularity possesses the advantage of solving different problems in parsing with corresponding mechanisms. Feature-oriented rule is able to express the complex lingual phenomena at the key point if needed. Evaluated on Penn Treebank corpus, we obtained 89.2% precision and 89.8% recall.
Empirical study of the syntax-prosody relation is hampered by the fact that current prosodic models are essentially linear, while syntactic structure is hierarchical. The present contribu-tion describes a syntax-prosody comparison heuristic based on two new algorithms: Time Tree Induction, TTI, for building a prosodic treebank from time-annotated speech data, and Tree Similarity Indexing, TSI) for comparing syntactic trees with the prosodic trees. Two parametrisations of the TTI algorithm, for different tree branching conditions, are applied to sentences taken from a read-aloud narrative, and compared with parses of the same sentences, using the TSI. In addition, null-hypotheses in the form of flat bracketing of the sentences are compared. A preference for iambic (heavy rightmost branch) grouping is found. The resulting quantitative evidence for syntax-prosody
<p>The OmniPaper project has implemented three information retrieval prototypes in the area of electronic news publishing. One prototype uses SOAP as communication protocol between the central system and a number of distributed news archives. The second prototype uses an RDF metadata database, enabling direct metadata queries to the central system. Finally the Topic Map prototype uses query expansion and semantic linking for smart metadata search. The Topic Map prototype enhances thesearch experience by implementing a knowledge layer that combines the semantic content of a lexical database, consisting of concepts and keywords, with a metadata-set of newspaper articles. The linking between both is currently implemented at the level of keywords but will be developed at the level of concepts in the final prototype. The knowledge layer has been designed from a Topic Map point of view, although the XTM syntax has not been used to avoid performance issues. The consortium’s adopted view on information publishing and retrieval considers querying and navigation as two very related actions that can both be captured under the name “search for relevant information”. Navigation forces the user to followpredefined paths whereas querying enables the user to look freely for a suitable starting point. The query and navigation functionality is provided through a web engine and is build on top of the information structure of the knowledge layer.</p>
In this paper we will present work carried out lately on the 50,000 words Italian Spontaneous Speech Corpus called AVIP, under national project API, made available for free download from the website of the coordinator, the University of Naples. We will concentrate on the tuning of the parser for Italian which had been previously used to parse 100,000 words corpus of written Italian within the National Treebank initiative coordinated by ILC in Pisa. The parser receives as input the adequately transformed orthographic transcription of the dialogues making up the corpus, in which pauses, hesitations and other disfluencies have been turned into most likely corresponding punctiation marks, interjections or truncation of the word underlying the uttered segment.\nThe most interesting phenomenon we will discuss is without any doubts "overlapping", i.e. a speech event in which two people speak at the same time by uttering actual words or in some cases nonwords, when one of the speakers, usually the one which is not the current turntaker, interrupts the current speaker.\nThis phenomenon takes place at a certain point in time where it has to be anchored to the speech signal but in order to be fully parsed and subsequently semantically interpreted, it needs to be referred semantically to a following turn.
This study examines expressional styles and plans for the education of Korean language in computer chatting rooms, which are relevant to hypertext, as a part of preparing the contents and methods of hypertext expression education\n\n First, because of the 'anonymity' of computer chatting rooms, people can expression their feelings without concealing. On the other hand, anonymity causes flaming and undermines linguistic morality. Conversations in computer chatting rooms occur through direct feedback using 'interaction' Because many meetings and conversations in computer chatting rooms are improvised, they are more impersonal than sincere, and politeness and rules necessary for conversation are often ignored Language in chatting rooms is exchanged by texts, but they contain oral elements, which are expressed through the mouth, as well as non-oral elements That is, the language. is mixed with oral words and textual words.\n\n Telecommunication language is used based on strategies for preserving efficiency, those for preserving expressivity, and psychological factors. Although some view it negatively saying that it destroys linguistic norms, its positive side of creative utilization of texts is not negligible, The reason that such telecommunication language is perceived negatively is its influence on everyday language. Concerning the influence, further research is required.\n\n Research related to the education of expression in computer chatting rooms is mostly focused on telecommunication language. As the present researcher investigated, in addition, contents in Writing and Korean Language Life in the 7th Education Curricula also deal mainly with the destruction of linguistic norms by telecommunication language and problems in linguistic morality.\n\n Thus, the researcher proposes plans for the education of Korean Language related to expressions in computer chatting rooms as follows. ① It is necessary to understand telecommunication language positively, regarding it as a social dialect. ② In a sense, abbreviations, emoticons (emotion icons) and symbolic words are language creation, which are necessary for efficient conversations in communication. ③ Except the examples presented in ②, telecommunication language should not be transferred to everyday language, and for this students must learn grammar intensively and have an ability to distinguish virtual worlds from the real world ④ Students must be given not only theoretic education on the characteristics of oral and textual words rot also one related to conversational expressions in computer chatting rooms in 'Speech' and 'Narration' classes.
The aim of this article is to show that the role of legal terminology in juridical discourse can adequately be determined in a framework of a "juridical textwork" model only. This implies to analyse their role under pragmatic and intertextual aspects. lt will be contended that neither semantics nor the wording of the law can be employed as a starting point for an adequate interpretation of legal language in public discourse. Of crucial importance is rather the question of how the legal text and social reality can combine to constitute the legal norm (which is more than just the legal text). This is illustrated with the example of German court decisions on the issue of sit-ins that were organised by the peace movement in the l970s and l980s to block the access to American army bases. lt will be demonstrated that in the process of putting the coercion law in concrete normative terms through different courts, specific legal terms are semanticly modified and adjusted to specific language use, hence constituting "semantic battles" and linguistic norm conflicts in the juridical discourse.
In den letzten Jahren ist die Zahl der verfgbaren linguistisch annotierten Korpora stndig gewachsen. Zu den bekanntesten gehren das Brown-Korpus, das Susanne-Korpus, die Penn-Treebank, das Negra-Korpus, das Tiger-Korpus und die im Zusam-
MLRy 98.4, 2003 1007 are either part of a system which just happens to be differentfrom the norm or an attempt to conjugate a lexical item which is simply alien to the system? Theoretically at least, Sablayrolles appears to give the status of neologism to all new forms, what? ever their source or motivation. (Perhaps intentional, conscious creation could be a useful criterion to constrain an otherwise over-generous approach to the concept?) A detailed list of contents helps to guide the reader through this tightly structured book, although one could have wished for a fuller index. (It is restricted to some of the technical terms used in the text.) The bibliography, too, is incomplete, curiously excluding reference to the works by Jorifand Meyer which provide the data for two of his corpora. The appendix, however, is generous in listing all of the neologisms discussed, with details of their function and context. This alone makes forfascinating reading. Particularly ingenious examples are la markethique as an ever-topical social issue, and multiconjugal, referringto a serial spouse, both from Le Monde. Queen Mary, London Hilary Wise A New Life ofDante. By Stephen Bemrose. Exeter: University of Exeter Press. 2000. xxi + 249pp.?42-5o(pbk?i4.99). ISBN0-85989-583-1 (pbk0-85989-584-x). This book aims to provide an account of Dante's life combined with discussion of his writings that is accessible to both university students and non-specialist readers. The eleven chapters take Dante's life in a series of chronological sequences, from childhood and the meeting with Beatrice, to his involvement in politics and exile, and on to his post-exilic years. In most chapters, the chronological segments encompass Dante's literary activities, and the relevant works are usually given detailed com? mentary. Chapter 8 provides a separate survey of the Comedy. The volume, which includes a bibliography and guide to further reading, is written in a fluent and often lively style, even if there is the occasional moment when the tone is a little stilted ('about which more anon'; 'So are all great men of learning'). At a documentary level, the writing of a biography of Dante is a forbidding task, not least because many of the sources are provided by Dante himself, or else mediated often in highly partial form by later chroniclers, commentators, and biographers. On the whole, Bemrose finds his way through the material with a straightforward but sure-footed approach that gives considerable emphasis to Dante's socio-political and intellectual context. Dante emerges as a thinker,a political animal, and a consummate literary artist. The socio-political emphasis in particular makes fora portrait of Dante which, in its general outlines, is at times reminiscent of Bruni's 'documentary' life. Bemrose adds nothing new to Petrocchi's authoritative modern life, does not find space for Padoan's more recent attempts to backdate the Inferno, and offersno real new evidence or arguments to resolve issues of dating and attribution. Bemrose is not, however, to be faulted in these respects, forthe book does fulfil(and often admirably) its own aims and it is very well tailored to its readership, especially at an undergraduate level. One of the greatest strengths of the volume is its acute awareness of the knowledge gaps in its intended readers: Bemrose repeatedly offershelpful points of clarification on general contexts (especially intellectual and historical), specific issues (Guelphs and Ghibellines), and important terminology (e.g. plenary indulgence, curia, vernacular) that are often taken for granted in other general and introductory works on Dante. What is more, the summaries that he gives of Dante's works, es? pecially the Comedy, the De vulgari eloquentia, and the Convivio, are clear, accurate, judicious, and often highly readable. Almost inevitably, of course, any general account of this kind will elicit quibbles and prompt specialists to point out omissions: some may well have expected more 1008 Reviews discussion of the Vita nuova and its relationship to the Comedy; others may have wished for a fuller account of the philosophical and scientific content of the Rime petrose (which is surprising, given Bemrose's expertise in such matters); still others a more focused and critically informed discussion of allegory...
Previous research (Aarts & Dijksterhuis, 2003) has shown that mental representations of situational norms (e.g., behaving quietly in libraries) and corresponding overt behaviors are capable of being automatically activated. Two experiments extended this line of research by investigating the conditional role of the tendency to conform to social norms in these effects. Participants explored a picture of a library and were given the goal to visit this library or not. Accessibility of representations of normative behavior was assessed in a lexical decision task. In the first experiment, individual differences in conformity to social norms were measured, whereas in the second experiment conformity was primed. Results indicated that the goal to visit the environment caused participants to automatically access representations of normative behavior. Importantly, in both experiments conformity was shown to moderate these accessibility effects: Automatic access to representations of normative behavior emerged when conformity tendencies were active.
Abstract. The problem of Prepositional Phrase (PP) attachment disambiguation consists in determining if a PP is part of a noun phrase, as in He sees the room with books, or an argument of a verb, as in He fills the room with books. Volk has proposed two variants of a method that queries an Internet search engine to find the most probable attachment variant. In this paper we apply the latest variant of Volk’s method to Spanish with several differences that allow us to attain a better performance close to that of statistical methods using treebanks. 1
Acronyms are a very dynamic area of the lexicon of many languages. A hybrid, modular methodology for the acquisition of acronyms is presented, which uses an existing acronym-expansion matching component, and machine learning in two separate phases for the identification of long-distance acronym definition patterns.The resulting system, using Support Vector Machines (SVM) is trained on 600 news stories from the Wall Street Journal component of the Penn Treebank corpus using a number of lexical, syntactic, and acronym-expansion matching features. Statistical cooccurrence information for acronym-expansion pairs is extracted from search engine hit counts.The system achieves Fβ=1=92.38% on 400 news stories from the same source and has good asymptotic efficiency, making it adequate for the automatic extraction of acronyms even from noisy sources, such as newspaper text.
This article explores the possibilities of automatic extraction of both surface and valency frames of Czech verbs. First, it is clearly documented that the data from Prague Dependency Treebank is not sufficient for collecting enough examples of verb frames to build a large scale lexicon. As a solution, an approach to pick nice examples of sentences from any texts is suggested and thoroughly described. A new scripting language to simplify the selection of sentences based on linguistic criteria was implemented and its main concepts are presented here, too. Also the problems of extracting surface and valency frames from the collected data are addressed and illustrated on real corpus data.
We explore learning prepositionalphrase attachment in Dutch, to use it as a filter in prosodic phrasing. From a syntactic treebank of spoken Dutch we extract instances of the attachment of prepositional phrases to either a governing verb or noun. Using cross-validated parameter and feature selection, we train two learning algorithms, IB1 and RIPPER, on making this distinction, based on unigram and bigram lexical features and a cooccurrence feature derived from WWW counts. We optimize the learning on noun attachment, since in a second stage we use the attachment decision for blocking the incorrect placement of phrase boundaries before prepositional phrases attached to the preceding noun. On noun attachment, IB1 attains an F-score of 82; RIPPER an F-score of 78. When used as a filter for prosodic phrasing, using attachment decisions from IB1 yields the best improvement on precision (by six points to 71) on phrase boundary placement.
Machine translation engines draw on various types of databases. This paper is concerned with Arabic as a source or target language, and focuses on lexical databases. The non-concatenative nature of Arabic morphology, the complex structure of Arabic word-forms, and the general use of vowel-free writing present a real challenge to NLP developers. We show here how and why a stem-grounded lexical database, the items of which are associated with grammar-lexis specifications – as opposed to a root-&-pattern database –, is motivated both linguistically and with regards to efficiency, economy and modularity. Arguments in favour of databases relying on stems associated with grammar-lexis specifications (such as DIINAR.1 or the Arabic dB under development at SYSTRAN), rather than on roots and patterns, are the following: (a) The latter include huge numbers of rule-generated word-forms, which do not actually appear in the language. (b) Rule-generated lemmas – as opposed to existing ones – are widely under-specified with regards to grammar-lexis relations. (c) In a Semitic language such as Arabic, the mapping of grammar-lexis specifications that need to be associated with every lexical entry of the database is decisive. (d) These specifications can only be included in a stem-based dB. Points (a) to (d) are crucial and in the context of machine translation involving Arabic.
This study presents results from a corpus-based analysis of the expression of attitude, emotion, certainty and doubt (stance) in a large corpus of British and American conversation. Stance marker frequencies were assessed through an automated procedure for identifying stanced lexical items occur-ring in particular grammatical frames. The frequencies were analyzed with a multi-variate statistical procedure known as factor analysis which identifies co-occurrence patterns (factors). These factors can be understood to be the most salient moods of stance. Three factors were identified as character-istic: 1) informal AFFECT (American dialect-based), 2) boulomaic planning (American work-based) versus small talk (British dialect-based), and 3) hedged opinion (British dialect-based). Social norms were identified by examining the factors in light of discourse context and interpersonal relationships among speakers. Cross-cultural misunderstandings seemed particularly likely in work contexts, where Americans preferred boulomaic verbs (want, need), and British preferred evidentials (know, maybe). Differences in informal adult conversations are also potentially important, where Americans used many more affect markers (such as love, crazy). More work on pragmatic or functional domains using multi-variate analysis is proposed in the conclusion.
This article reports the outcome of a publicly funded research project titled "Redesign of the British Sign Language (BSL) Notation System with a New Font for Use in ICT," which ran from September 2000 to December 2001. The aim of the project was to redesign the British Sign Language variant of Stokoe notation (as used in the BSL/English Dictionary) for practical use in information technology systems and software, such as lexical databases, word-processing packages, and teaching and learning applications. The project�s objectives tackled design issues, not sign linguistic system-level problems. The project resulted in two new type designs for writing BSL Stokoe, united into a larger type family called William C., in memory of William C. Stokoe. I anticipate that the new type design proposals will make the provision of searchable BSL Stokoe notation in database designs a possible next step in lexicographic and other sign linguistic research projects.
One morning each of us received a phone call from Ed Hovy. Are you sitting down? he asked. He told us that as a way to combat conference overload, and to promote interaction among communities, a joint conference had been proposed to combine HLT and NAACL. A diverse oversight committee had been formed, and according to Ed, this committee had been able to agree on two people -- and only two people -- as program co-chairs, because together we represented all of the vested interests. Marti was meant to represent the standards and tastes of the NAACL and the SIGIR crowds, and Mari the speech community, and both have been working on research contracts with HLT funders. Ed told us that if either of us said no, the entire enterprise would come crashing down. There are few better ways to convince busy people to become program co-chairs. Throughout the process, Ed provided the vision for and the drive behind this conference. We salute him for making this idea a reality, and for his enthusiastic and energetic phone calls that kept everything going. This is an exciting time for research in human language technologies. After years of relative calm, the field seems suddenly to be moving by leaps and bounds. Evidence of this can be found in our conference panel on Preparing for a Surprise Language (and as embodied in the short paper Desperately Seeking Cebuano). This panel will discuss the experiences of several groups of researchers, who at the behest of DARPA, acquired and developed language resources for an entirely new language within a span of only 10 days. This experiment took place in March of 2003, and the language in question was Cebuano, a language spoken in the Philippines. Participants successfully collected a large body of lexical and textual resources and developed a range of tools, including stemmers and POS taggers. (In June, DARPA will announce a new surprise language.) The existence of a variety of language resources, combined with advances in statistical analysis and modeling techniques, is resulting in fast-paced improvements in the field. parsers can now produce syntax trees for long sentences with high accuracy and great speed. Advances are starting to be made in automated semantic analysis. Great strides are being made in the sophistication and coverage of question answering systems. Speech recognition systems have achieved suficiently high accuracy that it is now possible to do retrieval, information extraction and topic tracking on spoken documents. Large and growing collections of text and speech corpora -- and the promise of much more from the web -- have enabled many of these advances. New developments in weakly supervised and unsupervised learning algorithms are critical for taking advantage of many new data sources, and hence this was chosen as a special theme of the conference. Lexical resources such as FrameNet, WordNet, PropBank, MeSH, and the Penn TreeBank also play prominent roles in HLT advances. As a field, human language technologies research should use, as motivation and guide, an understanding of the linguistic and cognitive bases of language. The invited talk by Dr. Elissa Newport, entitled Statistical language learning: Mechanisms for language acquisition in human learners, should help enlighten the community by informing us about the latest in psycholinguistic research. We received 162 submissions for full papers, of which 37 were accepted, resulting in a highly competitive acceptance rate of 22%. For the short (late-breaking) papers track, we received 80 submissions, of which 41 were accepted (2 later withdrawn). Some of these will be presented as short talks, and others as posters. Seventeen demonstrations will be shown. We were fortunate to be able to accept 15 papers that addressed the conference theme of unsupervised and weakly supervised methods. We also encouraged papers that described techniques that cross over or combine NLP, speech and/or IR, and several of the papers demonstrate this kind of crossover. The full paper reviewing was done using a two-tier system. First, two first-tier reviewers read every paper. Then a third reviewer, known as the meta-reviewer, wrote their own review. Finally, the meta-reviewer summarized these reviews and introduced additional comments. In some cases, the meta-reviewer instigated discussion among the first-tier reviewers to work out controversial issues. The meta-reviewers also attended the program committee meeting in which all the papers were discussed and acceptances were decided. For the short papers, each short paper received at least two reviews. Those papers whose reviewers disagreed, or which received middling scores, were subsequently reviewed by a member of the program committee and the program co-chairs. Paper submission and reviewing was done online using Marti's conference reviewing software (Conga), which she updated for this conference. Marti also maintained the conference website.
One morning each of us received a phone call from Ed Hovy. Are you sitting down? he asked. He told us that as a way to combat conference overload, and to promote interaction among communities, a joint conference had been proposed to combine HLT and NAACL. A diverse oversight committee had been formed, and according to Ed, this committee had been able to agree on two people -- and only two people -- as program co-chairs, because together we represented all of the vested interests. Marti was meant to represent the standards and tastes of the NAACL and the SIGIR crowds, and Mari the speech community, and both have been working on research contracts with HLT funders. Ed told us that if either of us said no, the entire enterprise would come crashing down. There are few better ways to convince busy people to become program co-chairs. Throughout the process, Ed provided the vision for and the drive behind this conference. We salute him for making this idea a reality, and for his enthusiastic and energetic phone calls that kept everything going. This is an exciting time for research in human language technologies. After years of relative calm, the field seems suddenly to be moving by leaps and bounds. Evidence of this can be found in our conference panel on Preparing for a Surprise Language (and as embodied in the short paper Desperately Seeking Cebuano). This panel will discuss the experiences of several groups of researchers, who at the behest of DARPA, acquired and developed language resources for an entirely new language within a span of only 10 days. This experiment took place in March of 2003, and the language in question was Cebuano, a language spoken in the Philippines. Participants successfully collected a large body of lexical and textual resources and developed a range of tools, including stemmers and POS taggers. (In June, DARPA will announce a new surprise language.) The existence of a variety of language resources, combined with advances in statistical analysis and modeling techniques, is resulting in fast-paced improvements in the field. parsers can now produce syntax trees for long sentences with high accuracy and great speed. Advances are starting to be made in automated semantic analysis. Great strides are being made in the sophistication and coverage of question answering systems. Speech recognition systems have achieved suficiently high accuracy that it is now possible to do retrieval, information extraction and topic tracking on spoken documents. Large and growing collections of text and speech corpora -- and the promise of much more from the web -- have enabled many of these advances. New developments in weakly supervised and unsupervised learning algorithms are critical for taking advantage of many new data sources, and hence this was chosen as a special theme of the conference. Lexical resources such as FrameNet, WordNet, PropBank, MeSH, and the Penn TreeBank also play prominent roles in HLT advances. As a field, human language technologies research should use, as motivation and guide, an understanding of the linguistic and cognitive bases of language. The invited talk by Dr. Elissa Newport, entitled Statistical language learning: Mechanisms for language acquisition in human learners, should help enlighten the community by informing us about the latest in psycholinguistic research. We received 162 submissions for full papers, of which 37 were accepted, resulting in a highly competitive acceptance rate of 22%. For the short (late-breaking) papers track, we received 80 submissions, of which 41 were accepted (2 later withdrawn). Some of these will be presented as short talks, and others as posters. Seventeen demonstrations will be shown. We were fortunate to be able to accept 15 papers that addressed the conference theme of unsupervised and weakly supervised methods. We also encouraged papers that described techniques that cross over or combine NLP, speech and/or IR, and several of the papers demonstrate this kind of crossover. The full paper reviewing was done using a two-tier system. First, two first-tier reviewers read every paper. Then a third reviewer, known as the meta-reviewer, wrote their own review. Finally, the meta-reviewer summarized these reviews and introduced additional comments. In some cases, the meta-reviewer instigated discussion among the first-tier reviewers to work out controversial issues. The meta-reviewers also attended the program committee meeting in which all the papers were discussed and acceptances were decided. For the short papers, each short paper received at least two reviews. Those papers whose reviewers disagreed, or which received middling scores, were subsequently reviewed by a member of the program committee and the program co-chairs. Paper submission and reviewing was done online using Marti's conference reviewing software (Conga), which she updated for this conference. Marti also maintained the conference website.
This mainly technological paper first provides a description of the web site called PapiLex. This first part shows how a file containing XML-structured lexical entries can be managed as a lexical database by using the Document Object Model (DOM) API. PapiLex offers the three essential management functions: creation, modification and deletion of a lexical entry. In a second part, two tools for entering Unicode-formatted text are presented: one for browsers having HTML 4 and JavaScript 1.2 capability and one for Microsoft Word. Such tools can be necessary for the minority languages which have no virtual keyboard embedded in the operating systems. 1 Starting point for building a lexical base Inside the Papillon project, the construction of a lexical base for a new language may take several different ways depending on where the author has to start. The following situations may occur regarding the availability of lexical resources1,2: • no dictionary exists, • a paper dictionary exists, • an electronic form of a dictionary exists, • a lexical database exists. In the last two cases, the question is to re-work the existing data so they meet the Papillon format and to fill the remaining fields. Tools are available for recycling electronic dictionary, (e.g. Nguyen 1998). Here, we will suppose that there is no preexisting dictionary or that its existence is limited to a paper dictionary. In such cases, the lexical entries have to be typed entirely. Among the 1: In addition to the existence of lexical resource, the script used for the language has also to be in Unicode and a font has to exist for it. Actually, the scripts of a number of minority languages are not in Unicode at the moment (e.g. Shan, Tai Dam, Mon). For some of them, fonts that really work are still missing as it is the case for Khmer. 2: In case there are existing data, property rights have to be looked at to say the resource is available. different ways in which this question can be handled, we chose a particular approach that consists in creating directly the Papillon formatted base by using generic and multiplatform Internet browsers. Section 2 will present how a standard browser can be used for this task3 (PapiLex mockup) and section 3 will show that a simple JavaScript program can provide a virtual keyboard that produces Unicode text. In section 4, another issue, less directly related to Papillon, will also be presented as it provides a very practical alternative for creating Unicodeencoded entries. It addresses a Windowsspecific tool for typing Unicode text in Microsoft Word when no standard keyboard is existing yet. The software was developed for the Lao language but can be applied to others. 2 The PapiLex mockup
This paper describes a national cooperational project STO, which has the aim of developing a large-scale Danish lexical database for computational use. We discuss some organisational aspects of the project and present the current project structure and main activities. Further, we discuss in more detail some of the linguistic issues that have required thorough consideration before a large-scale encoding could be initiated, encompassing topics such as the morphological encoding of compounds and proper names, as well as the syntactic encoding of there-constructions, phrasal verbs and reflexive verbs.
We present a procedure for building the lexicon of a transfer-based machine translation system for NPs and PPs from English to Basque. The agglutinative nature of Basque implies the need for morphosyntactic information, so the lexicon was created by automatically extracting this necessary information from a bilingual dictionary and a wide coverage lexical database. The system translates with 83% precision, 18% better than the first approach which used the raw bilingual dictionary.
Medical records have been evolving from the traditional paper-based records to digital ones, from the method of dictating reports and transcription to voice recognition systems. The transition to digital operations will not be complete until we have the ability to combine voice recognition with automated indexing of texts. This paper introduces the methods we used to evaluate existing voice recognition software programs and presents NOMINDEX, a system that turns a medical text into MeSH codes, using the French ADM lexical database. Those systems were applied to 28 patient discharge summaries in French, produced after a coronarography, and extracted from the MENELAS corpus of texts. Using the best configuration for voice recognition, the rate of accurate recognition exceeds 98 percent. Among the indexing concepts assigned by NOMINDEX, 25 percent were not pertinent and 12 percent of the relevant concepts were missing. Most errors were related to confusion between common language and medical language, and to the coverage of the ADM lexical database. Best results would be expected with a more comprehensive lexical resource In addition, only 3 percent of the errors generated by inadequate voice recognition that remained in the configuration that performed better, impacted on automatic indexing by NOMINDEX.
The database Profil has been set up tooffer readers studying modern literarymanuscripts a reference tool to identifywatermarked papers. In the study of writers'drafts as in artists' sketches, the differentkinds of papers used provide valuableinformation on the genesis of a work of art andwatermarks, when they exist, are the bestvisible hint allowing us to identify paper. Amultimedia database, with digitized images moreprecise than usual traced design, seems to beappropriate to register, visualize, and comparemodern watermarked papers. Besides itsusefulness for specialists, such a databasebearing on modern manuscripts should also beconceived in a didactic perspective, as it isoriented towards literary scholars who are notparticularly familiar with the history of modern paper. In this paper we present the database Profilwhich includes a set of digitized images from acollection of betagraphies made by thereproduction service of the National FrenchLibrary. Then we explain problems of databasenormalization when human sciences areinvolved.
This paper presents work which extends previous corpus-based work on training Machine Learning Algorithms to perform Prepositional Phrase attachment. Besides recreating others&apos; experiments to see how algorithms&apos; performance changes with the number of training examples and using n-fold cross-validation to produce more accurate error rates, we implemented our own vanilla Machine Learning Algorithms as a comparison. We also had people perform exactly the same task as the Machine Learning Algorithms to indicate whether the way forward lies in improving Machine Learning Algorithms or in improving the data sets used to train Machine Learning Algorithms. The results from all these experiments feed into our other work transforming the Penn TreeBank into a more useful resource for training Machine Learning Algorithms to do Prepositional Phrase attachment.
Word-level alignments of bilingual text (bitexts) are not an integral part of statistical machine translation models, but also useful for lexical acquisition, treebank construction. and part-of-speech tagging. The frequent occurrence of divergences, structural differences between languages, presents a great challenge to the alignment task. We resolve some of the most prevalent divergence cases by using syntactic parse information to transform the sentence structure of one language to bear a closer resemblance to that of the other language. In this paper, we show that common divergence types can be found in multiple language pairs (in particular, we focus on English-Spanish and English-Arabic) and systematically identified. We describe our techniques for modifying English parse trees to form resulting sentences that share more similarity with the sentences in the other languages; finally, we present an empirical analysis comparing the complexities of performing word-level alignments with an without divergence handling. Our results suggest that divergence-handling can improve word-level alignment.
The Prague Dependency Treebank (PDT, as described, e.g., in (Hajic, 1998) or more recently in (Hajic, Pajas and Vidova Hladka, 2001)) is a project of linguistic annotation of approx. 1.5 million word corpus of naturally occurring written Czech on three levels (“layers”) of complexity and depth: morphological, analytical, and tectogrammatical. The aim of the project is to have a reference corpus annotated by using the accumulated findings of the Prague School as much as possible, while simultaneously showing (by experiments, mainly of statistical nature) that such a framework is not only theoretically interesting but possibly also of practical use. In this contribution we want to show that the deepest (tectogrammatical) layer of representation of sentence structure we use, which represents “linguistic meaning” as described in (Sgall, Hajicova and Panevova, 1986) and which also records certain aspects of discourse structure, has certain properties that can be effectively used in machine translation1 for languages of quite different nature at the transfer stage. We believe that such representation not only minimizes the “distance” between languages at this layer, but also delegates individual language phenomena where they belong to whether it is the analysis, transfer or generation processes, regardless of methods used for performing these steps.
Comprehensive computational lexicons areessential to practical natural languageprocessing (NLP). To compile such computationallexicons by automatically acquiring lexicalinformation, however, we previously requiresufficiently large corpora. This study aims atpredicting the ideal size of suchautomatic-lexical-acquisition oriented corpora,focusing on six specific factors: (1) specificversus general purpose prediction, (2)variation among corpora, (3) base forms versus inflected forms, (4) open class items,(5) homographs, and (6) unknown words.Another important and related issue withregard to predictability has something to dowith data sparseness. Research using theTOTAL Corpus reveals serious datasparseness in this corpus. This, again, pointstowards the importance and necessity ofreducing data sparseness to a satisfactorylevel for the automatic lexical acquisition andreliable corpus predictions. The functions ofpredicting the number of tokens and lemmas in acorpus are based on the piecewisecurve-fitting algorithm. Unfortunately, thepredicted size of a corpus for automaticlexical acquisition is too astronomicalto compile it by using presently existingcompiling strategies. Therefore, we suggest apractical and efficient alternative method. Weare confident that this study will shed newlight on issues such as corpus predictability,compiling strategies and linguisticcomprehensiveness.
530 SEER, 8o, 3, 2002 I999). Zubova is well aware of the metatextual qualities of Russian postmodernism, and points to the intertextualgames of varioustexts. In addition to the more familiarnames, Zubova introducesher audience to lesser known authors, such as Vladimir Strochkov, Ian Satunovskii, and Vladimir Erl'.Unfortunately, Zubova's studydoes not contain any biographical detailsof the authorsshe quotes. Such an appendixwould be a usefultool for assessingthe spreadof linguisticdeviations, from the point of view of age groups, regional variations and the aesthetic preferences of the poets. It is difficultto assess whether some of the deviations from established linguistic norms were intentional, or derive from the contemporary sloppy usage of Russianlanguage that isparticularlynoticeable in post-Soviet Russianmedia. Krivulinand Shvarts,for example, are philologistsby training,and therefore are more inclined to have playful appropriationof some idioms, or absurd examplesof Soviet newspeak. As Zubova's study demonstrates, numerous poetic experiments reflect on the fluid state of the Russian language itself. In this respect, Zubova's discussion of the satirical elements relating to the concept of gender in contemporary Russian poetry is particularlyrewarding. Zubova's examples from Russian poetry reveal, for example, the uncertaintiesrelatingto gender of animals. Thus, some poets use the feminine form of the noun koshka (cat) with the additional note that it is used in their poem as a noun of masculine gender. Such examples are both amusing and obscure. Zubova suggeststhat contemporary Russian poets are struggling to revive the use of the neuter gender that otherwise has been steadily disappearing from the standard language (p. 301). Many poems quoted by Zubova use Church Slavonic constructions such as esi, bekh,izhe, byst', and sut', to name just a few (pp. 208-39). Another archaicelement that occurs in contemporarypoetry is the Double Nominative Case discussedon pages 368-70. Zubova also refers to the influence of the English language on the contemporary Russian language in relationto the use of nouns as describingwords (pp. 362-68): for example, Rus'-zemlia, shved-koroleva. Zubova'sbook mightbe seen as an attemptto mergelinguisticanalysiswith cultural anthropology (especially Durkheim's theory), since it claims that various linguistic experiments express an archetypal collective conscience (p. 7). The extensive bibliographywill be highly appreciatedby readersof the book, as well as Zubova's reassuring message that the poetic experiments reveal the dynamic process of language evolution and expose the mnemonic abilities of semiotic signs that enable the user to restore forgotten forms and contexts. Department ofFrench andRussian ALEXANDRA SMITH University ofCanterbugy, NewZealand Menzel,Birgit.Biirgerkrieg um Worte. Die russische Literaturkritik derPerestroika. Bohlau Verlag, Cologne, Weimar, Vienna, 200I. xi + 420 pp. Illustrations. Notes. Bibliography.Tables. Index. DM 89.80:?4s.9 I. How do you investigate a type of text like literary criticism? Not literary criticism in the general sense, but in a sense which native English speakers REVIEWS 53I normally do not use, and which the Russians, among others, do? Literary criticism in this particular sense means topical, applied writing that greatly influencesthe generalpublic and explicitlyevaluatesthe worksdiscussed. Two possibilities present themselves. We can analyse the contents for consistencyand substanceof the profferedargumentsand evaluations.Or we can focus on pragmatic aspects of literary criticism (for example, by underscoring, collating, and evaluating statistics illustrating the effects of criticismon the purchasingbehaviourof readers). In her Rostock habilitation thesis, Birgit Menzel chooses neither of the above possibilities.And with good reason, for she does not focus merely on the work of a single literary critic, or on the treatment of a single work by various schools or camps of literarycriticism. Rather, she assignsherselfthe more comprehensive task of presenting an 'overview of Russian literary criticism between I986 and I993' (p. I). The reader hoping for a detailed criticism of criticism in Menzel's book will thereforebe disappointed. What she does offer is a survey of groupings and argument-patternsalong with traditions and developments. In other words, Menzel seeks to grasp the changing structuresof literary criticism as an important form of literary communication underthe conditions of the disintegratingSoviet empire. The enticing path of marketingresearchremainsunfeasiblefor the simple reason that salesrecordsfor the period of the Soviet planned economy and the years which immediatelyfollowed do not at all provide an accuratepictureof what readersreallywanted. In her statisticalassessments,Menzel thereforeconfines herselfto extremely illuminatinginformationon developments regardingthe circulationnumbersfor the...
The term ‘standard language ideology’ as described by Milroy and Milroy (1998) and Lippi-Green (1994) characterises a particular set of beliefs about language. Such beliefs are typically held by populations of economically developed nation states where processes of standardisation have operated over a considerable time to produce an abstract set of norms-lexical, grammatical and (in spoken language) phonological-popularly described as constituting a standard language. The same beliefs also emerge, somewhat transformed by local histories and conditions, in these states’ colonies and ex-colonies. For example, in all the countries discussed by contributors to Cheshire (1991) where English has been imported (I confine my comments in this chapter to English-speaking hegemonies), beliefs about language similar to those discussed below can be found. These are reported in a sizeable literature; for example Platt and Weber (1980) and Gupta (1994) both describe the operation of a British-style standard language ideology in Singapore. Although debates about standard English are a staple of the British press (in the United States the most contentious ideological debates are usually slightly differently oriented, as we shall see), experts and laypersons alike have just about as much success in locating a specific agreed spoken standard variety in either Britain or the United States as have generations of children in locating the pot of gold at the end of the rainbow.
The increasing prestige of medicine as a science, accompanied by the social rise of the doctor, in eighteenth-century France is well documented. What I would like to argue here, however, is that there exists a correspondence between the establishment of medicine as an independent field of study in eighteenth-century France and the increasing use and influence of an autonomous form of medical discourse, namely, the aphorism, in this period. This is not so much a question of the language used by the more renowned doctors of the day but of a form of discourse deeply imbued and associated with medical practice. (It is nonetheless true that certain famous physicians combined medical and literary roles. For instance, Theophile Bordeu intervenes significantly in Diderot's Le Reve d'Alembert, and Vicq d'Azyr, Marie-Antoinette's doctor, was elected to the Académie Française in 1788 in a sort of social consecration or medical discourse, implicitly incorporating his medical figure and figures into the socio-linguistic norms of 'le bon usage’ promoted by the Académie itself.) Yet what interests me particularly here is the insinuation of the medical aphorism itself into other fields of late eighteenth-century discourse, notably those of literature and politics, the traditional domains of the maxim.
Cet article a pour objectif d’interroger les relations existant entre le système linguistique (entendu au sens large, comme englobant l’ensemble des règles ou régularités qui sous-tendent la production et l’interprétation des énoncés attestés), et la culture, et plus précisément les normes communicatives en vigueur dans une société donnée. Après avoir envisagé un certain nombre de faits langagiers (unités lexicales, formes honorifiques, actes de langage et formules rituelles) qui portent manifestement la trace de ces normes culturelles sous-jacentes, l’auteure en conclut qu’il est dans une certaine mesure possible de reconstituer à partir de ces traces l’« ethos communicatif » de la société considérée, mais que ce travail de reconstitution ne va pas sans rencontrer un certain nombre de difficultés, qui sont passées en revue.
We present statistical models for morphological disambiguation in agglutinative languages, with a specific application to Turkish. Turkish presents an interesting problem for statistical models as the potential tag set size is very large because of the productive derivational morphology. We propose to handle this by breaking up the morhosyntactic tags into inflectional groups, each of which contains the inflectional features for each (intermediate) derived form. Our statistical models score the probability of each morhosyntactic tag by considering statistics over the individual inflectional groups and surface roots in trigram models. Among the four models that we have developed and tested, the simplest model ignoring the local morphotactics within words performs the best. Our best trigram model performs with 93.95% accuracy on our test data getting all the morhosyntactic and semantic features correct. If we are just interested in syntactically relevant features and ignore a very small set of semantic features, then the accuracy increases to 95.07%.
Note: This is a personal view of the experience gathered, and does not necessarily reflect the opinion of the Floresta team The Floresta Sinta(c)tica http://acdc.linguateca.pt/treebank/ • A collaboration project between VISL (Southern Denmark University) and Linguateca (SINTEF); project leaders: Eckhard Bick & Diana Santos • Bosque: 1,427 syntactically analysed and revised trees (1,405 distinct sentences, 36,408 tokens, ca. 34,256 words), automatically created • Started October 2000, stopped December 2001, some research still being done as of today See Afonso et al. (2002) at LREC'2002
XML(eXtensible Markup Language)is a standard which was issued by W3C(wolld wide web consortium) in February,1998.It defines the data structure by means of an open self-description.Also a group of rules of XML can be used to set up Markup Language in accordance with specific applied fields.Furthermore,based on XML,cnXML is a linguistic norm of electronic business affairs,which is keeping with the commercial habits,tradition and circuit in the continent of China.cnXML supplies a set of unified,flexible,open and extensible exchange pattern of data,which makes all trade partners including such commercial organizations as buyers,sellers,runners and intermediary etc.carry out commercial activities conveniently through Internet.So cnXML is not only a standardized norm but also a premise of electronic business development in China.
Abstract There seems to be an increase in the use of foreignisms in Finnish business translation: a new linguistic norm, which incorporates foreignisms appears to be developing alongside traditional Finnish usage. Yet how justified is the use of foreignisms in business translations? Here is an analysis of a Nokia document: the author explores the status of this text as translation and offers criteria for justified and unjustified foreignisms in Finnish business translation.
This talk provides an overview of current work in my research group on the syntactic annotation of the Tubingen corpus of spoken German and of the German Reference Corpus (Deutsches Referenzkorpus: DEREKO) of written texts. Morpho-syntactic and syntactic annotation as well as annotation of function-argument structure for these corpora is performed automatically by a hybrid architecture that combines robust symbolic parsing with finite-state methods (&amp;quot;chunk parsing &amp;quot; in the sense Abney) with memory-based parsing (in the sense of Daelemans). The resulting robust annotations can be used by theoretical linguists, who are interested in large-scale, empirical data, and by computational linguists, who are in need of training material for a wide range of language technology applications. To aid retrieval of annotated trees from the treebank, a query tool VIQTORYA with a graphical user interface and a logic-based query language has been developed. VIQTORYA allows users to query the treebanks for linguistic structures at the word level, at the level of
This paper describes Grammar Learning by Partition Search, a general method for automatically constructing grammars for a range of parsing tasks. Given a base grammar, a training corpus, and a parsing task, Partition Search constructs an optimised probabilistic context-free grammar by searching a space of nonterminal set partitions, looking for a partition that maximises parsing performance and minimises grammar size. The method can be used to optimise grammars in terms of size and performance, or to adapt existing grammars to new parsing tasks and new domains. This paper reports an example application to optimising a base grammar extracted from the Wall Street Journal Corpus. Partition Search improves parsing performance by up to 5.29%, and reduces grammar size by up to 16.89%. Parsing results are better than in existing treebank grammar research, and compared to other grammar compression methods, Partition Search has the advantage of achieving compression without loss of grammar coverage.
We tested a computer-based procedure for assessing reader strategies that was based on verbal protocols that utilized latent semantic analysis (LSA). Students were given self-explanation—reading training (SERT), which teaches strategies that facilitate self-explanation during reading, such as elaboration based on world knowledge and bridging between text sentences. During a computerized version of SERT practice, students read texts and typed self-explanations into a computer after each sentence. The use of SERT strategies during this practice was assessed by determining the extent to which students used the information in the current sentence versus the prior text or world knowledge in their self-explanations. This assessment was made on the basis of human judgments and LSA. Both human judgments and LSA were remarkably similar and indicated that students who were not complying with SERT tended to paraphrase the text sentences, whereas students who were compliant with SERT tended to explain the sentences in terms of what they knew about the world and of information provided in the prior text context. The similarity between human judgments and LSA indicates that LSA will be useful in accounting for reading strategies in a Web-based version of SERT.