Papers reviewed and determined not to be word norm studies. Use the flag icon to report errors or suggest re-inclusion.
18265 papers
Recent evaluation techniques applied to corpus-based systems have been introduced that can predict quantitatively how well surface realizers will generate unseen sentences in isolation. We introduce a similar method for determining the coverage on the Fuf/Surge symbolic surface realizer, report that its coverage and accuracy on the Penn TreeBank is higher than that of a similar statistics-based generator, describe several benefits that can be used in other areas of computational linguistics, and present an updated version of Surge for use in the NLG community.
In the context of the Papillon project, which aims at creating a multilingual lexical database (MLDB), we have developed Jeminie, an adaptable system that helps automatically building interlingual lexical databases from existing lexical resources. In this article, we present a taxonomy of criteria for evaluating a MLDB, that motivates the need for arbitrary compositions of criteria to evaluate a whole MLDB. A quality measurement method is proposed, that is adaptable to different contexts and available lexical resources.
The architecture of a lexical database in which multilingual semantic networks would be stored requires the incorporation of e xible mechanisms and services, which would enable the efcient navigation within and across lexical data. We report on WordNet Management System (WMS), a system that functions as the interconnection and communication link between a user and a number of interlinked WordNets. Semantic information is being accessed through a distributed network of servers, forming a large-scale multilingual semantic network.
The current research explored the processes that predominate during the anticipation of an emotionally salient event. Experiment 1 (N536), employed three different conditional stimuli followed by pictorial pleasant, unpleasant or neutral unconditioned stimuli. Half the participants were trained with visual CSs, the other half with tactile CSs. In the group trained with visual CSs, startle eyeblinks were larger and faster during CSs that were paired with unpleasant pictures than CSs paired with neutral or pleasant pictures respectively, indicating an affect startle pattern. This linear trend was not found in the group trained with tactile CSs. Experiment 2 (N564) aimed to investigate whether the affective pattern found in the startle data in Experiment 1 could also be found using a behavioural measure of emotion. This time participants’ reaction time during a post-experimental affective priming taskwas used as dependantmeasure to assess the presence of emotional learning. Instead of a simple differential conditioning task, an occasion setting paradigm was employed and participants were trained using either a feature positive or feature negative design with pleasant or unpleasant picture USs. For participants trained with unpleasant USs, valence ratings collected before and after conditioning training suggested the presence of emotional learning, whereas no such pattern was found for participants trained with pleasant USs. These findings were not confirmed in the priming data.
The Multilingual Dictionary of Lexicographical Terms (MDLT) is a prototype of a linguistic database. The main goal of this project is to provide all users, especially student-linguists, with a tool for increasing lexicographical competence. We based our project on an investigation of users’ needs and demands. Through our rich entry structure (including translation equivalents, set phrases, synonyms, antonyms, etc.) the system combines large volume with quality of information and convenient query. Our initial motivation was a lack of dictionaries that serve to underline differences between Russian and English terms in the subject field of lexicography. Therefore, our dictionary contains entries in two languages – Russian and English. However, our system can also support other European languages such as German, French, Italian, Danish, etc.
Expert observers discriminate elementary features (e.g., colour, orientation) of salient objects outside their current focus of attention, but not complex features such as T/L shape or colour arrangement. Surprisingly, natural scene category behaves like an elementary feature in this context, in that scenes with vehicles, animals, etc. are successfully classified in a dual-task situation (Li et al., 2002). To confirm and extend this finding, we used a similar dual-task paradigm and presented a natural scene in the near periphery (8 to 14 eccentricity) simultaneously with an attention-demanding task near fixation (seven rotated Ts/Ls). Visual persistence was controlled by a particularly effective form of masking. For scenes of animals or vehicles (black and white with matched luminance scale), categorization performance exhibited little or no attentional cost, i.e., dual- and single-task performance were comparable. Thus, scene categorization under dual-task conditions can be explained neither by inadequate masking nor by trivial colour cues. In a second experiment, observers categorized scenes from the International Affective Image System database (IAPS, Lang et al., 1995) as “pleasant” or “unpleasant”. Although absolute performance was now lower (∼75% of the level reached under ideal viewing conditions), there was no significant attentional cost. As low-level features do not identify affective content, this implies some comprehension of scene gist. A control experiment highlighted the contrast between natural scenes and geometric stimuli in the dual-task situation (and also confirmed the peripheral absence of attention), in that observers failed to discriminate the colour arrangement of a peripherally presented ‘pill’ (half red, half green, inclined ±45 ). Li FF et al. (2002) Rapid natural scene categorization in the near absence of attention. PNAS 99: 9596ff. Lang PJ et al. (1995) IAPS: Technical Manual and Affective Ratings. Gainsville, FL.
The motivation of the Papillon project is to encourage the development of freely accessible Multilingual Lexical Resources by way of online collaborative work on the Internet. For this, we developed a generic community website originally dedicated to the diffusion and the development of a particular acception based multilingual lexical database.
An essential component of Language Engineering (LE) tools are verb class descriptors that provide information about the relations of the predicates to their arguments. The production of computationally tractable language resources necessitates the assignment of types of predicate-argument relations to a great variety of verb-centered structures: it is necessary to define not only the initial, canonical valency frame of a great number of verb lexemes, but also the diathesis alternations, which reflect the real-life usage of verbs. This paper describes the implementation of descriptors of the valency properties of Bulgarian verbs used in the production of a syntactic treebank of Bulgarian. The descriptors are based on available LE resources for Bulgarian: a verb subcategorization model implemented in the lexical data base that is used; a chunk grammar that recognizes verb form patterns. Predictive models are built and applied in a grammar that annotates grammatical relations inferred from the combination of morphosyntactic and shallow syntactic processing cues. The real significance of this particular processing is the resolution, in relation to the valency properties of many verbs, of the discrepancy or the contradiction between the verb lexicon specifications and the verb syntagmatic realization. 1.
This paper presents the construction of a manually annotated Chinese shallow Treebank, named PolyU Treebank. Different from traditional Chinese Treebank based on full parsing, the PolyU Treebank is based on shallow parsing in which only partial syntactical structures are annotated. This Treebank can be used to support shallow parser training, testing and other natural language applications. Phrase-based Grammar, proposed by Peking University, is used to guide the design and implementation of the PolyU Treebank. The design principles include good resource sharing, low structural complexity, sufficient syntactic information and large data scale. The design issues, including corpus material preparation, standard for word segmentation and POS tagging, and the guideline for phrase bracketing and annotation, are presented in this paper. Well-designed workflow and effective semiautomatic and automatic annotation checking are used to ensure annotation accuracy and consistency. Currently, the PolyU Treebank has completed the annotation of a 1-million-word corpus. The evaluation shows that the accuracy of annotation is higher than 98%. 1
We are developing a Natural Language Generation (NLG) system that generates texts tailored for the reading ability of individual readers. As part of building the system, GIRL (Generator for Individual Reading Levels), we carried out an analysis of the RST Discourse Treebank Corpus to find out how human writers linguistically realise discourse relations. The goal of the analysis was (a) to create a model of the choices that need to be made when realising discourse relations, and (b) to understand how these choices were typically made for "normal" readers, for a variety of discourse relations. We present our results for discourse relations: concession, condition, elaboration-additional, evaluation, example, reason and restatement. We discuss the results and how they were used in GIRL.
We present a detailed investigation of the challenges posed when applying parsing models developed against English corpora to Chinese. We develop a
A series of studies was conducted to analyse the relationships between rating behaviour, rater goals and contextual variables, namely unit or class climate. Cleveland and Murphy (1992) suggested that errors and inter-rater disagreements in performance ratings could be understood in terms of differences in unit climate and the goals pursued by raters. In two studies, substantial correlations were found between self-rated goals and evaluations of instructor's performance; our second study provided evidence that goals measured before raters have an opportunity to observe performance are related to ratings obtained after observing instructor performance. A third study investigated Murphy and Cleveland's (1995) proposal that contextual variables, specifically organisational climate, affect rating behaviour. To test this hypothesis, data reflecting perceptions of the climate of college level courses and ratings of instructor performance were collected. Ratings of participative and co-operative climates showed a strong relationship with student ratings of instructors' performance. Average correlations between climate and ratings are substantially larger than those between rater goals and ratings; the mediation hypotheses tested in this study were not supported. Results of the three studies are discussed in relation to past research.
This paper presents new methods for extracting semantic knowledge from collections of annotated images. The proposed methods include novel automatic techniques for extracting semantic concepts by disambiguating the senses of words in annotations using the lexical database WordNet, using both the images and their annotations, and for discovering semantic relations among the detected concepts based on WordNet. Another contribution of this paper is the evaluation of several techniques for visual feature descriptor extraction and data clustering in the extraction of semantic concepts. Experiments show the potential of integrating the analysis of both images and annotations for improving the performance of the word-sense disambiguation process. In particular, the accuracy improves 4-15% with respect to the baselines systems for nature images.
The paper describes ongoing work on the evaluation of methods for extracting collocation candidates from large text corpora. Our research is based on a German treebank corpus used as gold standard. Results are available for adjective+noun pairs, which proved to be a comparatively easy extraction task. We plan to extend the evaluation to other types of collocations (e.g., PP+verb pairs).
We present a formalization of the valency theory (Panevová, 1974) that fits the stratificational representation scheme used in the Prague Dependency Treebank. The notion of a lexicon as a repository of “static ” (invariable, or context-independent) source of information is formally presented; a different type of lexicon is used at every layer of sentence representation, with a formal link to this representation (and thus, annotation). In order to show how such a lexicon can be used in the annotation process itself, we describe also an automatic procedure using information from a valency lexicon for partial annotation of a corpus at the tectogrammatical layer. When adding nodes into tectogrammatical representation of sentences, we substantially increase recall at the cost of a small decrease of precision. 1
This paper presents an quantitative comparison of the se- mantic role inventories in two corpus-based resources for computational linguistics, TreeBank and FrameNet, along with the semantic roles from two knowledge representation frameworks, OpenCyc and Conceptual Graphs. In addition, experiments are done to illustrate the degree to which knowledge from the corpus-based resources can be transferred to the knowledge base resources. One benet of this is that new annotations in eect can be generated. In particular, there will eectiv ely be tagged data of the OpenCyc and Conceptual Graphs semantic roles, suitable for use in inferring the role usages.
現在入手可能な解析器と言語資源を用いて中国語解析を行った場合にどの程度の精度が得られるかを報告する. 解析器としては, サポートベクトルマシン (Support Vector Machine) を用いたYamChaを使用し, 中国語構文木コーパスとしては, 最も一般的なPenn Chinese Treebankを使用した. この両者を組み合わせて, 形態素解析と基本句同定解析 (base phrase chunking) の2種類の解析実験を行った. 形態素解析実験の際には, 一般公開されている統計的モデルに基づく形態素解析器MOZとの比較実験も行った. この結果, YamChaによる形態素解析精度は約88%でMOZよりも4%以上高いが, 実用的には計算時間に問題があることが分かった. また基本句同定解析精度は約93%であった.
A maximum entropy model in Chinese BaseNP recognition is used in this paper The open test on Chinese TreeBank, the public corpus, indicates the average recall and precision of 87 43% and 88 09% respectively with limited knowledge (text itself and its POS tag) Because of the incomparability of Chinese BaseNP recognition results, the same algorithm is applied in English BaseNP recognition The test on TREEBANK Ⅱ shows that the recall and precision are 93 31% and 93 04%, which are close to the state of the art This not only proves the availability of the algorithm, but also indicates its language independence
this paper I will discuss a framework for semantics which allows us to record truth-conditional and compositional analyses as dependency-style corpus annotations in a direct and fine-grained fashion. This method eliminates the need for a semantic representation formalism by decomposing semantic information into simple statements about (word or morpheme) tokens. A collection of such data would form a new kind of linguistic treebank. The main purpose of this article is to show that the present approach makes it possible to combine formal semantics and corpus-oriented study of language use in new and interesting ways. The methodology of this framework, which I call Token Dependency Semantics (TDS, Dahllf [4]), is in several respects different from the common one(s) in traditional formal semantics. TDS nevertheless delivers a fairly conventional (but ontologically restrained) analysis of truth-conditional meaning
In this paper, we describe an approach to annotate the propositions in the Penn Chinese Treebank. We describe how diathesis alternation patterns can be used to make coarse sense distinctions for Chinese verbs as a necessary step in annotating the predicate-structure of Chinese verbs. We then discuss the representation scheme we use to label the semantic arguments and adjuncts of the predicates. We discuss several complications for this type of annotation and describe our solutions. We then discuss how a lexical database with predicate-argument structure information can be used to ensure consistent annotation. Finally, we discuss possible applications for this resource.
Choosing the statistical model is the key problem in statistical parsing. Statistical model lies in the core of NLP parsing. This paper investigates 4 primary statistical parsing models, namely PCFG, history-based model, cascading parsing model and head-driven parsing model, and compares their performances in a 10000 Chinese treebank. The analysis based on the experiment were shown in the paper. The comparative study of these models can be exploited to build the practical and effective Chinese parser.
In this paper, we present a modular incremental statistical model for English full parsing. Unlike other full parsing approaches in which the analysis of the sentence is a uniform process, our model separates the full parsing into shallow parsing and sentence skeleton parsing. In shallow parsing, we finish POS tagging, Base NP identification, prepositional phrase attachment and subordinate clause identification. In skeleton parsing, we use a layered feature-oriented statistical method. Modularity possesses the advantage of solving different problems in parsing with corresponding mechanisms. Feature-oriented rule is able to express the complex lingual phenomena at the key point if needed. Evaluated on Penn Treebank corpus, we obtained 89.2% precision and 89.8% recall.
The article studies the particular features of the dynamics of lexical norms in Ukrainian language on the \nmaterials of dictionaries and mass-media. \nIt analyses the process of vocabulary enrichment by new lexical units, the phenomena of semantic \ntransformation and stylistic transposition.
In this paper we show how the trees in the Penn treebank can\nbe associated automatically with simple quasi-logical forms. Our approach is based on combining two independent strands of work: the first is the observation that there is a close correspondence between quasi-logical forms and LFG f-structures [van Genabith and Crouch, 1996]; the second is the development of an automatic f-structure annotation algorithm for the Penn treebank [Cahill et al, 2002a; Cahill\net al, 2002b]. We compare our approach with that of [Liakata and Pulman, 2002].
Acronyms are a very dynamic area of the lexicon of many languages. A hybrid, modular methodology for the acquisition of acronyms is presented, which uses an existing acronym-expansion matching component, and machine learning in two separate phases for the identification of long-distance acronym definition patterns.The resulting system, using Support Vector Machines (SVM) is trained on 600 news stories from the Wall Street Journal component of the Penn Treebank corpus using a number of lexical, syntactic, and acronym-expansion matching features. Statistical cooccurrence information for acronym-expansion pairs is extracted from search engine hit counts.The system achieves Fβ=1=92.38% on 400 news stories from the same source and has good asymptotic efficiency, making it adequate for the automatic extraction of acronyms even from noisy sources, such as newspaper text.
Machine translation engines draw on various types of databases. This paper is concerned with Arabic as a source or target language, and focuses on lexical databases. The non-concatenative nature of Arabic morphology, the complex structure of Arabic word-forms, and the general use of vowel-free writing present a real challenge to NLP developers. We show here how and why a stem-grounded lexical database, the items of which are associated with grammar-lexis specifications – as opposed to a root-&-pattern database –, is motivated both linguistically and with regards to efficiency, economy and modularity. Arguments in favour of databases relying on stems associated with grammar-lexis specifications (such as DIINAR.1 or the Arabic dB under development at SYSTRAN), rather than on roots and patterns, are the following: (a) The latter include huge numbers of rule-generated word-forms, which do not actually appear in the language. (b) Rule-generated lemmas – as opposed to existing ones – are widely under-specified with regards to grammar-lexis relations. (c) In a Semitic language such as Arabic, the mapping of grammar-lexis specifications that need to be associated with every lexical entry of the database is decisive. (d) These specifications can only be included in a stem-based dB. Points (a) to (d) are crucial and in the context of machine translation involving Arabic.
Among the most consistent findings in the warnings literature is the so-called 'familiarity effect.' Research has shown that the more familiar an individual is with a product or situation the less likely he or she is to notice, read, recall, or comply with hazard communications. The effect has been found across numerous product types and situations using various operational definitions of familiarity and measures of warning effectiveness. However, research has also shown that subjective familiarity ratings are not highly correlated with actual product experience. Thus, individuals must be capable of developing a false or exaggerated sense of familiarity. One possible source of this exaggerated familiarity is exposure to product advertising.\n\nThree experiments were conducted to investigate whether the familiarity effect can be produced from exposure to product advertising. The relationships between advertising exposure and perceived familiarity and between perceived familiarity, perceived safety and warning effectiveness were examined. Experiment 1 explored participants' attitudes and beliefs about well-known and obscure brands of household, consumer products and sought to determine how past, direct product experience influences those attitudes and beliefs. Experiments 2 and 3 examined how the number of advertising exposures and the safety-related content of advertisements influence attitudes and beliefs about the advertised products and the effectiveness of on product warnings.\n\nResults of Experiment 1 revealed that past experience can not fully explain consumers' attitudes and beliefs about household, consumer products. Experiments 2 and 3 showed that advertising influences perceived product familiarity and knowledge. While there was a trend of greater perceived safety with increased ad exposures, the effect was not significant. No effects of advertising on warning recall were found. Implications for the design of product advertisements and product packaging as well as directions for future research are discussed.
Crosslinguistically vocatives are an underexplored linguistic phenomenon and in different languages they can be highly idiosyncratic and complex (Levinson, 1987, p.71). Therefore, the problem, which is discussed in this paper, is not a language-specific one, in spite of the fact that most of the languages have their own repositories for marking the role of the addressee in the communicative utterances. In our opinion this linguistic phenomenon needs its adequate treatment in HPSG because of three main reasons: 
 
 The vocative is supposed to be present on two levels: syntax and pragmatics. Therefore it needs more elaborate interpretation on the interface side, which, in HPSG, is more developed for morphology/syntax and syntax/semantics than syntax/pragmatics. Note that a challenge for the theory is the semantic weight of the vocatives with respect to the head sentence. 
 It will be useful for HPSG-oriented implementations, especially treebanks and dialogue systems. 
 On prosodic grounds the vocatives are often viewed as being 'side or extended parts' of the sentence and therefore - very close to the parenthetical constructions. From our point of view, both phenomena are pragmatic and hence, the treatment of vocative, presented here, could be generalized to cover other phenomena of pragmatic nature. 
 
 In our work the vocatives are viewed through the possibility of the integration/separation of their pragmatic, syntactic and semantic properties.
In den letzten Jahren ist die Zahl der verfgbaren linguistisch annotierten Korpora stndig gewachsen. Zu den bekanntesten gehren das Brown-Korpus, das Susanne-Korpus, die Penn-Treebank, das Negra-Korpus, das Tiger-Korpus und die im Zusam-
We investigated the reliability and validity of a video-based method of measuring the magnitude of children’s emotion-modulated startle response when electromyographic (EMG) measurement is not feasible. Thirty-one children between the ages of 4 and 7 years were videotaped while watching short video clips designed to elicit happiness or fear. Embedded in the audio track of the video clips were acoustic startle probes. A coding system was developed to quantify from the video record the strength of the eye-blink startle response to the probes. EMG measurement of the eye blink was obtained simultaneously. Intercoder reliability for the video coding was high (Cohen’sκ = .90). The average within-subjects probe-by-probe correlation between the EMG- and video-based methods was .84. Group-level correlations between the methods were also strong, and there was some evidence of emotion modulation of the startle response with both the EMG- and the video-derived data. Although the video method cannot be used to assess the latency, probability, or duration of startle blinks, the findings indicate that it can serve as a valid proxy of EMG in the assessment of the magnitude of emotion-modulated startle in studies of children conducted outside of a laboratory setting, where traditional psychophysiological methods are not feasible.
El estudio experimental de la emoción requiere de estímulos que evoquen en una forma confiable reacciones psicológicas y fisiológicas que varien sistemáticamente sobre el rango de emociones de acuerdo a las dimensiones de valencia (agradable o desagradable), activación (excitado o calmado) y dominancia (alta y baja) (Lang, Bradley, Cuthbert, 1999). A pesar de que los correlatos neurales de las emociones básicas han sido investigados, la organización neural de las "emociones morales" en el cerebro humano no se conocen bien. El objetivo de la presente investigación fue obtener un grupo de estimulos diferenciados (fotografías) y caracterizarlos en términos de su valencia afectiva, activación, dominancia, y contenido moral, en una población mexicana. Se seleccionaron fotografías que representan escenas con una carga emocional amplia como violaciones morales (escenas de guerra, asaltos físicos, etc), escenas aversivas sin connotación moral (tumores, cuerpos mutilados) y escenas naturales (toallas, mesas, puertas, etc. ). Los sujetos evaluaron cada fotografía de acuerdo a su valencia, activación, dominancia y contenido moral (ausente o extremo). Para la evaluación, se utilizó la Escala Internacional Self-Assessment Maniki Affective Rating System desarrollada por Lang (1980). Se discute las implicaciones de los datos, para el estudio de las emociones y del juicio moral.
This paper describes the use of clustering at three stages within a larger research effort to identify semantic frames used in English automatically. The first of two tasks within this effort has been the identification of sets of semantically related verb senses that invoke a common semantic frame. Within this task, clustering has been used both to build sets of verb senses with the potential of invoking a common semantic frame and then to merge sets with a high degree of overlap. The paper is organized as follows: Section 2 introduces frame semantics. Section 3 outlines the methodology used to identify sets of semantically related verb senses that invoke a common semantic frame, while section 4 presents the specific clustering algorithm used within that process. Section 5 discusses the use of this clustering algorithm for the identification of semantically related verbs in two machine-readable lexical resources: the machine-readable version of the Longman Dictionary of Contemporary English (LDOCE, 1978 edition) and WordNet, an online lexical database (http://www.cogsci.princeton.edu/-wn; version 1.7.1 has been used for the work reported here). Section 6 presents the use of clustering to merge overlapping sets of verb senses formed in previous steps. Section 7 discusses the results of these clusterings, paying particular attention to the effect of LDOCE's restricted defining vocabulary on the clustering process.
In this paper we will present work carried out lately on the 50,000 words Italian Spontaneous Speech Corpus called AVIP, under national project API, made available for free download from the website of the coordinator, the University of Naples. We will concentrate on the tuning of the parser for Italian which had been previously used to parse 100,000 words corpus of written Italian within the National Treebank initiative coordinated by ILC in Pisa. The parser receives as input the adequately transformed orthographic transcription of the dialogues making up the corpus, in which pauses, hesitations and other disfluencies have been turned into most likely corresponding punctiation marks, interjections or truncation of the word underlying the uttered segment.\nThe most interesting phenomenon we will discuss is without any doubts "overlapping", i.e. a speech event in which two people speak at the same time by uttering actual words or in some cases nonwords, when one of the speakers, usually the one which is not the current turntaker, interrupts the current speaker.\nThis phenomenon takes place at a certain point in time where it has to be anchored to the speech signal but in order to be fully parsed and subsequently semantically interpreted, it needs to be referred semantically to a following turn.
A course that relies on open-source software for teaching introductory computer programming and Web development to psychology graduate and advanced undergraduate students is described. The rationale, content, learning goals and outcomes of the course are described, along with the specific software used. The advantages of relying on open-source solutions rather than commercial software for implementing such a course are discussed.
Natural language generation (NLG) is the task of formulating a fluent sequence of words in natural language to communicate information or ideas in applications like machine translation, human-computer dialogue, automatic summarization, and question-answering. Realization, a fundamental subtask of NLG, produces an individual sentence from a sentence plan specified in terms of linguistic relations between words and/or concepts. It involves determining the order of words, inserting function words like determiners and prepositions, performing morphological inflections, and ensuring grammaticality and agreement. An ultimate goal for natural language generation is to develop a large-scale, robust, general-purpose system. Two primary challenges are scaling up to broad coverage of syntax and producing high quality output. The irregularity of natural language makes it difficult to know how to combine linguistic primitives into fluent sentences. Also, the knowledge resources for making such a determination are time-consuming and labor-intensive to assemble, leading to a knowledge acquisition bottleneck. Evaluating whether a realizer performed appropriately is an additional challenge. There can often be more than one acceptable output, and no tools exist that can automatically assess grammaticality or fluency. This thesis takes an approach of using probabilistic models learned from text corpora to rank candidate sentences and output the most likely. It contributes (1) a symbolic mapping rule formalism and ruleset for mapping inputs to candidate outputs that achieves broad coverage through greater regularity, (2) a packed forest representation and efficient ranking algorithm that can manage the combinatorial growth in output candidates, and (3) an empirical evaluation of coverage, correctness, and the ability to handle underspecification. This evaluation is the first large-scale empirical evaluation of coverage and quality ever performed for sentence realization. The empirical evaluation is performed by automatically converting a set of 2400 hand-parsed sentences from the Penn Treebank corpus into system inputs, and then regenerating them using the system. The top-ranked output of the generator is compared to the original sentence. The results show better than 80% coverage of newspaper text and 94% precision (57% are exact matches) for almost fully-specified inputs, and the same coverage with 55% precision for minimally specified inputs.
This work deals with models used, or usable in the domain of Automatic Natural Language Processing, when one seeks a syntactic interpretation of a statement. This interpretation can be used as additional information for subsequent treatments, that can aim for instance at producing a semantic representation of the statement. It can also be used as a filter to select utterances belonging to a specific language, among several hypotheses, as done in Automatic Speech Recognition. As the syntactic interpretation of a statement is generally ambiguous with natural languages, the probabilisation of the space of syntactic trees can help in the analysis task: when several analyses are competing, one can then extract the most probable interpretation, or classify interpretations according to their probabilities. We are interested here in the probabilistic versions of Context-Free Grammars (PCFGs) and Substitution Tree Grammar (PTSGs). Syntactic treebanks, which as much as possible account for the language we wish to model, serve as the basis for defining the probabilistic parameters of such grammars. First, we exhibit in this thesis some drawbacks of the usual learning paradigms, due to the use of arbitrary heuristics (STSG DOP model), or to the use of learning criteria that consider these grammars as generative ones (creation of sentences from the grammar) rather than dedicated to analysis (creation of analyses from the sentence). In a second time, we propose new methods for training grammars, based on the traditional Maximum Entropy and Maximum Likelihood criteria. These criteria are instanciated so that they correspond to a syntactic analysis task rather than a language generation task. Specific training algorithms are necessary for their implementation, but traditional algorithms can cope with those models for the task of syntactic analysis. Lastly, we invest the problem of time complexity of syntactic analysis, which is a real issue for the effective use of PTSGs. We describe classes of PTSGs that allow the analysis of a sentence in polynomial complexity. We finally describe a method that enable the extraction of such a PTSG from the set of subtrees of a treebank. The PTSG produced by this method allows us to test our non-generative learning criterium on "realistic" data, and to give a statistical comparison between this criterium and the usual heuristic criterium in term of analysis performance.