Papers reviewed and determined not to be word norm studies. Use the flag icon to report errors or suggest re-inclusion.
18265 papers
The mood in a conversation is estimated by examining the speakers' affective states. Developing the automatic mood recognition system is one of the bigest issues in smooth communication research of human-human and human-computer/robot interactions. In this research, the affective rating data in UUDB (Utsunomiya University Spoken Dialogue Database) is used as the speakers' affective states. Twenty participants heard speech data and they were asked to rate the mood for each five utterances of the dialogue in the UUDB. The head-counts of participant, who considered the mood as bad/negative for the utterance blocks, are modeled by the Poisson regression modeling technique. The results demonstrate that the model is able to estimate the mood in a conversation by utilizing the speakers' affective states before the target utterance block. Thus the current study indicates that it is possible for a computer/robot to infer mood and it is possible to have effective human-to-computer/robot communications.
The article discusses the issues of teachers’ speech culture which is an important component of pedagogical culture. Mastering the linguistic norms is one of the topical questions of increasing the teachers’ pedagogical culture which are considered at the professional development courses “Russian as State Language” realized by department of language and literary education of the Chelyabinsk Institute of Retraining and Improvement of Professional Skill of Educators the framework of implementing the Federal targeted program “Russian language”. These courses are designed for various categories of educators: to specialists of educational governances of the Russian Federation; to specialists of educational governances of municipal education; to heads and deputy heads of educational institutions; to teachers of Russian language and literature; to teachers of other subjects, specialists of the preschool education system. The work aimed at developing the educators’ culture of speech should be considered as an important means contributing to the functioning of the Russian language as a state language of the Russian Federation.
Many algorithms for natural language processing rely on manual feature engineering. However, manually finding effective features is a labor-intensive task. Moreover, whenever these algorithms are applied on new types of content, they do not perform that well anymore and new features need to be engineered. For example, current algorithms developed for Part-of-Speech (PoS) tagging of news articles with Penn Treebank tags perform poorly on microposts posted on social media. As an example, the state-of-the-art Stanford tagger trained on news article data reaches an accuracy of 73% when PoS tagging microposts. When the Stanford tagger is retrained on micropost data and new micropost-specific features are added, an accuracy of 88.7% can be obtained. We show that we can achieve state-of-the-art performance for PoS tagging of Twitter microposts by solely relying on automatically inferred distributed word representations as features and a neural network. To automatically infer the distributed word representations, we make use of 400 million Twitter microposts. Next, we feed a context window of distributed word representations around the word we want to tag to a neural network to predict the corresponding PoS tag. To initialize the weights of the neural network, we pre-train it with large amounts of automatically high-confidence labeled Twitter microposts. Using a data-driven approach, we finally achieve a state-of-the-art accuracy of 88.9% when tagging Twitter microposts with Penn Treebank tags.
The various views to the sovietness and the opposite tendencies of desovietization and resovietization are observed in different spheres of modern Russian society, including in language. This article is devoted to the consideration of the (anti-) sovietness in modern Russian discourse and meta-discourse about modern Russian language. In the modern Russian language there is a clear tendency to deviate from the previous linguistic norms, and the substandard units widely intruded in the literary language. This article analyzes the desovietization shifts in the system of functional styles, language play and the parodic use of sovietisms as a precedent text. As the desovietization of modern Russian language and the violation of norms consistently increases, the protection of language purity is called for to a great extent. Under puristic approach toward Russian language inviolability of linguistic norms and ideality of language of the past are emphasized, and any language innovations are rejected. Especially the aspiration to the ideal language and non-recognition of language variation remind of the Soviet puristic ideology. Soviet puristic ideology was aimed at compliance of spontaneous oral speech with standardized written language and thereby at the unification of the mass language use.
The Contradictions of Samuel Beckett Andre Furlani (bio) Beyond New Critical Paradox “He could have shouted and could not,” begins Samuel Beckett’s first published story, “Assumption.”1 It appeared in transition in 1929 just as Ludwig Wittgenstein began an epochal reestimation of the peculiar sense such a contradiction can make. The philosopher told his students at Trinity College, Cambridge in 1933 that the purported “law” of contradiction (the rule forbidding statements of the order “p & ~p”) is really a set of norms, “which may recommend itself highly. This does not mean that we cannot use a contradiction. In fact it is used, for example, in the statement ‘I like it and don’t like it.’”2 This is not mere semantics. “If we say a thing can’t at the same time be both red and not-red, we mean that in our system we have not given this any meaning” (Wittgenstein’s Lectures, 72). Beckett continues this rehabilitation of contradiction, showing its legitimate operation in specific language games and the extent to which language games subtend even the most verifiable utterances. Wittgenstein proposes that the agreement in getting a result is the justification of a technique, be it a mathematical proof, logical proposition, or empirical statement; this agreement is not logical but grammatical.3 Contradiction in Beckett similarly need be neither an absurdist or existential device, nor an instance of aporia and infinite undecidability; instead, it can be a paradigmatic manifestation of what Wittgenstein calls the “elasticity” of linguistic norms (Wittgenstein’s Lectures, 72). While recent scholarship continues to elucidate Beckett’s philosophical sources and complements, Wittgenstein is seldom [End Page 449] addressed, despite the fact he was one of the few modern philosophers Beckett was interested in.4 Scholars have long surmised an acquaintance with the arguments of the Tractatus, as suggested by the novels Murphy and Watt, but Beckett’s relationship with Wittgenstein proves to be more intensive and prolonged.5 Beckett’s Paris library, catalogued and clarified by Mark Nixon and Dirk Van Hulle, contains a wide range of books both by and about Wittgenstein that has no equivalent among his collections of modern philosophy. These include German and English editions of the Tractatus, Lectures and Conversations on Aesthetics, Psychology, and Religious Belief, and the 1960 two-volume edition of the collected works (Schriften), published by his own German publisher Siegfried Unseld at Suhrkamp. 6 In addition to the Tractatus, it contains Philosophische Untersuchungen (Philosophical Investigations), Philosophische Bemerkungen (Philosophical Remarks), and Tagebücher 1914–1916 (Notebooks 1914–1916), an early version of the Tractatus.7 The Schriften was no mere bookshelf embellishment, for Beckett acquired a secondary literature on the philosopher. In addition to owning the supplementary Suhrkamp Beiheft, which he annotated, Ulrich Steinvorth’s edited collection(?), Über Ludwig Wittgenstein, Beckett read David Pole’s and the Suhrkamp volume Über Ludwig Wittgenstein, Beckett read David Pole’s The Later Philosophy of Wittgenstein, writing to Barbara Bray on December 21, 1962, that he was “reading Pole on Wittgenstein again.”8 He extensively annotated Bertrand Russell’s introduction to the Tractatus. He also read memoirs containing much explication of the philosopher’s earlier and later thinking, including Ludwig Wittgenstein, Personal Recollections, edited by Rush Rhees, and Paul Engelmann’s Letters from Ludwig Wittgenstein with a Memoir. A September 17, 1967, letter to Bray reports that he has received the German translation of Norman Malcolm’s Ludwig Wittgenstein: A Memoir (Samuel Beckett Papers, MS 10948/1/402), while a New Year’s Day 1971 letter thanks Mary Hutchinson for its English edition: “Wittgenstein book safely arrived. Very glad to have it” (quoted in Nixon and Van Hulle, Samuel Beckett’s Library, 167). Associates of Beckett’s also testified to his interest in Wittgenstein. John Fletcher recalled Beckett telling him that he had been reading Wittgenstein since the late 1950s, the theater technician Duncan Scott recalled a conversation with Beckett about the Tractatus in the 1970s, and André Bernold recalled Beckett telling him in 1984 that he had been reading Wittgenstein.9 The philosopher E. M. Cioran published a memoir of Beckett in 1976 that stressed his friend’s similarity to Wittgenstein, while Bray, who met Beckett when she...
In this paper, we show an approach to extracting \ndifferent types of constraint rules \nfrom a dependency treebank. Also, we \nshow an approach to integrating these constraint \nrules into a dependency data-driven \nparser, where these constraint rules inform \nparsing decisions in specific situations \nwhere a set of parsing rule (which is \ninduced from a classifier) may recommend \nseveral recommendations to the parser. \nOur experiments have shown that parsing \naccuracy could be improved by using different \nsets of constraint rules in combination \nwith a set of parsing rules. Our parser \nis based on the arc-standard algorithm of \nMaltParser but with a number of extensions, \nwhich we will discuss in some detail.
This article deals with the regularization of non-standard spellings of the verbal forms extracted from a corpus. It addresses the question of what the limits of regularization are when lemmatizing Old English weak verbs. The purpose of such regularization, also known as normalization, is to carry out lexicological analysis or lexicographical work. The analysis concentrates on weak verbs from the second class and draws on the lexical database of Old English Nerthus, which has incorporated the texts of the Dictionary of Old English Corpus. As regards the question of the limits of normalization, the solution adopted are, in the first place, that when it is necessary to regularize, normalization is restricted to correspondences based on dialectal and diachronic variation and, secondly, that normalization has to be unidirectional.
Developing a practical and accurate statistical parser for low-resourced languages is a hard problem, because it requires large-scale treebanks, which are expensive and labor-intensive to build from scratch. Unsupervised grammar induction theoretically offers a way to overcome this hurdle by learning hidden syntactic structures from raw text automatically. The accuracy of grammar induction is still impractically low because frequent collocations of non-linguistically associable units are commonly found, resulting in dependency attachment errors. We introduce a novel approach to building a statistical parser for low-resourced languages by using language parameters as a guide for grammar induction. The intuition of this paper is: most dependency attachment errors are frequently used word orders which can be captured by a small prescribed set of linguistic constraints, while the rest of the language can be learned statistically by grammar induction. We then show that covering the most frequent grammar rules via our language parameters has a strong impact on the parsing accuracy in 12 languages.
This paper is meant as a brief description of the Romanian syntax within the dependency framework, more specifically within the Universal Dependency (UD) framework, and is the result of a volunteer activity of mapping two independently created Romanian dependency treebanks to the UD specifications. This mapping process is not trivial, as concessions have to be made and solutions need to be found for various language specific phenomena. We highlight the specific characteristics of the UD relations in Romanian and argument the need for other relations. If they have already been defined for (an)other language(s) in the UD project, we adopt them.
Paninian Grammar framework provides a better solution for parsing free word order languages and Stanford Parser gives the dependencies for English language (Fixed word order language). In this paper, we map the Stanford parser dependencies to karaka relations. By using VerbNet, we capture the syntax and semantics of verb. We present the issues that encounter while doing adaptation and proposed solution to overcome these problems. We are using Hindi Dependency parser for verification of results. With this adaptation of Stanford Parser, an English-Hindi parallel treebank can be created.
BACKGROUND: As lip augmentation becomes more popular, validated measures of lip fullness for quantification of outcomes are needed. OBJECTIVE: Develop a scale for rating lip fullness and establish its reliability and sensitivity for assessing clinically meaningful differences. METHODS: The initial Allergan Lip Fullness Scale (iLFS; a four-point photographic scale with verbal descriptions) was validated by eight physicians rating 55 live subjects during two rounds, conducted on one day. In addition, subjects performed self-evaluations. The revised Allergan Lip Fullness Scale (LFS), a five-point scale with a broader range of lip presentations, was validated by 21 clinicians in two online image rating sessions, ≥14 days apart, in which they used the LFS to rate overall, upper, and lower lip fullness of 144 3-dimensional (3D) images. Physician inter- and intra-rater agreement, subject intra-rater agreement (iLFS), and subject-physician agreement (iLFS) were evaluated. Additionally, during online rating session 1, raters ranked 38 pairs of 3D images, taken before and after lip augmentation, as "clinically different" or "not clinically different." The median LFS score difference for clinically different pairs was calculated to determine the clinically meaningful difference. RESULTS: Clinician inter- and intra-rater agreement for the iLFS and LFS was substantial to almost perfect. Subject self-assessments (iLFS) had substantial intra-rater reliability and a high level of agreement with physician assessments. Median LFS score differences for overall, upper, and lower lip fullness were 1 (mean: 0.63-0.69) for "clinically different" and 0 (mean: 0.28-0.36) for "not clinically different" image pairs; thus, clinical significance of a 1-point difference in LFS score was established. CONCLUSIONS: The LFS is a reliable instrument for physician classification of lip fullness. A 1-point score difference can detect clinically meaningful differences in lip fullness.
The potential of processing user-generated texts freely available on the web is widely recognized, but due to the non-canonical nature of the language used in the web, it is not possible to process these data using conventional methodologies designed for well-edited formal texts. Procedures for properly annotating raw web data have not been as extensively researched as those for annotating well-edited texts, as also evident from the viewpoint of Turkish language processing. Moreover, there is a considerable shortage of human-annotated corpora derived from Turkish web data. The ITU Web Treebank is the first attempt for a diverse corpus compiled from Turkish texts found on the web. In this paper, we first present our survey of the non-canonical aspects of the language used in the Turkish web. Next, we discuss in detail the annotation procedure followed in the ITU Web Treebank, revised for compatibility with the language of the web. Finally, we describe the web-based annotation tool following this procedure, on which the treebank was annotated.
We propose a linguistically driven approach to represent discourse relations in Chinese text as sequences. We observe that certain surface characteristics of Chinese texts, such as the order of clauses, are overt markers of discourse structures, yet existing annotation proposals adapted from formalism constructed for English do not fully incorporate these characteristics. We present an annotated resource consisting of 325 articles in the Chinese Treebank. In addition, using this annotation, we introduce a discourse chunker based on a cascade of classifiers and report 70% top-level discourse sense accuracy.
This paper presents the IULA Spanish LSP Treebank, an open-source treebank of over 40,000 sentences, developed in the framework of the European project METANET4U. The IULA Spanish LSP Treebank is the first technical corpus of Spanish annotated at surface syntactic level, following the dependency grammar theory. We present the method we used to create the resource and the linguistic annotations that the treebank provides, using examples and comparing with similar resources. We also provide the statistics of the treebank and the evaluation results.
ABSTRACT This research encompasses the construction of a multilingual lexical database for cross-lingual information retrieval in the Indonesian legal domain. Multilingual lexical database featuring lexically and legally grounded conceptual representation can fit the cross-lingual information retrieval. Lexical database use Ontology Web Language (OWL) representation language. This representation is useful to provide application developers a high-quality resource and to promote interoperability.
Alexithymia is believed to involve deficits in emotion processing and imagery ability. Previous findings suggest that it is especially related to deficits in processing the arousal dimension of emotion, and that discordance may exist between self-report and physiological responses to emotional stimuli in alexithymia. The current study used a well-established emotional imagery paradigm to examine emotion processing deficits and discordance in participants (N = 86) selected based on their extreme scores on the Toronto Alexithymia Scale-20. Physiological (skin conductance, heart rate, and corrugator and zygomaticus electromyographic responses) and self-report (valence, arousal ratings) responses were monitored during imagery of anger, fear, joy, and neutral scenes and emotionally neutral high arousal (action) scenes. Results from regression analyses indicated that alexithymia was largely unrelated to responses on valence-based measures (facial electromyography, valence ratings), but that it was related to arousal-based measures. Specifically, alexithymia was related to higher heart rate during neutral and lower heart rate during fear imagery. Alexithymia did not predict differential responses to action versus neutral imagery, suggesting specificity of deficits to emotional contexts. Evidence for discordance between physiological responses and self-report in alexithymia was obtained from within-person analyses using multilevel modeling. Results are consistent with the idea that alexithymic deficits are specific to processing emotional arousal, and suggest difficulties with parasympathetic control and emotion regulation. Alexithymia is also associated with discordance between self-reported emotional experience and physiological response to emotion, consistent with prior evidence.
See http://clld.org
<h3>Introduction</h3><br> RST Signalling Corpus was developed at Simon Fraser University and contains annotations for signalling information added to RST Discourse Treebank (<a href="../../../LDC2002T07">LDC2002T07</a>). RST Discourse Treebank (RST-DT) is a collection of English news texts annotated for rhetorical relations under the RST (Rhetorical Structure Theory) framework. In RST Signalling Corpus, information about textual signals -- such as <em>although</em>, <em>because,</em> <em>thus</em> -- and signals such as tense, lexical chains or punctuation were added as an annotation layer to examine how rhetorical relations are signalled in discourse. <br> <h3>Data</h3><br> The source data consists of 385 Wall Street Journal news articles from the <a href="../../../LDC99T42"> Penn Treebank</a> annotated for rhetorical relations in RST Discourse Treebank. As in RST-DT, the data in this release is divided into a training set (347 articles) and a test set (38 articles). <br> The signalling annotation in this data set was performed using the <a href="http://www.wagsoft.com/CorpusTool/">UAM CorpusTool </a> version 2.8.12. Files are presented as UTF-8 encoded XML and plain text. The corpus is divided into three annotation sub-directories: training, test and full. All sub-directories include source, metadata, signalling annotation, and dtd files. <br> <h3>Samples</h3><br> Please view the following samples: <br> <ul><br> <li><a href="desc/addenda/LDC2015T10.metadata.xml">Metadata Sample</a></li><br> <li><a href="desc/addenda/LDC2015T10.signal.xml">Signal Sample</a></li><br> <li><a href="desc/addenda/LDC2015T10.txt">Text Sample</a></li><br> </ul><br> <h3>Updates</h3><br> None at this time. </br> Portions © 1987-1989 Dow Jones & Company, Inc., © 2015 Depobam Das, © 2015 Maite Taboada, © 1995, 1999, 2002, 2015 Trustees of the University of Pennsylvania
Bone et al. recently proposed an unsupervised signal-derived vocal arousal score (VC-AS) based on fusion of three intuitive acoustic features, i.e., pitch, intensity, and HF500, and have shown the effectiveness of quantifying humans perceptual rat-ings of arousal robustly across multiple corpora. Due to the readily-applicable nature of the system, this objective quantifi-cation scheme could foresee-ably be used in multiple fields of behavioral science as an objective measure of affect. In this work, we investigate in detail the relationship of this signal-derived measure to both intended arousal expression (i.e., pro-duction aspect) and perceived arousal rating (i.e., perception as-pect). On the perception side, our results in three databases (EMA, VAM, and IEMOCAP) indicate that VC-AS agrees with mean perception at least as well as an average individ-ual rater does. Regarding production, we demonstrate that in-tended arousal correlates more with VC-AS than mean percep-tion (EMA and IEMOCAP). We also show that VC-AS corre-lates more with intended arousal than perceived arousal (EMA); this finding is quite surprising given that the framework is sup-ported by extensive affective perception studies, although there is physiological motivation as well. Implication for utilizing VC-AS for novel scientific study in engineering (e.g., to miti-gate subjectivity) is further discussed. Index Terms: vocal arousal rating, affective perception, affec-tive production
We highlight the main changes recently undergone by the Italian Dependency Treebank in the transition to an extended and revised edition, compliant with the annotation schema of Universal Dependencies. We explore how these changes affect the accuracy of dependency parsers, performing comparative tests on various versions of the treebank. Despite significant changes in the annotation style, statistical parsers seem to cope well and mostly improve.
In this report we present the results obtained analysing the use, frequency of use and the position of adverbial clauses. This analysis has been performed in the Basque Dependency Treebank (BDT). We also have used the descriptive grammars of Euskaltzaindia, the Royal Academy of the Basque.
Syntactic-semantic treebank for domain ontology creationThis paper focuses on the creation of a domain treebank for the purposes of compiling a domain ontology. The domain treebank is viewed as a suitable resource for extracting of semantic relations from syntactic structures. First, the steps for ontology building are considered. Then, the processing over glossaries and standards is described with regard to their syntactic annotation. The utility of deriving semantic knowledge from the Treebank is also illustrated via the basic phrases. The idea is that the domain knowledge is represented in the domain data, but via treebanking more linguistic patterns can be extracted, which to be mapped to concepts and relations in a domain ontology.
International Affective Picture System (IAPS) is a database of photographs, which is used for studying human emotions, cognition, behavior and other areas. Although IAPS is considered suitable for using in different countries, there are some cross-cultural differences.Purpose. The aim of the present study was to determine the valence, arousal and dominance ratings of IAPS pictures in the sample of Lithuanian students and compare them with original United States (US) norms.Methods and Results. 103 Lithuania’s psychology students rated valence, arousal and dominance of 59 images from IAPS system. The results showed a high correlation between ratings of Lithuanian and US samples of all three dimensions. However, there were significant differences in the mean ratings of emotional valence and arousal – Lithuanian participants’ ratings were closer to neutral value. Moreover, some gender differences were found. Our study showed that men are more aroused by pleasant pictures compared to women, whereas an opposite tendency was observed with unpleasant pictures – women are more aroused by such images compared to men.Conclusions. The study findings suggested that IAPS can be reliably used as stimuli for studies of emotion in Lithuania.
The growth of data and information has encouraged researchers to overcome information overloaded. One of the convincing methods is through utilization of WordNet, a large lexical database of English. Generally, WordNet is used in Natural Language Processing (NLP) area and increasingly manipulated in the information retrieval field. Some search engines have using the WordNet in acquiring better meaning of user-entered keyword. Its usage might help in addressing the emergence of huge numbers of information on the web. This paper discusses the WordNet as an alternative database for semantic search engine (SSE). The SSE is then used for information retrieval. We tested our semantic search engine and compare it using Google and Yahoo search engines. Results show that our SSE gives better information retrieval when it comes towards personalization using user profiling to produce more relevant results.
The present research sought to determine how skin color, facial shape, and facial width to height ratio (fWHR) affect ratings of 10 Black male facial shapes. Based on evolutionary theory and prior research, the rectangular, quadratic, inverted trapezium, and pentagonal faces were hypothesized to receive the highest attractiveness, dominance, maturity, masculinity, strength, and social competence ratings. Additionally, faces with higher fWHRs were expected to receive higher dominance, strength, and masculinity ratings. Smaller, round or oval faces were hypothesized to receive highest warmth ratings. The results were partially consistent with these hypotheses. The examination of the effect of skin color was exploratory. Skin color did not affect ratings of the faces. These findings are discussed in terms of evolutionary adaptations and prior research.
The availability of large multi-parallel corpora offers an enormous wealth of material to contrastive corpus linguists, translators and language learners, if we can exploit the data properly. Necessary preparation steps include sentence and word alignment across multiple languages. Additionally, linguistic annotation such as part-of-speech tagging, lemmatisation, chunking, and dependency parsing facilitate precise querying of linguistic properties and can be used to extend word alignment to sub-sentential groups. Such highly inter-connected data is stored in a relational database to allow for efficient retrieval and linguistic data mining, which may include the statistics-based selection of good example sentences. The varying information needs of contrastive linguists require a flexible linguistic query language for ad hoc searches. Such queries in the format of generalised treebank query languages will be automatically translated into SQL queries.
Emotion research has benefited from standardized sets of emotional expressions. Compared to the wealth of research on the recognition of static displays of emotion, only a few studies have examined dynamic information presentation. The Amsterdam Dynamic Facial Expression Set (ADFES) is a set of emotion stimuli that contains video clips of standardized expressions of emotions (Van der Schalk et al., 2011). In the Phase 1 pilot study, psychometric properties of the ADFES were examined with a sample of 110 college students. The most accurate and reliable emotion stimuli were used in Phase 2, in which the link between blood pressure and emotion response and recognition was examined. One perspective of this link emphasizes an overall dampening of emotional responses to affect-laden stimuli (McCubbin et al., 2011), whereas another emphasizes perceiving emotions in a way that maximizes positive affect and minimizes negative affect (Jorgensen Johnson, Kolodziej, & Schreer, 1996). To compare these theoretical perspectives, Phase 2 sought to examine the relationships between resting BP and the recognition and response to dynamic facial stimuli. Analyses showed that systolic BP (β =.26; t = 2.30, p =.02) was a significant predictor of the valance rating assigned to surprise stimuli, and diastolic BP (β =.22; t = 2.02, p =.05) was a significant predictor of the arousal rating assigned to joy stimuli. Further, regression analyses revealed that diastolic BP (β = -.30; t = -2.84, p =.006) was inversely associated with the accuracy of the appraisal fear stimuli. Finally, regression analyses also revealed that systolic BP (β =.34; t = 2.92, p <.01) was a significant predictor of reaction time when viewing anger stimuli and disgust stimuli (β =.29; t = 2.59, p =.01). These results are consistent with a bias toward accentuating positive emotions while minimizing acknowledgement of negative emotions. Research directions related to response dampening versus positive response bias are discussed.
Recent years have seen a serious increase in the number and the quality of available online linguistic databases. Yet the use of databases for research purposes is still in its infancy. I intend to show how, by using different databases, one can test scientific hypotheses and how different databases can be used for different purposes. The examples will be drawn from phonology, the domain where the most comprehensive datasets are to be found. Specifically, I will examine in some details the distribution of labial-velars in Africa, since this feature has been considered typical of the 'Macro-Sudan Belt' (Clements & Rialland 2008, Guldemann 2008), hence showing an areal distribution pattern instead of a genealogical one. The following online databases have been explored: WALS (World Atlas of Language Structures), PHOIBLE, LAPSyD (Lyon-Albuquerque Phonological Systems Databases, Version 1.0.). All of them are freely available and provide maps for selected features. First I will show how these databases differ from each other: scope and quantity of data, their quality (i.e. reliability), presence vs absence of explicit curation, query interface, etc. Then the specific case of labial-velars will be explored through these databases. The geographical distribution of labial-velars is known to be restricted to an area that has been labelled ‘Macro-Sudan Belt’ (Guldemann 2008). Languages that have one or more labial-velar consonant(s) in their inventory are all situated within this area, and no language outside the area seem to exhibit any labial-velar consonant (though there probably are a handful of exceptions in the Pacific region). The use of various online databases to assess this claim yields no big surprise, if one considers the general pattern only. In the details, though, lie a few interesting things, not every one of which are captured by online databases. The most important is the local prevalence of labial-velars, i.e. their status in each language. If the actual distribution of labial-velars is the result of contact and diffusion, one would expect that their status be more marginal at the edges of the domain. The goals of this talk will therefore be: i) to present arguments to verify (or not) the above prediction; to show that even when the use of databases is not in itself enough to answer some scientific questions, they can be most helpful in suggesting directions that could have been very difficult to take without these tools.
Abstract. This paper proposes a hierarchical model to parse both En-glish and Chinese sentences. This is done by iteratively constructing simple constituents first, so that complex ones could be detected reliably with richer contextual information in the following processes. Evalua-tion on the Penn WSJ Treebank and the Penn Chinese Treebank using maximum entropy models shows that our method can achieve a good performance with more flexibility for future improvement.
The ontological approach to the creation of the learning process support systems is proposed. We research the method for automated creation of learning ontologies based on computational linguistics algorithms, the method of analysis of lexical-semantic fields corps of texts in Russian and English, frequency dictionaries of terms. Using the lexical database WordNet and terminological dictionaries the prototype ontology was developed and uploaded into the software tool "OntoMASTER-Ontology" for refining by experts. The developed method is used for creating ontologies to support the learning process of students in the field of "Information systems and technologies".
For many students, coming to learn mathematics is as much about the pedagogical relay through which concepts are conveyed as it is about the mathematics per se. This relay comprises social, cultural and linguistic norms as well as the mathematical discourse. In this study, I outline the practices of one remote school and how the teaching practices scaffold Indigenous learners whose home language is different from the language of instruction (Standard Australian English) as they come to learn mathematics. Through various strategies, teachers have created positive learning environments that celebrate the home languages of the students while supporting the transition into Standard Australian English and the discourse of mathematics.
This thesis presents open source resources in the form of annotated corpora and modules for automatic morphosyntactic processing and analysis of Persian texts. More specifically, the resources consist of an improved part-of-speech tagged corpus and a dependency treebank, as well as tools for text normalization, sentence segmentation, tokenization, part-of-speech tagging, and dependency parsing for Persian. In developing these resources and tools, two key requirements are observed: compatibility and reuse. The compatibility requirement encompasses two parts. First, the tools in the pipeline should be compatible with each other in such a way that the output of one tool is compatible with the input requirements of the next. Second, the tools should be compatible with the annotated corpora and deliver the same analysis that is found in these. The reuse requirement means that all the components in the pipeline are developed by reusing resources, standard methods, and open source state-of-the-art tools. This is necessary to make the project feasible. Given these requirements, the thesis investigates two main research questions. The first is how can we develop morphologically and syntactically annotated corpora and tools while satisfying the requirements of compatibility and reuse? The approach taken is to accept the tokenization variations in the corpora to achieve robustness. The tokenization variations in Persian texts are related to the orthographic variations of writing fixed expressions, as well as various types of affixes and clitics. Since these variations are inherent properties of Persian texts, it is important that the tools in the pipeline can handle them. Therefore, they should not be trained on idealized data. The second question concerns how accurately we can perform morphological and syntactic analysis for Persian by adapting and applying existing tools to the annotated corpora. The experimental evaluation of the tools shows that the sentence segmenter and tokenizer achieve an F-score close to 100%, the tagger has an accuracy of nearly 97.5%, and the parser achieves a best labeled accuracy of over 82% (with unlabeled accuracy close to 87%).
We are proposing an extension of the recursive neural network that makes use of a variant of the long short-term memory architecture. The extension allows information low in parse trees to be stored in a memory register (the 'memory cell') and used much later higher up in the parse tree. This provides a solution to the vanishing gradient problem and allows the network to capture long range dependencies. Experimental results show that our composition outperformed the traditional neural-network composition on the Stanford Sentiment Treebank.