Papers reviewed and determined not to be word norm studies. Use the flag icon to report errors or suggest re-inclusion.
16504 papers
La segmentation d'un texte en Unites Discursives Minimales (UDM) a pour but de decouper le texte en segments qui ne se chevauchent pas. Ces segments sont ensuite relies entre eux afin de construire la structure discursive d'un texte. La plupart des approches existantes utilisent une analyse syntaxique extensive. Malheureusement, certaines langues ne disposent pas d'analyseur syntaxique robuste. Dans cet article, nous etudions la faisabilite de la segmentation discursive de textes arabes en nous basant sur une approche d'apprentissage supervisee qui predit les UDM et les UDM imbriques. La performance de notre segmentation a ete evaluee sur deux genres de corpus: des textes de livres de l'enseignement secondaire et des textes du corpus Arabic Treebank. Nous montrons que la combinaison de traits typographiques, morphologiques et lexicaux permet une bonne reconnaissance des bornes de segments. De plus, nous montrons que l'ajout de traits syntaxiques n'ameliore pas les performances de notre segmentation.
This article analyses the structure of Yoruba numerals and their derivation. Data are collected from the compilation of Yoruba numerals and observation of its use coupled with the researcher's intuitive knowledge of the language. The work dwells on the existing literature on numerals too. The author adopts a descriptive method in analysing the data. The work looks at the roles of affixes in realising odd numbers, multiples of 20, centenary, bicentenary, and so on in their order of increase. It is discovered that the direction of counting in Yoruba is largely progressive. Besides, the language adopts base 5, decimal (base 10) and vigesimal (base 20) systems of counting. It is equally discovered that the choice of either of the two variations is largely dependent on the articulatory parameter of the first vowel (V1) of the root word. It is noted that the Yoruba numeral system offers a suitable linguistic database for both the theoretical and empirical domains of linguistic study especially documentary linguistics. The current study has general pedagogic implications for the teaching and learning of Yoruba numerals.
In this paper, we introduce experiment results of a Vietnamese sentence parser which is built by using the Chomsky's subcategorization theory and PDCG (Probabilistic Definite Clause Grammar). The efficiency of this subcategorized PDCG parser has been proved by experiments, in which, we have built by hand a Treebank with 1000 syntactic structures of Vietnamese training sentences, and used different testing datasets to evaluate the results. As a result, the precisions, recalls and F-measures of these experiments are over 98%.
Traditional information retrieval systems rely on keywords to index documents and queries. In such systems, documents are retrieved based on the number of shared keywords with the query. This lexicalfocused retrieval leads to inaccurate and incomplete results when different keywords are used to describe the documents and queries. Semantic-focused retrieval approaches attempt to overcome this problem by relying on concepts rather than on keywords to indexing and retrieval. The goal is to retrieve documents that are semantically relevant to a given user query. This paper addresses this issue by proposing a solution at the indexing level. More precisely, we propose a novel approach for semantic indexing based on concepts identified from a linguistic resource. In particular, our approach relies on the joint use of WordNet and WordNetDomains lexical databases for concept identification. Furthermore, we propose a semantic-based concept weighting scheme that relies on a novel definition of concept centrality. The resulting system is evaluated on the TIME test collection. Experimental results show the effectiveness of our proposition over traditional IR approaches.
We present an automatic animacy classier for Dutch that can determine the animacy status of nouns | how alive the noun’s referent is (human, inanimate, etc.). Animacy is a semantic property that has been shown to play a role in human sentence processing, felicity and grammaticality. Although animacy is not marked explicitly in Dutch, we expect knowledge about animacy to be helpful for parsing, translation and other NLP tasks. Only a few animacy classiers and animacyannotated corpora exist internationally. For Dutch, animacy information is only available in the Cornetto lexical-semantic database. We augment this lexical information with context information from the Dutch Lassy Large treebank, to create training data for an animacy classier that uses a novel kind of context features. We use the k-nearest neighbour algorithm with distributional lexical features, e.g. how frequently the noun occurs as a subject of the verb ‘to think’ in a corpus, to decide on the (predominant) animacy class. The size of the Lassy Large corpus makes this possible, and the high level of detail these word association features provide, results in accurate Dutch-language animacy classication.
Discourse structure and discourse relations are an important ingredient in systems for the analysis of text that go beyond the boundary of single clauses. Discourse relations often indicate important additional information about the connection between two clauses, such as causality, and are widely believed to have an influence on aspects of reference resolution.In this article, we first present the general design choices that are to be made in the design of an annotation scheme for discourse structure and discourse relations. In a second part, we present the scheme used in our annotation of selected articles from the TüBa-D/Z treebank of German (Telljohann et al., 2009). The scheme used in the annotation is theory-neutral, but informed by more detailed linguistic knowledge in the way of linguistic tests that can help disambiguate between several plausible relations.
Second language learners’ (L2ers’) perception and production of consonant clusters is influenced by the syllable structure of the native language (L1). This study investigates whether the perception of epenthetic vowels is partially responsible for why Spanish speakers have difficulty producing /s/ + Consonant (“sC”) clusters in English, and whether it affects word recognition in continuous speech. Spanish, German L2ers of English, and native English speakers completed: (i) an AXB task with (/ə/)sC-initial nonce words (e.g., [əsman]-[sman]); (ii) a word monitoring task with (/ə/)sC-initial words in semantically ambiguous sentences (e.g., I have lived in that (e)state for a long time); and (iii) a production task with the same sentences as in (i). L2ers also took a word-familiarity rating task and a cloze test to assess their proficiency. For (i) and (ii), accuracy rates were recorded, and response times were measured from target onset. For (iii), acoustic analyses showed whether the L2ers’ productions of sC-initial words contained an epenthetic vowel. Preliminary results suggest that perception difficulties may be partially responsible for Spanish speakers’ production and word-recognition difficulties with sC-clusters in English, but production data suggest that articulatory problems may also play an important role. Proficiency does not seem to help overcome this difficulty.
This paper presents a new method of analysis by which structural similarities between brain data and linguistic data can be assessed at the semantic level. It shows how to measure the strength of these structural similarities and so determine the relatively better fit of the brain data with one semantic model over another. The first model is derived from WordNet, a lexical database of English compiled by language experts. The second is given by the corpus-based statistical technique of latent semantic analysis (LSA), which detects relations between words that are latent or hidden in text. The brain data are drawn from experiments in which statements about the geography of Europe were presented auditorily to participants who were asked to determine their truth or falsity while electroencephalographic (EEG) recordings were made. The theoretical framework for the analysis of the brain and semantic data derives from axiomatizations of theories such as the theory of differences in utility preference. Using brain-data samples from individual trials time-locked to the presentation of each word, ordinal relations of similarity differences are computed for the brain data and for the linguistic data. In each case those relations that are invariant with respect to the brain and linguistic data, and are correlated with sufficient statistical strength, amount to structural similarities between the brain and linguistic data. Results show that many more statistically significant structural similarities can be found between the brain data and the WordNet-derived data than the LSA-derived data. The work reported here is placed within the context of other recent studies of semantics and the brain. The main contribution of this paper is the new method it presents for the study of semantics and the brain and the focus it permits on networks of relations detected in brain data and represented by a semantic model.
This paper presents a novel method using graph-based semi-supervised learning (SSL) to improve the syntax parsing of unknown words. Different from conventional approaches that uses hand-crafted rules, rich morphological features, or a character-based model to handle unknown words, this method is based on a graph-based label propagation technique. It gives greater improvement on grammars trained on a smaller amount of labeled data and a large amount of unlabeled one. A transductiv 1 graph-based SSL method is employed to propagate POS and derive the emission distributions from labeled data to unlabeled one. The derived distributions are incorporated into the parsing process. The proposed method effectively augments the original supervised parsing model by contributing 2.28 % and 1.72 % absolute improvement on the accuracy of POS tagging and syntax parsing for Penn Chinese Treebank respectively. 1
The Penn Discourse Treebank (PDTB) was released to the public in 2008 and remains the largest corpus of manually annotated discourse relations — both relations that are signaled explicitly (e.g., by a coordinating or subordinating conjunction, or by a discourse adverbial or other construction) and ones that otherwise appear implicit. The Penn Discourse TreeBank also diverges from other discourse-annotated corpora in permitting more than one discourse relation to be annotated as holding concurrently. Annotators could indicate this by assigning multiple sense labels to an explicit connective. Or, in those cases where adjacent sentences had no explicit connective, annotators could indicate concurrent discourse relations by either annotating a single implicit connective that concurrently conveyed multiple senses or annotating multiple implicit connectives, each conveying one of the concurrent relation(s). Subsequent experiments carried out using Mechanical Turk showed that, when a discourse adverbial explicitly signalled a discourse relation, there was often a separate concurrent relation that could be associated with an implicit coordinating or subordinating conjunction. There are different circumstances in which different sets of concurrent discourse relations are taken to hold. I will go through these, and conclude with what I take the implications of this to be for various language technologies, including statistical machine translation. Bonnie Webber. 2013. Concurrent Discourse Relations. In Proceedings of Australasian Language Technology Association Workshop, page 3.
OBJECTIVES: To intraindividually evaluate the potential of 4th generation iterative reconstruction (IR) on brain CT with regard to subjective and objective image quality. METHODS: 31 consecutive raw data sets of clinical routine native sequential brain CT scans were reconstructed with IR level 0 (= filtered back projection), 1, 3 and 4; 3 different brain filter kernels (smooth/standard/sharp) were applied respectively. Five independent radiologists with different levels of experience performed subjective image rating. Detailed ROI analysis of image contrast and noise was performed. Statistical analysis was carried out by applying a random intercept model. RESULTS: Subjective scores for the smooth and the standard kernels were best at low IR levels, but both, in particular the smooth kernel, scored inferior with an increasing IR level. The sharp kernel scored lowest at IR 0, while the scores substantially increased at high IR levels, reaching significantly best scores at IR 4. Objective measurements revealed an overall increase in contrast-to-noise ratio at higher IR levels, which was highest when applying the soft filter kernel. The absolute grey-white contrast decreased with an increasing IR level and was highest when applying the sharp filter kernel. All subjective effects were independent of the raters' experience and the patients' age and sex. CONCLUSION: Different combinations of IR level and filter kernel substantially influence subjective and objective image quality of brain CT.
In this paper, we describe our experiments in preposition disambiguation based on a – compared to a previous study – revised annotation scheme and new features derived from a matrix factorization approach as used in the field of distributional semantics. We report on the annotation and Maximum Entropy modelling of the word senses of two German prepositions, mit (‘with’) and auf (‘on’). 500 occurrences of each preposition were sampled from a treebank and annotated with syntacto-semantic classes by three annotators. Our coarse-grained classification scheme is geared towards the needs of information extraction, it relies on linguistic tests and it strives to separate semantically regular and transparent meanings from idiosyncratic meanings (i.e. of collocational constructions). We discuss our annotation scheme and the achieved inter-annotator agreement, we present descriptive statistical material e.g. on class distributions, we describe the impact of the various features on syntacto-semantic and semantic classification and focus on the contribution of semantic classes stemming from distributional semantics.
The study of acoustic ecology is concerned with the manner in which life interacts with its environment as mediated through sound. As such, a central focus is that of the soundscape: the acoustic environment as perceived by a listener. This dissertation examines the application of several computational tools in the realms of digital signal processing, multimedia information retrieval, and computer music synthesis to the analysis of the soundscape. Namely, these tools include a) an open source software library, Sirens, which can be used for the segmentation of long environmental field recordings into individual sonic events and compare these events in terms of acoustic content, b) a graph-based retrieval system that can use these measures of acoustic similarity and measures of semantic similarity using the lexical database WordNet to perform both text-based retrieval and automatic annotation of environmental sounds, and c) new techniques for the dynamic, realtime parametric morphing of multiple field recordings, informed by the geographic paths along which they were recorded.
This paper reports on the annotation and maximum-entropy modeling of the semantics of two German prepositions, mit (‘with’) and auf (‘on’). 500 occurrences of each preposition were sampled from a treebank and annotated with syntacto-semantic classes by two annotators. The classification is guided by a perspective of information extraction, relies on linguistic tests and aims at the separation of semantically transparent and opaque meanings (that is of collocational constructions). Apart from descriptive statistical material, we present results of experiments using monolingual and multilingual evidence (the latter from informative English and Spanish translations) in order to predict the semantic classes.
The Sensor Web is a growing phenomenon where an increasing number of sensors are collecting data in the physical world, to be made available over the Internet. To help realize the Sensor Web, the Open Geospatial Consortium (OGC) has developed open standards to standardize the communication protocols for sharing sensor data. Spatial Data Infrastructures (SDIs) are systems that have been developed to access, process, and visualize geospatial data from heterogeneous sources, and SDIs can be designed specifically for the Sensor Web. However, there are problems with interoperability associated with a lack of standardized naming, even with data collected using the same open standard. The objective of this research is to automatically group similar sensor data layers. We propose a methodology to automatically group similar sensor data layers based on the phenomenon they measure. Our methodology is based on a unique bottom-up approach that uses text processing, approximate string matching, and semantic string matching of data layers. We use WordNet as a lexical database to compute word pair similarities and derive a set-based dissimilarity function using those scores. Two approaches are taken to group data layers: mapping is defined between all the data layers, and clustering is performed to group similar data layers. We evaluate the results of our methodology.
This paper examines the morphophonological violation of the linguistic rules of Yoruba especially by the literary artists in their attempt to achieve communicative aesthetics. Through observational method, we discovered that they manipulate the morphophonological resources of the language without a second thought on its implication for communication. In the work, we discover that this deliberate deviation from the linguistic norms of the everyday language do have some consequences (i) it may lead to ambiguity (ii) it often leads to derivation of a new words which may be out of context with the discussion at stake (iii) it could be for the ease of speech production (iv) it is also noticeable in ordinary discourse as against some views that it is only manifested in the literary discourse (v) it is capable of constituting communicative difficulty to the language learners. Keywords: morphology, phonology, deviation process, derivation, communicative aesthetics, literary discourse
The present paper focuses on ways in which the pragmatic (functional) meaning that arises from various contextual features, known in corpus linguistics as semantic prosody (Sinclair 1996, 2004; Louw 1993, etc.) can become an integral part of lexicographic descriptions. This is especially important for the treatment of phraseology and idiomatics. The workings of semantic prosody are a good example of the ways pragmatic meaning exploits linguistic means to be codified in the text. We thus investigate the meaning that can only be studied in context, as it is completely dependent on collocation, i.e., syntagmatic relations, and therefore cannot be attributed solely to a concrete word form. Corpus analysis has yielded significant results in areas such as the lexicographic treatment of semantic prosody. We believe that in order to improve teaching pragmatics in all its complexity, it is necessary to recognise and assess various aspects of pragmatic meaning both in written and spoken language. Second/foreign language teaching/learning in particular has been strongly dependent on the inclusion of relevant information in dictionaries, in which, traditionally, pragmatic aspects of meaning have been largely neglected. Language technologies have enabled us both to study the subtleties of pragmatic meaning and to design accurate and more user-friendly (pedagogical) dictionaries. We will attempt to demonstrate the value of explicit description of functional pragmatic meaning, i.e. semantic prosody, as implemented in the Slovene Lexical Database (2008-2012). A brief overview of the theoretical background is first provided, after which we describe the definition strategies employed to include pragmatics, as well as presenting a case study and arguing that explicating semantic prosody is crucial in developing pragmatic competence in (young/foreign) language learners.
Bracketing induction is the unsupervised learning of hierarchical constituents without labeling their syntactic categories such as verb phrase (VP) from natural raw sentences. Constituent Context Model (CCM) is an effective generative model for the bracketing induction, but the CCM computes probability of a constituent in a very straightforward way no matter how long this constituent is. Such method causes severe data sparse problem because long constituents are more unlikely to appear in test set. To overcome the data sparse problem, this paper proposes to define a non-parametric Bayesian prior distribution, namely the Pitman-Yor Process (PYP) prior, over constituents for constituent smoothing. The PYP prior functions as a back-off smoothing method through using a hierarchical smoothing scheme (HSS). Various kinds of HSS are proposed in this paper. We find that two kinds of HSS are effective, attaining or significantly improving the state-of-the-art performance of the bracketing induction evaluated on standard treebanks of various languages, while another kind of HSS, which is commonly used for smoothing sequences by n-gram Markovization, is not effective for improving the performance of the CCM.
This paper describes the design of question answering system that participates in the maintask of QA4MRE at CLEF 2013. This system will initially perform preprocessing stage of the document and the documents related questions. Then, it identifies the type of questions in order to be able to search the answers with the most appropriate approach. In order to finding the answers, the system uses eight heuristics features and two query expansion techniques, namely Pseudo Relevance feedback and WordNet as the lexical database. There are nine types of runs were submitted at CLEF 2013, where the difference lies in the combination of query expansion techniques and features used. Evaluation will be based on the c@1 measure. Best results are obtained when it uses both query expansion techniques and all features with certain weights.
Recent research has found it useful to distinguish between the form and meaning of sounds. To investigate the relevance of meaning, naïve students and professional drivers listened to four levels of meaning neutralisation and four levels of spectral slope of recorded truck sound. Self-assessment of emotional reactions showed that professional drivers did not vary much in activation and rated over all lower activation than naïve participants whose affect ratings moved more or less along the annoyance correlation line in the upper left quadrant of the affect map. This gives some information about the importance of the source being recognisable and of previous user experience for product sound quality. It is further supported by that the overall difference between naïve participants’ and professional drivers’ ratings decreased with increasing meaning neutralisation. The methodology applied in the current study may be adopted to form homogenous panels of experts for sound evaluation.
This paper builds six dependence syntactic networks based on six treebanks of different styles and gives a comparative analysis of overall characteristics of the networks, including the number of edges, the number of the nodes, the average degree, the clustering coefficient, the average path length, the centralization, the diameter, and the index of power-law, coefficient of determination. After that, the paper uses the Euclideanthe shortest distancemethod, with characteristics as variables, to do clustering analysis of these networks. The results show that using some main parameters of networks, namely the number of the nodes, the clustering coefficient, the average path length, the centralization and the index of power-law, can do cluster analysis on texts. Compared with the traditional text clustering, the results are easier to explain in linguistic angle.
Conventional statistics-based methods for joint Chinese word segmentation and part-of-speech tagging (S&T) have generalization ability to recognize new words that do not appear in the training data. An undesirable side effect is that a number of meaningless words will be incorrectly created. We propose an effective and efficient framework for S&T that introduces features to significantly reduce meaningless words generation. A general lexicon, Wikepedia and a large-scale raw corpus of 200 billion characters are used to generate word-based features for the wordhood. The word-lattice based framework consists of a character-based model and a word-based model in order to employ our word-based features. Experiments on Penn Chinese treebank 5 show that this method has a 62.9% reduction of meaningless word generation in comparison with the baseline. As a result, the F1 measure for segmentation is increased to 0.984.
This study undertook a critical appraisal of the correlation between the intractable social conflicts like the Boko-Haram and the Niger Delta crises, where youths are the key players, on the international image of Nigeria and tourism development in the country. It is motivated by the avalanche of media reports that the country’s image is being seriously battered abroad by these internal social problems. The specific objectives sought were to: ascertain the correlation between the Boko-Haram crisis and the nation’s image ratings abroad; the Niger Delta crisis and the nation’s image ratings abroad and their impacts on tourism development in the country. Survey design was adopted in the study, where electronic questionnaires (E-questionnaire) via the Internet were used to gather the primary data. The data so sourced were statistically presented/analyzed with Likert’s 5-points scale, Spearman’s correlation coefficient and Friedman chi-square. Results obtained show that both the Boko Haram crisis and the Niger Delta crisis have adverse impact on the country’s international image and tourism development, consequently on youths’ unemployment rate. It was then recommended that proactive public relations crisis management strategies should be used in nipping such crisis in their buds in future. Keywords: Boko Haram crisis, Niger Delta crisis, National Image, Tourism Development.
The paper presents a research tool for studying semantic change and polysemy patterns in Russian adjectives and adverbs. It is based on a corpus analysis of high-freqency polysemous units. For each of them we describe the meanings it can have, assign to each meaning a corresponding taxonomic class, identify types of semantic shifts between individual meanings (metaphor, metonymy; besides, a full-scale approach reveals non-canonical cases of semantic shifts), describe context conditions of these shifts (semantic and grammatical restrictions on co-occurring words). The results gained from this analysis are implemented in a database, which allows for various generalizations on the regularities of change in adjective and adverb meaning. Several examples are given to illustrate what kinds of queries can be performed on the database. Keywords: polysemy; semantic shift; metaphor; metonymy; semantics of adjectives; Russian language; lexical database
The increasing significance of political communication in the functioning of the modern socium actualizes the problem of the communicative essence of power, forms and methods of its manifestation, the correlation of communicative and power components. The author researches the communicative specificity of political interaction, its moral-ethical principles, and gives the interpretation of the linguo-communicative code as a system of linguistic and extra-linguistic norms and rules that determine the success of modern political interaction.
Recognizing and classifying implicit discourse relations is a challenging task since hardly any strong indicators exist, and a variety of weak indicators has to be harnessed to yield evidence for a particular discourse relation or another. Most current approaches rely on a combination of shallow, surface-based features and rather specialized hand-crafted features, with a considerable gap in between which is partly due to the sheer complexity of combining evidence from different levels of linguistic description. As a way to avoid both the shallowness of word-based representations and the lack of coverage of specialized linguistic features, we use a graph-based representation of discourse segments, which allows for a more abstract (and hence generalizable) notion of syntactic (and partially of semantic) structure, we propose an approach to use a graph-structured representation of discourse units in order to improve the classification of implicit discourse relations. We validate this approach using implicit discourse relation data from the TüBa-D/Z treebank, providing an extended discussion and error analysis that looks at the impact of the graph-based representation on the different kinds of discourse relations. The empirical evaluation shows that our graph-based approach not only provides a suitable representation for the linguistic factors that are needed in disambiguating discourse relations, but also improves results over a strong state-of-the-art baseline by more accurately identifying Temporal, Comparison and (for the German data) Reporting discourse relations. 1.
An MT-oriented system using Conditional Random Fields (CRFs) is presented to identify English Prepositional Phrases (PPs) within business domain. For the purpose of English-Chinese Machine Translation (MT), we, under the guidance of the theory of Syntactic Functional Grammar (SFG), refine PP function chunks into four types instead of the binary attachment. In order to improve the identification of these chunk types, we revise the Penn Treebank tagset with four major changes being made. A small size of 998k English annotated corpus in business domain is semi-automatically built based on our new tagset employing the Maximum Entropy model. Experiments show that our system achieves an accuracy of 88.45%, higher than other reported approaches. The adjustments made in the PP chunk types and POS tagset give rise to 4.11%, 4.25% and 4.15% increase in the precision, recall and F-score respectively.
This paper addresses semantic search of Web services using natural language processing. First we survey various existing approaches, focusing on the fact that the expensive costs of current semantic annotation frameworks result in limited use of semantic search for large scale applications. We then propose a service search framework based on the vector space model to combine the traditional frequency weighted term-document matrix, the syntactical information extracted from a lexical database and a dependency grammar parser. In particular, instead of using terms as the rows in a term-document matrix, we propose using synsets from WordNet to distinguish different meanings of a word under different contexts as well as clustering different words with similar meanings. Also based on the characteristics of Web services descriptions, we propose an approach to identifying semantically important terms to adjust weightings. Our experiments show that our approach achieves its goal well.
Information management is an important requirement in today"s world. The anaphoric references hide the important information. The identification of anaphorically referred information is called as anaphora resolution which has significant impact on increasing the efficiency of information management, techniques including text summarization, information extraction etc. In this paper we have proposed a method for anaphora resolution to engineer information management. The proposed method acceptably determines the potential referents of the anaphora specially of the verb phrase form, distinguishes between pleonastic "it" and the anaphoric "it" and resolve the anaphora which is referred to after an interval of multiple sentences. The referents are stored in a list in the order of their occurrence in the discourse and eliminated from the list if they are not referred for to long. "Recency" is used as a salience factor to select the correct referent if other information like gender, number and type are not suffices to estimate the correct referent for an anaphor. To achieve a more precise resolution system WordNet lexical database is exploited to compare the synonyms of the anaphor with its possible referents.
Inspired by robust generalization and adversarial learning we describe a novel approach to learning structured perceptrons for part-ofspeech (POS) tagging that is less sensitive to domain shifts. The objective of our method is to minimize average loss under random distribution shifts. We restrict the possible target distributions to mixtures of the source distribution and random Zipfian distributions. Our algorithm is used for POS tagging and evaluated on the English Web Treebank and the Danish Dependency Treebank with an average 4.4 % error reduction in tagging accuracy. 1
SYNTAGMA is a rule-based parsing system, structured on two levels: a general parsing engine and a language specific grammar. The parsing engine is a language independent program, while grammar and language specific rules and resources are given as text files, consisting in a list of constituent structuresand a lexical database with word sense related features and constraints. Since its theoretical background is principally Tesniere's Elements de syntaxe, SYNTAGMA's grammar emphasizes the role of argument structure (valency) in constraint satisfaction, and allows also horizontal bounds, for instance treating coordination. Notions such as Pro, traces, empty categories are derived from Generative Grammar and some solutions are close to Government&Binding Theory, although they are the result of an autonomous research. These properties allow SYNTAGMA to manage complex syntactic configurations and well known weak points in parsing engineering. An important resource is the semantic network, which is used in disambiguation tasks. Parsing process follows a bottom-up, rule driven strategy. Its behavior can be controlled and fine-tuned.
The paper describes a broadly applicable method of designing multilingual semantics-syntactic analyzers of recommender systems. The user inputs may include the questions of many kinds formed with the help of interrogative words (or without interrogative words), verbs, nouns, attributes, prepositions, the designations of the digital values of various parameters. For the queries in English and German, the developed algorithm of semantic-syntactic analysis processes the questions of many kinds, the commands, and the statements from a restricted sublanguage of NL. For the queries in Russian, the algorithm is additionally able to process the requests with participle constructions and attributive clauses. As a semantic intermediary language, the algorithm uses the SK-language determined by the considered linguistic database. The class of SK-languages is introduced by the theory of K-representations (knowledge representations), its current version is mainly stated in a monograph of the author published by Springer in 2010. The developed algorithm is implemented by means of the programming language PYTHON.
Financial literacy education has traditionally never been a major subject area in most public schools in Maryland. With the unanticipated struggling of the United States of America (USA) Economy, Maryland (one of the richest states in the USA and its residents, especially of its major urban city, Baltimore, is facing severe family and personal financial crisis. This unanticipated financial crisis characterized by declining family and personal savings, mortgage defaults and foreclosures, and tenant evictions due to rent defaults is motivating educators, business entities, and politicians to develop education strategies to educate children from future financial crisis. Although there is an overwhelming consensus for providing financial education as a major curriculum to our schoolchildren, the practicability of most curriculums is obscure. To develop a practical approach to teaching financial literacy in elementary schools, an assessment using a Grid Familiarity Rating Chart (Very Familiar -Not Familiar) of basic financial and economics concepts was assigned to 210 students (4th -5th graders) in four inner city schools. Students received information to place a check mark by each concept if they are very familiar or familiar with the concept. Results revealed that 97% of the students were very familiar or familiar with the basic financial concepts as compared to 49% of the students who were very familiar or familiar with basic economic concepts. Even when students were encouraged to explain their very familiar or familiar concepts, 80% provided explanations that were accurate or almost accurate for basic financial concepts as compared to 20% accuracy for explanations with basic economic concepts. Conclusively, the most basic economic concepts (scarcity, choice, and opportunity cost), which are essential decision making tools for most financial decisions should be simplified and included in financial education lesson plans.
This experiment investigated the role of acoustic correlates of the singing voice in the perception of broad affect dimensions using the two-dimensional model of affect. The dataset consisted of vocal and glottal recordings of a sung vowel interpreted in different singing expressions. Listeners were asked to rate the sounds according to four perceived affect dimensions. A cross-tabulation was done between the singing expressions and affect judgments. A one-way ANOVA was performed for 11 acoustic cues with the affect ratings. It was found that the singing power ratio (SPR), mean intensity, brightness, mean pitch, jitter, shimmer, mean harmonic-to-noise ratio (HNR), and mean autocorrelation discriminate broad affect dimensions. Principal component analysis (PCA) was performed on the acoustic correlates. Two components were retained that explained 78.1% of the total variance of vocal cues and 73.5% of that of the glottal cues.
Automatic syntactic analysis of natural language is one of the fundamental problems in natural language processing. Dependency parses (directed trees in which edges represent the syntactic relationships between the words in a sentence) have been found to be particularly useful for machine translation, question answering, and other practical applications. For English dependency parsing, we show that models and features compatible with how conjunctions are represented in treebanks yield a parser with state-of-the-art overall accuracy and substantial improvements in the accuracy of conjunctions. For languages other than English, dependency parsing has often been formulated as either searching over trees without any crossing dependencies (projective trees) or searching over all directed spanning trees. The former sacrifices the ability to produce many natural language structures; the latter is NP-hard in the presence of features with scopes over siblings or grandparents in the tree. This thesis explores alternative ways to simultaneously produce crossing dependencies in the output and use models that parametrize over multiple edges. Gap inheritance is introduced in this thesis and quantifies the nesting of subtrees over intervals. The thesis provides O( n6) and O(n 5) edge-factored parsing algorithms for two new classes of trees based on this property, and extends the latter to include grandparent factors. This thesis then defines 1-Endpoint-Crossing trees, in which for any edge that is crossed, all other edges that cross that edge share an endpoint. This property covers 95.8% or more of dependency parses across a variety of languages. A crossing-sensitive factorization introduced in this thesis generalizes a commonly used third-order factorization (capable of scoring triples of edges simultaneously). This thesis provides exact dynamic programming algorithms that find the optimal 1-Endpoint-Crossing tree under either an edge-factored model or this crossing-sensitive third-order model in O(n 4) time, orders of magnitude faster than other mildly non-projective parsing algorithms and identical to the parsing time for projective trees under the third-order model. The implemented parser is significantly more accurate than the third-order projective parser under many experimental settings and significantly less accurate on none.
1 Kateřina Rysová Annotation The presented thesis is focused on the Czech word order of contextually non-bound verbal modifications. It monitors whether there is a basic order in the contextually non-bound part of the sentence (significantly predominant in frequency) in the surface word order (cf. narodit se v Brně v roce 1950 vs. narodit se v roce 1950 v Brně; literally to be born in Brno in 1950 vs. to be born in 1950 in Brno). At the same time, we try to find out the factors influencing the word order (such as the form of modifications, their lexical expression or the effect of verbal valency). Finally, we briefly compare the word order tendencies in Czech and German. For the verification of the objectives, mainly the data from the Prague Dependency Treebank are used. The work is based on the theoretical principles of Functional Generative Description. Research results demonstrate that, at least in some cases, it is possible to detect certain general tendencies to use preferably one of two possible surface word order sequences in Czech. Abstract The aim of the doctoral thesis is to describe particular aspects of the Czech (and partly also German) word order in the sentences coming mainly from journalistic texts. The first part examines the role of different types of verbal modifications in sentence...
Historical linguistics, among other things, aims at understanding the principles and factors that cause changes in languages. The Dravidian comparative linguistics in the last few decades has arrived at excellent results at different levels of language change: phonology, morphology and etymology. However, the field of historical syntax remains to be explored in detail. The linguistic analysis of Tamil inscriptions and classical and ancient literary texts will shed light on the historical linguistics of Tamil and will try to fill a gap in the historical linguistics of the Dravidian family of languages. An in-depth linguistic analysis of Tamil epigraphic texts will show how the Tamil language used in Tamil inscriptions constitutes an important diachronic evidence of both sociolinguistic and linguistic evolution. I will concentrate here on the following three aspects: 1) Historical sociolinguistics: Maṇipravāḷa style and the development of Tamil as Inscriptional Language, 2) Historical linguistics, Syntax and Information structure, and 3) Construction of a fine-grained linguistic database and demonstrate how ‘corpus analysis’ can help us mapping the process of language change and language use.
The dissertation is a three-paper investigation of potential methods for improving exposure therapy outcomes for anxiety disorders. If proven effective, these methods can be utilized to modify exposure therapy for anxiety disorders in order to enhance treatment effects and reduce relapse rates. Study 1 tested a prediction of the Rescorla-Wagner model: that presenting two fear-provoking stimuli simultaneously (compound extinction) would maximize learning during extinction. Participants were presented with single extinction trials only or single extinction trials followed by compound extinction trials. Additionally, participants were randomized to caffeine or placebo ingestion prior to extinction. Results indicated participants presented with compound trials demonstrated significantly less fear responding at spontaneous recovery whereas ingestion of caffeine provided limited protection (only on valence ratings). At reinstatement, only compound extinction trials predicted attenuated fear responding. Study 2 investigated whether occasional reinforced trials during extinction enhances learning, as has been demonstrated in recent literature in animal models. Participants were randomly assigned to typical exposure procedures (i.e., no CS-US pairings) or occasional CS-US pairings during extinction. Results indicated participants presented with occasional reinforced trials maintained elevated fear responding throughout exposure but demonstrated attenuated spontaneous recovery and rapid reacquisition effects at follow-up testing. Study 3 was based on recent evidence in animal models suggesting that sustained arousal and enhanced fear responding throughout extinction predicts better long-term outcomes, which is contrary to traditional exposure therapy in which reduction of fear responding is used as an index of learning. Participants completed exposure with or without the presence of additional excitatory stimuli which were intended to enhance arousal and fear responding throughout exposure. A set of regression analyses investigating whether any exposure process measures predicted outcome indicated that sustained arousal throughout exposure and variability in subjective fear responding throughout exposure predicted lower levels of fear at follow-up testing. In sum, these studies indicate that exposure therapy will be most effective when exposure sessions are unpredictable, variable, include multiple fear-provoking stimuli, and include some aversive events (e.g., social rejection, panic attack). Although fear responding may remain elevated throughout exposure, such procedures may predict better treatment outcomes and reduce the likelihood of relapse.
The topicality of material presented in the article is conditioned with the lack of research aimed at the formation of a culture of dialogue speech of students of nonphilological specialties. The author of the article has considered the dialog, voice, speech, cultural, communicative culture in forming the dialogues, the functioningof which depends on the ability to listen, hear and understand the interlocutor, which is aimed at the development ofskills to produce dialogic speech. Keywords: culture of dialogic speech, linguistic features of dialogs, vocal and linguistic norms, vocal etiquette, communication.
Lucas (3;6) is playing with Lego animals on the carpet. He takes the lions and puts them into the compound he prepared for them. He comments: “Kuck, e Léiw!” (Look, a lion!) Then he whirls them all around and says: “Tout mélanger!” (Mixing everything!) Lucas here seems not only to mix the animals but the languages, too… Linguistic diversity is not only an integral part of Luxembourgian society in general but also of the everyday practice in early childcare settings. This is nothing exceptional. Rather, the increasing diversification of languages, cultures, and identities – or the increasing acknowledgement that these concepts have never been simple and fixed – is a central characteristic of contemporary societies worldwide. Nevertheless, education political and media discourses keep on positing multilingualism as a special challenge that pedagogical practice has to cope with. They call for early language promotion, school preparation, and for the advancement of social integration and equality. While these discourses instantly turn to the programmatic question how multilingualism should be dealt with, thus presupposing a normative understanding of language in education, the present paper asks how these complex demands are actually met in everyday pedagogical practice and how, along the way, linguistic norms are practically accomplished as well as caught into question. The paper draws on ethnographic material from three Luxembourgian daycare centers that were investigated during 18 months as part of my doctoral research. The choice and interpretation of field notes is guided by the central question how linguistic diversity is dealt with in the centers’ everyday routines. The empirical exploration reveals how pedagogical practice is itself constituted within a field of tension between monolingualist agendas and the actors’ translingual practices. The three centers manage this tension differently, thus demonstrating the multiplicity of possible pathways when dealing with multilingualism in early education.
Chinese divides into simplified Chinese and traditional Chinese. Some treebank resources like Penn Chinese Treebank: CTB had been built for training simplified Chinese parser (Yu, et al. 2010) while Sinica Treebank was developed for parsing traditional Chinese (Chen et al., 1999). Limit to our knowledge, there are still not grammatical resources that analyze both simplified Chinese and traditional Chinese. A rule-based Chinese grammatical resource --Chinese Sentence Structure Grammar: CSSG had been developed based on the idea of Sentence Structure Grammar: SSG (Wang et al., 2012). We assume that a rule-based grammatical resource should analyze both simplified Chinese and traditional Chinese if there are no obvious differences between their grammatical constructions. Aiming at verifying this assumptions, we parse the test sentences from the simplified Chinese parsing task (task 3) and the traditional Chinese parsing task (task 4) of CLP 2012 with the same rule-based parser that was implemented the grammatical resource CSSG. CSSG includes two parts of resources: the grammatical rules and a simplified Chinese morphological dictionary. We transfer the simplified Chinese characters of the dictionary to traditional Chinese characters for obtaining a traditional Chinese morphological dictionary. We parse the test sentences of task 3 and task 4 with the same CSSG rules but different morphological dictionaries (simplified or traditional Chinese characters). We convert CSSG parsing trees to TCT-style trees and Sinica-style trees to participate in the evaluations of the two tasks. The experiments show that the CSSG rules can parse both simplified Chinese and traditional Chinese, but the performance of the latter is lower than the former. We noticed that a few traditional Chinese constructions are different from simplified Chinese.
Objective – To assess how the age, gender, and race characteristics of library users affect their perceptions of the approachability of reference librarians with similar or different demographic characteristics.
 
 Design – Image rating survey.
 
 Setting – Large, three-campus university system in the United States.
 
 Subjects – There were 449 students, staff, and faculty of different ages, gender, and race.
 
 Methods – In an online survey respondents were presented with images of hypothetical librarians and asked to evaluate their approachability, using a scale from 1 to 10. The images showed librarians with neutral emotional expressions against a standardized, neutral background. The librarians’ age, gender, and race were systematically varied. Only White, African American, and Asian American librarians were shown. Afterwards respondents were asked to identify their own age, gender, race, and status.
 
 Main Results – Respondents perceived female librarians as more approachable than male librarians, maybe due to expectations caused by the female librarian stereotype. They found librarians of their own age group more approachable. African American respondents scored African American librarians as more approachable, whereas Whites expressed no significant variation when rating the approachability of librarians of different races. Thus, African Americans demonstrated strong in-group bias but Whites manifested colour blindness – possibly a strategy to avoid the appearance of racial bias. Asian Americans rated African American librarians lower than White librarians.
 
 Conclusion – This study demonstrates that visible demographic characteristics matter in people’s first impressions of librarians. Findings confirm that diversity initiatives are needed in academic libraries to ensure that all users feel welcome and are encouraged to approach librarians. Regarding gender, programs that deflate the female librarian stereotype may help improve the approachability image of male librarians. Academic libraries should staff the reference desk with individuals covering a wide range of ages, including college-aged interns, whom traditional age students find most approachable. Libraries should also build a racially diverse staff to meet the needs of a racially diverse user population. Since first impressions have lasting effects on the development of social relationships, structural diversity should be a priority for libraries’ diversity programs.
Introduction The processing of nouns and verbs and their differences have been an area of intense interest among psycholinguists (see review by Vigliocco et al., 2010), for the obvious reasons that they are two major word classes across languages and convey the most basic information in communication. Word class effects are often reflected in response latency and/or accuracy in naming tasks. However, single word production does not resemble daily communication in which linguistic contexts may facilitate word finding. Previous studies directly comparing lexical retrieval between naming and narrative tasks have obtained mixed results (e.g. Berndt et al., 2002; Pashek & Tompkins, 2002), despite the fact that nouns and verbs were rarely matched for relevant psycholinguistic variables. This study minimized the influence of confounding factors and employed neuropsychological data to examine retrieval of nouns and verbs in confrontation naming and connected speech. Method The participants were 19 Cantonese-speaking adults with anomic aphasia and 19 age-, gender- and education-matched controls. Production of nouns and verbs was obtained from confrontation picture naming and narrative tasks from the Cantonese AphasiaBank database (Kong et al., 2009). At least 20 items in each condition were chosen with comparable age of acquisition and familiarity estimates; however, imageability ratings were higher in naming than narrative tasks and higher for nouns than verbs, and verbs in naming were longer than those in narrative task. Results and Discussion Significant main effects of speaker group, word class, and task, as well as a two-way interaction between task and word class were found (p < 0.01). Better performance in nouns than verbs and naming in picture than narrative tasks? was observed in normal speakers. The difference in accuracy between word classes was greater in naming than narrative tasks. A hierarchical multiple regression was also carried out to assess the effects of word class and task after the influence of imageability had been taken into consideration. Only “task” remained a significant predictor (p < 0.05). Our results have shown that when the influence of confounding factors is reduced, there is no evidence for word class specific deficits among fluent aphasic speakers (but see Matzig et al., 2009) or contextual support for word production.
The Electronic Health Record (EHR) contains information useful for clinical, epidemiological and genetic studies. This information of patient symptoms, history, medication and treatment is not completely captured in the structured part of the EHR but is often found in the form of freetext narrative. A major obstacle for clinical studies is finding patients that fit the eligibility criteria of the study. Using EHR in order to automatically identify relevant cohorts can help speed up both clinical trials and retrospective studies (Restificar, Korkontzelos et al. 2013). While the clinical criteria for inclusion and exclusion from the study are explicitly stated in most studies, automating the process using the EHR database of the hospital is often impossible as the structured part of the database (age, gender, ICD9/10 medical codes, etc.’) rarely covers all of the criteria. Many resources such as UMLS (Bodenreider 2004), cTakes (Savova, Masanz et al. 2010), MetaMap (Aronson and Lang 2010) and recently richly annotated corpora and treebanks (Albright, Lanfranchi et al. 2013) are available for processing and representing medical texts in English. Resource poor languages, however, suffer from lack in NLP tools and medical resources. Dictionaries exhaustively mapping medical terms to the UMLS medical meta-thesaurus are only available in a limited number of languages besides English. NLP annotation tools, when they exist for resource poor languages, suffer from heavy loss of accuracy when used outside the domain on which they were trained, as is well documented for English (Tsuruoka, Tateishi et al. 2005; Tateisi, Tsuruoka et al. 2006). In this work we focus on the problem of classifying patient eligibility for inclusion in retrospective study of the epidemiology of epilepsy in Southern Israel. Israel has a centralized structure of medical services which include advanced EHR systems. However, the free text sections of these EHR are written in Hebrew, a resource poor language in both NLP tools and handcrafted medical vocabularies. Epilepsy is a common chronic neurologic disorder characterized by seizures. These seizures are transient signs and/or symptoms of abnormal, excessive, or hyper synchronous neuronal activity in the brain. Epilepsy is one of the most common of the serious neurological disorders (Hirtz, Thurman et al. 2007).
Objective To explore the parental perception of school-age children's body images and its influence factors among parents of school children in Changsha city,Hunan province,and to provide reference for childhood obesity prevention. Methods A survey with anthropometric measurement(height,w eight) was conducted among 2 224 elementary school students of grade 4-6 in Changsha city in April,2012. Body image was categorized based on the WHO 2007 body mass index(BMI) reference standard. The parental perception of the body image was examined with a questionnaire among the parents of the school children and the influence factors of the perception were analyzed with logistic regression. Results The body image rating scale results show ed that 56. 1% of the parents underestimated their children' s body image size. Multiple-factors logistic regression show ed that the risk factors for underestimating the children's body image was child's big body image(odds ratio [OR]= 35. 763,95% confidence interval [95% CI]= 23. 745-53. 863), w hile the protective factor was female parent(OR = 0. 623,95% CI = 0. 400-0. 969). Residing in the countryside(OR =1. 464,95%CI =1. 090-1. 966),with the child in senior grade(OR =1. 272,95%CI =1. 067-1. 517),and with the child having big body images(OR = 32. 089,95% CI = 16. 810-61. 255) were risk factors for parental underestimation of children's body image,w hile high parental education(OR = 0. 870,95% CI = 0. 773-0. 979) and overw eight or obese of the parents(OR = 0. 578,95% CI = 0. 403-0. 830) were protective factors. Conclusion Parental perception of school children's body image prensents an underestimation trend among the parents in Changsha. We should reinforce the communication and education in the parents to promote the correct evaluation on their children's body images for the control of overw eight and obesity in the school children.
We describe in this paper how different learning strategies can be applied on the same NLP task, namely chunking. The reference corpus is extracted from the French Treebank, the symbolic learning strategy used is grammatical inference and the statistical one is CRFs (Conditional Random Fields). As expected, the symbolic approach allows readability but is less effective than the statistical one. We then propose two distinct ways to combine both approaches and show that in both cases they benefit from one another.