Papers reviewed and determined not to be word norm studies. Use the flag icon to report errors or suggest re-inclusion.
16504 papers
It is clear that the conventions which govern the use of ‘rude’ language in public discourse have altered over a generation. It is possible to imagine a Tory patriarch like Ted Heath or a Labour leader like Harold Wilson referring to members of their governing cabinet as ‘bastards’ in private conversation. But it is difficult to imagine either of these British Prime Ministers using this description in public. Perhaps more importantly, it is open to doubt whether the use of such a term, even if it slipped out by mistake, would have been reported by the leading newspapers and media outlets of the day. It is more likely that the desire on the part of the gatekeepers of culture to protect the linguistic propriety of the political field would have outweighed the temptation to print a controversial story.1 Yet John Major's position as Prime Minister in 1993 appeared wholly unaffected by his leaked admission that he didn't sack rebel ministers after a parliamentary vote of confidence because he didn't want ‘three more of the bastards’ conspiring against him. Indeed the comment may well have enhanced Major's weak image in the eyes of the electorate. His successor, Tony Blair, at least in the early days of his premiership, actively cultivated an association with the ‘bad-mouthed’ boys and girls of Cool Britannia and his Press Secretary and confidante Alastair Campbell gained a reputation for his use of expletives in his dealings with the media.2 There seems then to have been a modification to the ‘structure of feeling’ associated with this aspect of rudeness in British society. But there is one place in Britain which has been almost automatically linked with forms of rudeness which are socially unacceptable; a location where offensiveness, crudity, insulting behaviour and nastiness constitute not so much the exception as the norm. Or at least this is how it appears in the social imaginary. The aim of this chapter will be to explore this arena in order to determine what it reveals about both British society and its boundaries of rudeness. The site to be considered is the Premiership football ground, when Saturday comes.3
У статті проаналізовано лексичні та граматичні особливості українських\nділових документів, зокрема грамот різних типів ХІV–ХV століть. Описанотематичні групи лексики, розглянуто слова й словотвірні варіанти, які за\nзначенням і формою відрізняються від сучасних. Простежено процеси\nформування у студентів-філологів історичної пам’яті, позитивного\nмнемонічного простору, відтворення ментальної історії України під час\nопрацювання текстів документів.The article was devoted to analysis of lexical and grammatical\nfeatures of Ukrainian business documents, including letters of different types of the\nXIV and XV centuries. It is established that the texts of letters designed in the style\nbusiness using the vocabulary of home and work life. Described the thematic group\nof lexicon words and considered word formation variants that differ in meaning and\nform of the modern ones. We characterized the local dialect features charters, which\nare especially spelling Ukrainian literary language of this period, denotes the value of\nlanguage literacy in the development and establishment of norms of the Ukrainian\nlanguage. In addition, traces the process of the formation of students-philologists of\nhistorical memory, the positive mnemonic space, reproduce the mental history of\nUkraine during word processing documents.
La medecine actuelle se base sur les principes de l'EBM dans l'optique destandardiser les pratiques. Les soins palliatifs sont eux aussi soumis a unereglementation et il apparait alors une norme dans les soins palliatifs. Or les soinspalliatifs sont une discipline a part qui a pour objet principal la prise en chargeindividualisee. A partir d'une revue de la litterature, les normativites theorique etpratique ont ete definies. L'objectif de ce travail est de voir quelle est la conceptiondes soignants du travail en soins palliatifs et de voir comment normativite et ideauxs'articulent au quotidien dans le travail en USP. Pour repondre a cette problematique, une etude qualitative par entretiens semi-diriges de soignants exercant en USP a ete realisee. Ils ont ete soumis a une double analyse: une analyse par le logiciel d'analyse lexical Alceste© et une analyse manuelle. Le travail en equipe, le dialogue, l'accompagnement ressortent comme des points importants du travail en USP. Il existe egalement une grande adaptabilite au patient meme s'il existe un certain rythme dans l'organisation de la journee. La norme est donc necessaire en soins palliatifs pour creer un cadre et assurer un acces egal a tous et garantir l'individualisation des prises en charge. Cette norme est en constante evolution pour s'adapter au patient
The article focuses on the moral and ethical code of the Abkhaz people, Apsuara, and on the meaning that the Abkhazians invest in such concepts as alamys, anamys and apatu as core components of the Apsuara, and how valuable these concepts are for them. The Apsuara has no lexical equivalent in Russian language. The term is difficult to describe, to define and to analyze, and its literal meaning — «abhazstvo» — is very conventional and fi gurative. The article analyzes the main components of the Apsuara structure: worldview of Abkhazians, norms of social behaviour, rules of life, a system of upbringing and etiquette, and prohibiting categories. It compares the Apsuara with the Circassians` ethical system, the Adygage (“dejstvo”). Since 1960 Abkhaz scholars such as Sh D. and Inal-Ipa, EK Adzhindzhal and others have been trying to interpret the concept of Apsuara and its system by providing their own interpretations. This article analyzes several interpretations of the concepts Apsuara.
One of the most important but easily overlooked aspects of expression of a source language into a target language is the writing norms of the target language. Every language has its own unique way of writing. Because translation involves two distinct languages and cultures, the interference of one of the forms, whether at the lexical or syntactic level, can be considered as inherent in this process, and thus unavoidable. The mediation of the translator is therefore essential in reducing the distance between the author of the source text and the reader of the final translated text. Given the existence of hegemony and a sort of power relation between cultures, this is especially true for translation from a “minor” language like Korean into a “major” language like French. Problems arise when the literary style of the source language conflicts with the writing norms of the target language. This paper seeks to find ways to cope with this problem by analyzing the French translation of Korean literary works.
In emotional speech research, it has been suggested that loudness, along with other prosodic features, may be an important cue in communicating high activation affects. In earlier studies, we found different voice quality stimuli to be consistently associated with certain affective states. In these stimuli, as in typical human productions, the different voice qualities entailed differences in loudness. To examine the extent to which the loudness differences among these voice qualities might influence the affective coloring they impart, two experiments were conducted with the synthesized stimuli, in which loudness was systematically manipulated. Experiment 1 used stimuli with distinct voice quality features including intrinsic loudness variations and stimuli where voice quality (modal voice) was kept constant, but loudness was modified to match the non-modal qualities. If loudness is the principal determinant in affect cueing for different voice qualities, there should be little or no difference in the responses to the two sets of stimuli. In Experiment 2, the stimuli included distinct voice quality features but all had equal loudness to test the hypothesis that equalizing the perceived loudness of different voice quality stimuli will have relatively little impact on affective ratings. The results suggest that loudness variation on its own is relatively ineffective whereas variation in voice quality is essential to the expression of affect. In Experiment 1, stimuli incorporating distinct voice quality features consistently obtained higher ratings than the modal voice stimuli with varied loudness. In Experiment 2, non-modal voice quality stimuli proved potent in affect cueing even with loudness differences equalized. Although loudness per se does not seem to be the major determinant of perceived affect, it can contribute positively to affect cueing: when combined with a tense or modal voice quality, increased loudness can enhance signaling of high activation states.
Turkish is an agglutinative language with rich morphology-syntax interactions. As an extension of this property, the Turkish Treebank is designed to represent sublexical dependencies, which brings extra challenges to parsing raw text. In this work, we use a joint POS tagging and parsing approach to parse Turkish raw text, and we show it outperforms a pipeline approach. Then we experiment with incorporating morphological feature prediction into the joint system. Our results show statistically significant improvements with the joint systems and achieve the state-ofthe-art accuracy for Turkish dependency parsing.
This paper investigates the impact on French dependency parsing of lexical generalization methods beyond lemmatization and morphological analysis. A distributional thesaurus is created from a large text corpus and used for distributional clustering and WordNet automatic sense ranking. The standard approach for lexical generalization in parsing is to map a word to a single generalized class, either replacing the word with the class or adding a new feature for the class. We use a richer framework that allows for probabilistic generalization, with a word represented as a probability distribution over a space of generalized classes: lemmas, clusters, or synsets. Probabilistic lexical information is introduced into parser feature vectors by modifying the weights of lexical features. We obtain improvements in parsing accuracy with some lexical generalization configurations in experiments run on the French Treebank and two out-of-domain treebanks, with slightly better performance for the probabilistic lexical generalization approach compared to the standard single-mapping approach. 1
Part-of-speech (POS) tagging is a fundamental task in natural language processing (NLP). It provides useful information for many other NLP tasks, including word sense disambiguation, text chunking, named entity recognition, syntactic parsing, semantic role labeling, and semantic parsing. In this paper, we present a new method for Vietnamese POS tagging using dual decomposition. We show how dual decomposition can be used to integrate a word-based model and a syllable-based model to yield a more powerful model for tagging Vietnamese sentences. We also describe experiments on the Viet Treebank corpus, a large annotated corpus for Vietnamese POS tagging. Experimental results show that our model using dual decomposition outperforms both word-based and syllable-based models.
Four different patterns of biased ratings of facial expressions of emotions have been found in socially anxious participants: higher negative ratings of (1) negative, (2) neutral, and (3) positive facial expressions than nonanxious controls. As a fourth pattern, some studies have found no group differences in ratings of facial expressions of emotion. However, these studies usually employed valence and arousal ratings that arguably may be less able to reflect processing of social information. We examined the relationship between social anxiety and face ratings for perceived trustworthiness given that trustworthiness is an inherently socially relevant construct. Improving on earlier analytical strategies, we evaluated the four previously found result patterns using a Bayesian approach. Ninety-eight undergraduates rated 198 face stimuli on perceived trustworthiness. Subsequently, participants completed social anxiety questionnaires to assess the severity of social fears. Bayesian modeling indicated that the probability that social anxiety did not influence judgments of trustworthiness had at least three times more empirical support in our sample than assuming any kind of negative interpretation bias in social anxiety. We concluded that the deviant interpretation of facial trustworthiness is not a relevant aspect in social anxiety.
El objetivo del trabajo consiste en reutilizar el Treebank de dependencias EPEC-DEP (BDT) para construir el gold standard de la sintaxis superficial del euskera. El paso basico consiste en el estudio comparativo de los dos formalismos aplicados sobre el mismo corpus: el formalismo de la Gramatica de Restricciones (Constraint Grammar, CG) y la Gramatica de Dependencias (Dependency Grammar, DP). Como resultado de dicho estudio hemos establecido los criterios linguisticos necesarios para derivar la funciones sintacticas en estilo CG. Dichos criterios han sido implementados y evaluados, asi en el 75% de los casos somos capaces de derivar automaticamente las funciones sintacticas para construir el gold standard.
This article considers dictionaries as lexical information / knowledge sources to be derived from a deeper, underlying, lexical database. These dictionary-tokens or -instantiations are inter alia specified by the users' needs. As a case in point of such a derivation meeting the needs of a multilingual society, a bidirectional bilingual learner dictionary is presented. Specific tools, such as editors with reversal function, and models, such as the hub-and-spoke model, are discussed as means to function within the lexicographical infrastructure of a multilingual society.
Online content analysis employs algorithmic methods to identify entities in unstructured text. Both machine learning and knowledge-base approaches lie at the foundation of contemporary named entities extraction systems. However, the progress in deploying these approaches on web-scale has been been hampered by the computational cost of NLP over massive text corpora. We present SpeedRead (SR), a named entity recognition pipeline that runs at least 10 times faster than Stanford NLP pipeline. This pipeline consists of a high performance Penn Treebank- compliant tokenizer, close to state-of-art part-of-speech (POS) tagger and knowledge-based named entity recognizer.
Empty elements (EEs) play a critical role in Chinese syntactic, semantic and discourse analysis. Previous studies employ a language-independent sentence-level approach to EE recovery, by casting it as a linear tagging or structured parsing problem. In comparison, this paper proposes a clause-level hybrid approach to address specific problems in Chinese EE recovery, which recovers EEs in Chinese language from the clause perspective and integrates the advantages of both linear tagging and structured parsing. In particular, a comma disambiguation method is employed to improve syntactic parsing and help determine clauses in Chinese. In this way, the noise introduced by sentence-level syntactic parsing and multiple EEs in the same position of a linear sentence can be well addressed. Evaluation on Chinese Treebank 6.0 shows the significant performance improvement of our clause-level hybrid approach over the state-of-the-art sentence-level baselines, and its great impact on a state-of-the-art Chinese syntactic parser.
Participants viewed dynamic facial expressions that moved from a neutral expression to varying degrees of angry, happy, or sad or from these emotionally expressive faces to neutral.A contrast effect was observed for expressions that moved to a neutral state. That is, a neutral expression that began as angry was rated as having a mildly positive expression, whereas the same neutral expression was rated as negatively valenced when it began with a smile. In Experiment 2, static expressions presented sequentially elicited contrast effects, but they were weaker than those following dynamic expressions. Experiment 3 assessed a broad range of facial movements across varying degrees of angry and happy expressions. We observed momentum effects for movements that ended at mildly expressive points (25% and 50% expressive). For such movements, affect ratings were higher, as if the perceived expression moved beyond their endpoint. Experiment 4 assessed sad facial expressions and found both contrast and momentum effects for dynamic expressions to and from sad faces. These findings demonstrate new and potent contextual influences on dynamic facial expressions and highlight the importance of facial movements in social-emotional communication.
ABSTRACT. In this article, we describe our research on wide-coverage semantics for Frenchlanguage texts and on its application to produce detailed semantic descriptions of itineraries. Using a categorial grammar semi-automatically extracted from the French Treebank and a manually constructed semantic lexicon, the resulting parser computes discourse representation structures representing the meaning of arbitrary text. The main goal of this paper is to apply and specialize this general framework of wide-coverage semantics to the spatial and temporal organization of the Itipy corpus — a set of 19th century texts discussing voyages through the Pyrenees mountains. The implemented system gives satisfying results and opens the door to the integration with specialized extensions, such as a separate module computing the discourse relations between the textual units. RÉSUMÉ. Dans cet article, nous donnons une description de notre recherche sur la sémantique à large couverture pour le français et, plus précisément, de la façon dont ces expressions sémantiques sont utilisées pour donner l’interprétation spatio-temporelle des itinéraires. En utilisant une grammaire catégorielle extraite semi-automatiquement du French Treebank et un lexique sémantique construit manuellement, l’analyseur convertit les analyses syntaxiques en DRS (discourse representation structures). Le but principal de cet article est l’application et la spécialisation de cette méthodologie générale au corpus Itipy contenant des récits de voyage du 19e siècle. La chaîne de traitement complète donne des résultats tout à fait satisfaisants et laisse l’opportunité d’y ajouter des extensions spécialisées, comme un composant qui calculerait les relations discursives entre les parties du texte.
We propose a new variant of Tree-Adjoining Grammar that allows adjunction of full wrapping trees but still bears only context-free expressivity. We provide a transformation to context-free form, and a further reduction in probabilistic model size through factorization and pooling of parameters. This collapsed context-free form is used to implement efficient grammar estimation and parsing algorithms. We perform parsing experiments the Penn Treebank and draw comparisons to Tree-Substitution Grammars and between different variations in probabilistic model design. Examination of the most probable derivations reveals examples of the linguistically relevant structure that our variant makes possible. 1
Nous présenterons les différentes couches d'annotation du treebank Rhapsodie, un corpus de français parlé richement annoté. Le corpus contient plusieurs niveaux de segmentation indépendants: en unités illocutoires pour la macrosyntaxe, en unités rectionnelles pour la microsyntaxe, en périodes, paquets intonatifs et groupes accentuels pour la prosodie. Les unités rectionnelles sont analysées en dépendance, avec un traitement fin des phénomènes d'entassements (coordination, reformulation, négo...
Large-scale unlabeled data contains abundant lexical information for NLP tasks such as Chinese word segmentation and POS tagging.This work extracted high-dimensional distributional lexical information from a largescale unlabeled Chinese corpus.An auto-encoder then performed the unsupervised dimension reduction.The learned low-dimensional lexicon features were used as new lexical features for a joint Chinese word segmentation and POS tagging task.Experiments on the Chinese Treebank 5corpus showed that the additional lexicon features improve the performance and are better than those features learned by using the principal component analysis and the k-means algorithm.
A grammar-driven dependency parsing has been attempted for Bangla (Bengali). The free-word order nature of the language makes the development of an accurate parser very difficult. The Paninian grammatical model has been used to tackle the free-word order problem. The approach is to simplify complex and compound sentences and then to parse simple sentences by satisfying the Karaka demands of the Demand Groups (Verb Groups). Finally, parsed structures are rejoined with appropriate links and Karaka labels. The parser has been trained with a Treebank of 1000 annotated sentences and then evaluated with un-annotated test data of 150 sentences. The evaluation shows that the proposed approach achieves 90.32% and 79.81% accuracies for unlabeled and labeled attachments, respectively.
This paper describes SUC-CORE, a subset of the Stockholm Ume°a Corpus and the Swedish Treebank annotated with noun phrase coreference. While most coreference annotated corpora consist of exts of similar types within related domains, SUC-CORE consists of both informative and imaginative prose and covers a wide range of literary genres and domains. This allows for exploration of coreference cross different text types, but it also means that there are limited amounts of data within each type. Future work on coreference resolution for Swedish should include making more annotated data vailable for the research community.
In this paper, with the help of corpus which was changed by the Harbin Institute of Technology Information Retrieval Laboratory's Chinese dependency-based Treebank and by using the parser MaltParser as an implementation tool, this paper proposed and implementation a Chinese dependency parsing algorithm, in order to effectively identify the right boundary of the prepositional phrases that contains a verb. Empirical results show that this algorithm has improved the error when analyze the prepositional phrase containing a verb.
Depressive symptomatology is associated with impaired recognition of emotion. Previous investigations have predominantly focused on emotion recognition of static facial expressions neglecting the influence of social interaction and critical contextual factors. In the current study, we investigated how youth and maternal symptoms of depression may be associated with emotion recognition biases during familial interactions across distinct contextual settings. Further, we explored if an individual's current emotional state may account for youth and maternal emotion recognition biases. Mother-adolescent dyads (N = 128) completed measures of depressive symptomatology and participated in three family interactions, each designed to elicit distinct emotions. Mothers and youth completed state affect ratings pertaining to self and other at the conclusion of each interaction task. Using multiple regression, depressive symptoms in both mothers and adolescents were associated with biased recognition of both positive affect (i.e., happy, excited) and negative affect (i.e., sadness, anger, frustration); however, this bias emerged primarily in contexts with a less strong emotional signal. Using actor-partner interdependence models, results suggested that youth's own state affect accounted for depression-related biases in their recognition of maternal affect. State affect did not function similarly in explaining depression-related biases for maternal recognition of adolescent emotion. Together these findings suggest a similar negative bias in emotion recognition associated with depressive symptoms in both adolescents and mothers in real-life situations, albeit potentially driven by different mechanisms.
Compared to well-resourced languages such as English and Dutch, NLP tools for linguistic analysis in Afrikaans are still not abundant. In order to facilitate corpus-based linguistic research for Afrikaans, we are creating a treebank based on the Taalkommissie corpus. We adapted a tokenizer and a shallow parser, while using a TnT tagger to do part-of-speech annotation. A first linguistic phenomenon we are investigating is the occurrence of infinitivus pro participio (IPP) in Afrikaans. IPP refers to constructions with a perfect auxiliary, in which an infinitive appears instead of the expected past participle. The phenomenon has been studied extensively in Dutch and German, but studies on Afrikaans IPP triggers are sparse. In contrast to the former two languages, it is often mentioned in the literature that in Afrikaans, IPP occurs optionally. We want to check this statement doing a corpus analysis.
It has been observed that the inclusion of morphosyntactic information in dependency treebanks is crucial to obtain high results in dependency parsing for some languages. In this paper we explore in depth to what extent it is useful to include morphological features, and the impact of diverse morphosyntactic annotations on statistical dependency parsing of Spanish. For this, we give a detailed analysis of the results of over 80 experiments performed with MaltParser through the application of MaltOptimizer. Our goal is to isolate configurations of morphosyntactic features which would allow for optimizing the parsing of Spanish texts, and to evaluate the impact that each feature has, independently and in combination with others. 1
Using neural networks to estimate the probabilities of word sequences has shown significant promise for statistical language modeling. Typical modeling methods include multi-layer neural networks, log-bilinear networks and recurrent neural networks, etc. In this paper, we propose the temporal kernel neural network language model, a variant of models mentioned above. This model explicitly captures long-term dependencies of words with exponential kernel, where the memory of history is decayed exponentially. Additionally, several sentences with variable lengths as a mini-batch are efficiently implemented for speeding up. Experimental results show that the proposed model is very competitive to the recurrent neural network language model and obtains the lower perplexity of 111.6 (more than 10% reduction) than the state-of-the-art results reported in the standard Penn Treebank Corpus. We further apply this model to Wall Street Journal speech recognition task, and observe significant improvements in word error rate.
This work is based on examples taken from tape-recordings made for a project entitled ‘Projeto da linguagem dos idosos velhos (LIV)’ (The speech of the old-old), comprising some fifty studies involving the interaction between young and old speakers in a wide variety of contextsat home, in old people’s homes and in rest homes. We have also used, exceptionally, some recordings from the ‘Projeto de estudo da norma lingüística urbana culta de São Paulo (NURC/SP)’ (Study of the educated urban linguistic norm of São Paulo, Brazil).
In this paper, we consider the problem of cross-formalism transfer in parsing. We are interested in parsing constituencybased grammars such as HPSG and CCG using a small amount of data specific for the target formalism, and a large quantity of coarse CFG annotations from the Penn Treebank. While all of the target formalisms share a similar basic syntactic structure with Penn Treebank CFG, they also encode additional constraints and semantic features. To handle this apparent discrepancy, we design a probabilistic model that jointly generates CFG and target formalism parses. The model includes features of both parses, allowing transfer between the formalisms, while preserving parsing efficiency. We evaluate our approach on three constituency-based grammars — CCG, HPSG, and LFG, augmented with the Penn Treebank-1. Our experiments show that across all three formalisms, the target parsers significantly benefit from the coarse annotations. 1 1
The Penn Discourse Treebank (PDTB) was released to the public in 2008 and remains the largest corpus of manually annotated discourse relations — both relations that are signaled explicitly (e.g., by a coordinating or subordinating conjunction, or by a discourse adverbial or other construction) and ones that otherwise appear implicit. The Penn Discourse TreeBank also diverges from other discourse-annotated corpora in permitting more than one discourse relation to be annotated as holding concurrently. Annotators could indicate this by assigning multiple sense labels to an explicit connective. Or, in those cases where adjacent sentences had no explicit connective, annotators could indicate concurrent discourse relations by either annotating a single implicit connective that concurrently conveyed multiple senses or annotating multiple implicit connectives, each conveying one of the concurrent relation(s). Subsequent experiments carried out using Mechanical Turk showed that, when a discourse adverbial explicitly signalled a discourse relation, there was often a separate concurrent relation that could be associated with an implicit coordinating or subordinating conjunction. There are different circumstances in which different sets of concurrent discourse relations are taken to hold. I will go through these, and conclude with what I take the implications of this to be for various language technologies, including statistical machine translation. Bonnie Webber. 2013. Concurrent Discourse Relations. In Proceedings of Australasian Language Technology Association Workshop, page 3.
OBJECTIVES: To intraindividually evaluate the potential of 4th generation iterative reconstruction (IR) on brain CT with regard to subjective and objective image quality. METHODS: 31 consecutive raw data sets of clinical routine native sequential brain CT scans were reconstructed with IR level 0 (= filtered back projection), 1, 3 and 4; 3 different brain filter kernels (smooth/standard/sharp) were applied respectively. Five independent radiologists with different levels of experience performed subjective image rating. Detailed ROI analysis of image contrast and noise was performed. Statistical analysis was carried out by applying a random intercept model. RESULTS: Subjective scores for the smooth and the standard kernels were best at low IR levels, but both, in particular the smooth kernel, scored inferior with an increasing IR level. The sharp kernel scored lowest at IR 0, while the scores substantially increased at high IR levels, reaching significantly best scores at IR 4. Objective measurements revealed an overall increase in contrast-to-noise ratio at higher IR levels, which was highest when applying the soft filter kernel. The absolute grey-white contrast decreased with an increasing IR level and was highest when applying the sharp filter kernel. All subjective effects were independent of the raters' experience and the patients' age and sex. CONCLUSION: Different combinations of IR level and filter kernel substantially influence subjective and objective image quality of brain CT.
In this paper, we describe our experiments in preposition disambiguation based on a – compared to a previous study – revised annotation scheme and new features derived from a matrix factorization approach as used in the field of distributional semantics. We report on the annotation and Maximum Entropy modelling of the word senses of two German prepositions, mit (‘with’) and auf (‘on’). 500 occurrences of each preposition were sampled from a treebank and annotated with syntacto-semantic classes by three annotators. Our coarse-grained classification scheme is geared towards the needs of information extraction, it relies on linguistic tests and it strives to separate semantically regular and transparent meanings from idiosyncratic meanings (i.e. of collocational constructions). We discuss our annotation scheme and the achieved inter-annotator agreement, we present descriptive statistical material e.g. on class distributions, we describe the impact of the various features on syntacto-semantic and semantic classification and focus on the contribution of semantic classes stemming from distributional semantics.
The study of acoustic ecology is concerned with the manner in which life interacts with its environment as mediated through sound. As such, a central focus is that of the soundscape: the acoustic environment as perceived by a listener. This dissertation examines the application of several computational tools in the realms of digital signal processing, multimedia information retrieval, and computer music synthesis to the analysis of the soundscape. Namely, these tools include a) an open source software library, Sirens, which can be used for the segmentation of long environmental field recordings into individual sonic events and compare these events in terms of acoustic content, b) a graph-based retrieval system that can use these measures of acoustic similarity and measures of semantic similarity using the lexical database WordNet to perform both text-based retrieval and automatic annotation of environmental sounds, and c) new techniques for the dynamic, realtime parametric morphing of multiple field recordings, informed by the geographic paths along which they were recorded.