Papers reviewed and determined not to be word norm studies. Use the flag icon to report errors or suggest re-inclusion.
18265 papers
This article presents an online dictionary environment, with enhanced sorting and searching functionalities and a text to speech feature, for hearing the pronunciation of the words. The online dictionary environment has been developed as part of the ‘Syntychies’ research program. ‘Syntychies’ online environment is a pioneering webservice for Greek dialectal lexicography and it is the first of its kind for Cypriot Greek.
This paper describes a method to convert existing treebanks with syntactic information into banks of meaning representations. The central component is a system of evaluation for a small formal language with respect to an information state. Inputs to the evaluation system are formal language expressions obtained from the conversion of parsed representations conforming to (Penn Treebank Project) guidelines. Outputs from the evaluation system are Davidsonian (higher-order) predicate logic meaning representations. Having a system of evaluation as the basis for generating meaning representations makes possible accepting input with minimal conversion from existing treebanks and from the tools used to construct treebanks. Results of having built corresponding banks of meaning representations from available treebanks are discussed.
Language resources are essential for linguistic research and the development of NLP applications. Low-density languages, such as Irish, therefore lack significant research in this area. This paper describes the early stages in the development of new language resources for Irish – namely the first Irish dependency treebank and the first Irish statistical dependency parser. We present the methodology behind building our new treebank and the steps we take to leverage upon the few existing resources. We discuss language-specific choices made when defining our dependency labelling scheme, and describe interesting Irish language characteristics such as prepositional attachment, copula and clefting. We manually develop a small treebank of 300 sentences based on an existing POS-tagged corpus and report an inter-annotator agreement of 0.7902. We train MaltParser to achieve preliminary parsing results for Irish and describe a bootstrapping approach for further stages of development.
We also show ensemble dependency parsing and self training approaches applicable to under-resourced languages using our manually annotated dependency structures. We show that for an under-resourced language, the use of tuning data for a meta classifier is more effective than using it as additional training data for individual parsers. This meta-classifier creates an ensemble dependency parser and increases the dependency accuracy by 4.92% on average and 1.99% over the best individual models on average. As the data sizes grow for the the under-resourced language a meta classifier can easily adapt. To the best of our knowledge this is the first full implementation of a dependency parser for Indonesian. Using self-training in combination with our Ensemble SVM Parser we show additional improvement. Using this parsing model we plan on expanding the size of the corpus by using a semi-supervised approach by applying the parser and correcting the errors, reducing the amount of annotation time needed.
The annotation of large corpora is usually restricted to syntactic structure and word class. Pure lexical information and information on the structure of words are stored in specialized dictionaries (Baayen et al., 1995). Both data structures ‐ dictionary and text corpus ‐ can be matched to get e.g. a distribution of certain (restricted) lexical information from a text. This procedure works fine for synchronic corpora. What is missing, however, is either a special mark-up in texts linking each of the items to a certain time or a diachronic lexical database that allows for the matching of the items over time. In what follows, we take the latter approach and present a tool set (MoreXtractor, Morphilizer, MorQuery), a database (Morphilo-DB) and the architecture of a platform (Morphorm) for a sustainable use of diachronic linguistic data for Middle English, Early Modern English and Modern English.
Mood is an important aspect of music and knowledge of mood can be used as a basic feature in music recommender and retrieval systems. A listening experiment was carried out establishing ratings for various moods and a number of attributes, e.g., valence and arousal. The analysis of these data covers the issues of the number of basic dimensions in music mood, their relation to valence and arousal, the distribution of moods in the valence-arousal plane, distinctiveness of the labels, and appropriate (number of) labels for full coverage of the plane. It is also shown that subject-averaged valence and arousal ratings can be predicted from music features by a linear model.
The focus of this article is on the creation of a collection of sentences manually annotated with respect to their sentence structure. We show that the concept of linear segments—linguistically motivated units, which may be easily detected automatically—serves as a good basis for the identification of clauses in Czech. The segment annotation captures such relationships as subordination, coordination, apposition and parenthesis; based on segmentation charts, individual clauses forming a complex sentence are identified. The annotation of a sentence structure enriches a dependency-based framework with explicit syntactic informa- tion on relations among complex units like clauses. We have gathered a collection of 3,444 sentences from the Prague Dependency Treebank, which were annotated with respect to their sentence structure (these sentences comprise 10,746 segments forming 6,341 clauses). The main purpose of the project is to gain a development data—promising results for Czech NLP tools (as a dependency parser or a machine translation system for related languages) that adopt an idea of clause segmentation have been already reported. The collection of sentences with annotated sentence structure provides the possibility of further improvement of such tools.
Despite of rapid progress in Southern Africa in the direction of multifunctionality of lexical databases through the advent of generic lexicographic software, a considerable number of lexicographic projects — especially in Khoe and Saan languages — still use or have recently used a word processor with the sole objective of compiling a printed dictionary. Hence the present paper expounds on the case of the Khoekhoegowab Dictionary Project, how in the early 1990s some off-the-shelf DOS-based database software was configured as part of a "home-grown" custom-made dictionary writing system. It is demonstrated in a non-technical way that the use of a structured database with fully-fledged retrieval facilities allows for the far-reaching elimination of human error in a dictionary, for the automatisation of processes like language reversal and sorting, and, finally, for the significantly enhanced usability of the data for purposes other than fixed media dictionary compilation. Compiling a dictionary without extensive query facilities as offered by tabular databases, is argued to be a lost opportunity, as it should be possible to utilise lexicographic data for more than just lexicography. By 2010 the data was accommodated in open source software to ensure its optimal survival in digital form for future use. Keywords: automatisation; compilation software; data retrieval; database configuration; database report; flat-file database; form; information generation; khoekhoe; khoesaan dictionaries; lexicography; lookup facilities; multifunctionality; query facilities; retrieval facilities; software; tones
We present the Prague Dependency Treebank 2.5, the newest version of PDT and the first to be released under a free license. We show the benefits of PDT 2.5 in comparison to other state-of-the-art treebanks. We present the new features of the 2.5 release, how they were obtained and how reliably they are annotated. We also show how they can be used in queries and how they are visualised with tools released alongside the treebank.
In this paper we describe semi-automatical extending of the Czech WordNet lexical database (48,000 literals in 28,000 synsets) by translation of English literals from existing synsets in Princeton WordNet. We make use of a machine-readable bilingual dictionary to extract English-Czech translation pairs, search the English literals in Princeton WordNet and in case of a high-confidence match we transfer the literal into Czech WordNet. Along with literals, new synsets parallel to the English ones and identified by ILI are introduced into CzechWordNet, including information on their ILR (Internal Language Relations) such as hypernymy/hyponymy. The paper describes the parsing of the dictionary data, extraction of translation pairs and the criteria used for estimating the confidence level of a match. Results of the work are 36,228 added literals and 12,403 created synsets. An overview of previous similar attempts for other languages is also included.
Annotated corpora such as treebanks are important for the development of parsers, language applications as well as understanding of the\nlanguage itself. Only very few languages possess these scarce resources. In this paper, we describe our efforts in syntactically annotating\na small corpora (600 sentences) of Tamil language. Our annotation is similar to Prague Dependency Treebank (PDT) and consists of\nannotation at 2 levels or layers: (i) morphological layer (m-layer) and (ii) analytical layer (a-layer). For both the layers, we introduce\nannotation schemes i.e. positional tagging for m-layer and dependency relations for a-layers. Finally, we discuss some of the issues in\ntreebank development for Tamil.
We propose HamleDT – HArmonized Multi-LanguagE Dependency Treebank. HamleDT is a compilation of existing dependency treebanks (or dependency conversions of other treebanks), transformed so that they all conform to the same annotation style. While the license terms prevent us from directly redistributing the corpora, most of them are easily acquirable for research purposes. What we provide instead is the software that normalizes tree structures in the data obtained by the user from their original providers. Keywords:dependency treebank, annotation scheme, harmonization 1.
Learning vocabulary and understanding texts present difficulty for language learners due to, among other things, the high degree of lexical ambiguity. By developing an intelligent tutoring system, this dissertation examines whether automatically providing enriched sense-specific information is effective for vocabulary learning and reading comprehension of second language learners. The system developed in this study contributes to an extended understanding of how NLP techniques can be applied more effectively in an educational environment. The system allows learners to upload texts and click on any content word in order to obtain sense-appropriate lexical information for unfamiliar or unknown words during reading. The system consists of three components: (1) the system manager controls the interaction among each learner, the NLP server, and the lexical database; (2) the NLP server converts a raw input text to a linguistically-analyzed text; (3) the lexical database is used to provide a sense-appropriate definition and example sentences of a word to the learner. To obtain the sense-appropriate information, the system first performs word sense disambiguation (WSD) on the input text. Pointing to appropriate examples tuned for language learners, however, is complicated by the fact that the database of examples is from one repository (COBUILD), while automatic WSD systems generally rely on senses from another (WordNet). The lexical database, then, is indexed by WordNet senses, each of which points to an appropriate corresponding COBUILD sense. The fact that every sense inventory has its own standards of sense distinction poses a serious problem in integrating these inventories into one. To redirect an input WordNet sense to a corresponding COBUILD sense, thus, a word sense alignment algorithm was developed, following a heuristic of favoring flatter alignment structures. With this system, an empirical study was conducted with 60 intermediate learners of English as a second language to examine whether this system can lead learners to improve their vocabulary acquisition and reading comprehension. The findings show that learners demonstrated higher performance when receiving sense-specific information. Furthermore, the qualitative examination of the effect of automatic system errors show that, although learners showed learning regardless of the appropriateness of lexical information, they still showed relatively greater learning when given appropriate lexical information.
In the past few years, much attention has been paid on extending phrase-based statistical machine translation with syntactic structures. In this paper we introduce a novel syntax encapsulated phrase(SEP) model, in which treebank tag sequences are employed to decorate the bilingual phrase pairs. We use tag sequences, instead of phrase pairs, to train the lexicalized reordering model. Since the number of treebank tags is much smaller than the number of words, the tag sequence based reordering model is smaller and more accurate than the phrase based reordering model. Experiments were carried out on four types of models: the phrase model, the hierarchical phrase model, the POS tag encapsulated phrase(PTEP) model and the syntactic tag encapsulated phrase(STEP) model. The STEP model obtained higher BLEU-4 score than other models on NIST 2005 MT task.
Syntactic analysis is a valuable addition to corpus annotation, since it enables the retrieval of structural information which is otherwise difficult to access. The Norwegian Infrastructure for the Exploration of Syntax and Semantics is adding syntactic information to the Norwegian Newspaper Corpus and is effectively producing a treebank as a parsed corpus.
The degree of translation adequacy and full-value depends on its compliance with the existing general linguistic norms. Vocabulary potential of a translator is determined by the proficiency of language to translate into. Key moments in course of transferring means of another language text are those three main features: context, word-collocations, the knowledge of ethnic specifications. Meanings of words and sentences and even whole abstracts are not autonomous, and depend on the general distributions and surroundings.
Starting from the definition of treebanks and considering that treebanks are theory dependent, we propose an annotation scheme for Romanian using several approaches ranging from phrase structure to dependency grammars and property grammars. The annotation has its starting point in a generative grammar study of the Romanian AP and validates the data of the linguistic study using an annotation scheme consisting of a constraint based approach.
Annotation of discourse relations is a project related to the Prague Dependency Treebank 2.5. It represents a new manually annotated layer of language description, above the existing layers of the PDT, and it portrays linguistic phenomena from the perspective of discourse structure and coherence.
Texts The Prague Czech-English Dependency Treebank 2.0 (PCEDT 2.0) is a major update of the Prague Czech-English Dependency Treebank 1.0 (LDC2004T25). It is a manually parsed Czech-English parallel corpus sized over 1.2 million running words in almost 50,000 sentences for each part. Data The English part contains the entire Penn Treebank - Wall Street Journal Section (LDC99T42). The Czech part consists of Czech translations of all of the Penn Treebank-WSJ texts. The corpus is 1:1 sentence-aligned. An additional automatic alignment on the node level (different for each annotation layer) is part of this release, too. The original Penn Treebank-like file structure (25 sections, each containing up to one hundred files) has been preserved. Only those PTB documents which have both POS and structural annotation (total of 2312 documents) have been translated to Czech and made part of this release. Each language part is enhanced with a comprehensive manual linguistic annotation in the PDT 2.0 style (LDC2006T01, Prague Dependency Treebank 2.0). The main features of this annotation style are: dependency structure of the content words and coordinating and similar structures (function words are attached as their attribute values) semantic labeling of content words and types of coordinating structures argument structure, including an argument structure ("valency") lexicon for both languages ellipsis and anaphora resolution. This annotation style is called tectogrammatical annotation and it constitutes the tectogrammatical layer in the corpus. For more details see below and documentation. Annotation of the Czech part Sentences of the Czech translation were automatically morphologically annotated and parsed into surface-syntax dependency trees in the PDT 2.0 annotation style. This annotation style is sometimes called analytical annotation; it constitutes the analytical layer of the corpus. The manual tectogrammatical (deep-syntax) annotation was built as a separate layer above the automatic analytical (surface-syntax) parse. A sample of 2,000 sentences was manually annotated on the analytical layer. Annotation of the English part The resulting manual tectogrammatical annotation was built above an automatic transformation of the original phrase-structure annotation of the Penn Treebank into surface dependency (analytical) representations, using the following additional linguistic information from other sources: PropBank (LDC2004T14) VerbNet NomBank (LDC2008T23) flat noun phrase structures (by courtesy of D. Vadas and J.R. Curran) For each sentence, the original Penn Treebank phrase structure trees are preserved in this corpus together with their links to the analytical and tectogrammatical annotation.
This article focuses on the relationship between a lexical database and a derivational map, a hierarchical representation of the lexicon that specifies lexical and morphological inheritance by means of graph theory. In order to take steps towards construing a three-dimensional lexicon, this article also puts forward the concept of semantic pole. A semantic pole is a pivot of lexical organization defined as the area of lexical space comprised of the intersection of the lexical areas of one or more derivational paradigms and the major exponents of a semantic prime. This proposal is applied to the semantic pole sōð-trēowe in Old English and two main conclusions are reached. Firstly, a semantic pole constitutes a panchronic representation of lexical relations that contributes to the development of the third-generation Internet, which aims, among other things, at compiling databases and representing contents in 3D. Secondly, the concept of semantic pole constitutes an explanatory principle of derivational morphology and lexical semantics because it explains the degree of convergence between morphological and lexical inheritance, accounts for the clustering of lexical items around certain semantic poles and predicts the rise of polysemy.
This paper presents an ongoing project whose goal is to create a freely available dependency treebank for Persian. The data is taken from the Bijankhan corpus, which is already annotated for parts of speech, and a syntactic dependency annotation based on the Stanford Typed Dependencies is added through a bootstrapping procedure involving the open-source dependency parser MaltParser. We report preliminary parsing experiments with promising results after training the parser on a manually annotated seed data set of 215 sentences.
The present study focuses on the characteristics of parental child-directed communication and its relationship with child language development. For this purpose, thirty-six toddlers (18 males and 18 females) and their parents were observed in a laboratory during triadic free play at ages 1; 3 and 1; 9. The characteristics of the maternal and paternal child-directed language (characteristics of communicative functions and lexicon as reported in psycholinguistic norms for Italian language) were coded during free play. Child language development was assessed during free play and at ages 2; 6 and 3; 0 using the Italian version of the MacArthur-Bates Communicative Development Inventory (2; 6) and the revised Peabody Picture Vocabulary Test (PPVT-R) (3; 0). Data analysis indicated differences between mothers and fathers in the quantitative characteristics of communicative functions and language, such as the mean length of utterances (MLU), and the number of tokens and types. Mothers also produced the more frequent nouns in the child lexicon. There emerged a relation between the characteristics of parental child-directed language and child language development.
Existing logic-based querying tools for dependency treebanks use first order logic or monadic second order logic. We introduce a very fast model checker based on hybrid logic with operators ↓, @ and A and show that it is much faster than an existing querying tool for dependency treebanks based on first order logic, and much faster than an existing general purpose hybrid logic model checker. The querying tool is made publicly available.
We address the issue of consuming heterogeneous annotation data for Chinese word segmentation and part-of-speech tagging. We empirically analyze the diversity between two representative corpora, i.e. Penn Chinese Treebank (CTB) and PKU’s People’s Daily (PPD), on manually mapped data, and show that their linguistic annotations are systematically different and highly compatible. The analysis is further exploited to improve processing accuracy by (1) integrating systems that are respectively trained on heterogeneous annotations to reduce the approximation error, and (2) re-training models with high quality automatically converted data to reduce the estimation error. Evaluation on the CTB and PPD data shows that our novel model achieves a relative error reduction of 11 % over the best reported result in the literature. 1
We present a system for cross-lingual parse disambiguation, exploiting the assumption that the meaning of a sentence remains unchanged during translation and the fact that different languages have different ambiguities. We simultaneously reduce ambiguity in multiple languages in a fully automatic way. Evaluation shows that the system reliably discards dispreferred parses from the raw parser output, which results in a pre-selection that can speed up manual treebanking. 1
A morphological analyser only recognizes words that it already knows in the lexical database. It needs, however, a way of sensing significant changes in the language in the form of newly borrowed or coined words with high frequency. We develop a finite-state morphological guesser in a pipelined methodology for extracting unknown words, lemmatizing them, and giving them a priority weight for inclusion in a lexicon. The processing is performed on a large contemporary corpus of 1,089,111,204 words and passed through a machine-learning-based annotation tool. Our method is tested on a manually-annotated gold standard of 1,310 forms and yields good results despite the complexity of the task. Our work shows the usability of a highly non-deterministic finite state guesser in a practical and complex application. 1
State-of-the-art dependency representations such as the Stanford Typed Dependencies may represent the grammatical relations in a sentence as directed, possibly cyclic graphs. Querying a syntactically annotated corpus for grammatical structures that are represented as graphs requires graph matching, which is a non-trivial task. In this paper, we present an algorithm for graph matching that is tailored to the properties of large, syntactically annotated corpora. The implementation of the algorithm is built on top of the popular IMS Open Corpus Workbench, allowing corpus linguists to re-use existing infrastructure. An evaluation of the resulting software, CWB-treebank, shows that its performance in real world applications, such as a web query interface, compares favourably to implementations that rely on a relational database or a dedicated graph database while at the same time offering a greater expres-sive power for queries. An intuitive graphical interface for building the query graphs is available via the Treebank.info project.
This paper presents the work of the Hong Kong Polytechnic University (PolyUCOMP) team which has participated in the Semantic Textual Similarity task of SemEval-2012. The PolyUCOMP system combines semantic vectors with skip bigrams to determine sentence similarity. The semantic vector is used to compute similarities between sentence pairs using the lexical database WordNet and the Wikipedia corpus. The use of skip bigram is to introduce the order of words in measuring sentence similarity. 1
This paper describes the support for mouth activity annotation provided by the iLex annotation workbench on a holistic level connected to the lexical database, on a feature level, as well as in the context of semi-automatic annotation.
Parallel treebanks have received increasing attention in the past few years,\nprimarily due to their potential use in statistical machine translation. Creating\nparallel treebanks manually is a time-consuming and expensive task\nand for this reason there is considerable interest in creating treebanks automatically.\nThis task can be solved using standard tools such as parsers and\naligners. However, because parallel treebanks are based on parallel corpora,\nwe are in a special situation where the same meaning is represented\nin two different ways. This thesis is about how we can exploit this information\nto create better parallel treebanks than we can by using standard\ntools....
Treebanking a large corpus of relatively structured speech transcribed from various Arabic Broadcast News (BN) sources has allowed us to begin to address the many challenges of annotating and parsing a speech corpus in Arabic. The now completed Arabic Treebank BN corpus consists of 432,976 source tokens (517,080 tree tokens) in 120 files of manually transcribed news broadcasts. Because news broadcasts are predominantly scripted, most of the transcribed speech is in Modern Standard Arabic (MSA). As such, the lexical and syntactic structures are very similar to the MSA in written newswire data. However, because this is spoken news, cross-linguistic speech effects such as restarts, fillers, hesitations, and repetitions are common. There is also a certain amount of dialect data present in the BN corpus, from on-the-street interviews and similar informal contexts. In this paper, we describe the finished corpus and focus on some of the necessary additions to our annotation guidelines, along with some of the technical challenges of a treebanked speech corpus and an initial parsing evaluation for this data. This corpus will be available to the community in 2012 as an LDC publication.
Continuous self-reported emotion expressed by four pieces of music were collected on a two-dimensional (valence and arousal) emotion space in a repeated measures (test-retest conditions) design. Initial orientation time (IOT), test-retest reliability and afterglow were examined. Median IOT was 8 seconds. Valence ratings took up to 25 (median 4), and for arousal up to 35 (median 12) seconds. Slower tempi seemed to require longer IOT. Test-retest reliability examined correlation coefficients, and compared periods of sample-by-sample good agreement in response between Test and Retest condition. About 80% of responses were reliable in both the Test and Retest conditions regardless of response dimension. Pearson correlations demonstrated better test-retest reliability for arousal responses than for valence. Retest condition ratings were within 8% of Test condition rating within participant. Average standard deviations for ratings collapsed across dimension, stimulus and conditions was 12.2% of the ratings scale range. Afterglow effects – large outliers in spread of scores just after the end of a piece – were identified. The reliability of continuous emotional response is therefore considered to be quite good, but caution must be taken as to how to deal with the opening and ending of continuous emotional response data.
While humans are capable of mentally transcending the here and now, this faculty for mental time travel (MTT) is dependent upon an underlying cognitive representation of time. To this end, linguistic, cognitive and behavioral evidence has revealed that people understand temporal constructs by mapping them to concrete spatial domains (e.g. past = backward, future = forward). However, very little research has investigated factors that may determine the topographical characteristics of these spatiotemporal maps. Guided by the imperative role of episodic content for retrospective and prospective thought (i.e., MTT), here we explored the possibility that the spatialization of time is influenced by the amount of episodic detail a temporal unit contains. In two experiments, participants mapped temporal events along mediolateral (Experiment 1) and anterioposterior (Experiment 2) spatial planes. Importantly, the temporal units varied in self-relevance as they pertained to temporally proximal o)
LOLspeak is a complex and systematic reimagining of the English language. It is most often associated with the popular, productive and long-lasting Internet meme ‘LOLcats’. This style of English is characterised by the simultaneous playful manipulation of multiple levels of language. Using community-generated web content as a corpus, we analyse some of the common language play strategies (Sherzer 2002) used in LOLspeak, which include morphological reanalysis, atypical sentence structure and lexical playfulness. The linguistic variety that emerges from these manipulations displays collaboratively constructed norms and tendencies providing a standard which may be meaningfully adhered to or subverted by users. We conclude with a discussion of why people may choose to participate in such language play, and suggest that the language play strategies used by participants allow for the construction of complex identity.
Background: Extraction of linguistically relevant auditory features is critical for speech comprehension in complex auditory environments, in which the relationships between acoustic stimuli are often abstract and constant while the stimuli per se are varying. These relationships are referred to as the abstract auditory rule in speech and have been investigated for their underlying neural mechanisms at an attentive stage. However, the issue of whether or not there is a sensory intelligence that enables one to automatically encode abstract auditory rules in speech at a preattentive stage has not yet been thoroughly addressed. Methodology/Principal Findings: We chose Chinese lexical tones for the current study because they help to define word meaning and hence facilitate the fabrication of an abstract auditory rule in a speech sound stream. We continuously presented native Chinese speakers with Chinese vowels differing in formant, intensity, and level of pitch to construct a complex and)
Context-free grammars are fundamental for the description of linguistic syntax. However, most artificial grammar learning experiments have explored learning of simpler finite-state grammars, while studies exploring context-free grammars have not assessed awareness and implicitness. This paper explores the implicit learning of context-free grammars employing features of hierarchical organization, recursive embedding and long-distance dependencies. The grammars also featured the distinction between left- and right-branching structures, as well as between centre- and tail-embedding, both distinctions found in natural languages. People acquired unconscious knowledge of relations between grammatical classes even for dependencies over long distances, in ways that went beyond learning simpler relations (e.g. n-grams) between individual words. The structural distinctions drawn from linguistics also proved important as performance was greater for tail-embedding than centre-embedding structures)
Principal component analysis identifies uncorrelated components from correlated variables, and a few of these uncorrelated components usually account for most of the information in the input variables. Researchers interpret each component as a separate entity representing a latent trait or profile in a population. However, the components are guaranteed to be independent and uncorrelated only when the multivariate normality of the variables is assumed. If the normality assumption does not hold, components are guaranteed to be uncorrelated, but not independent. If the independence assumption is violated, each component cannot be uniquely interpreted because of contamination by other components. Therefore, in the present study, we introduced independent component analysis, whose components are uncorrelated and independent even when the multivariate normality assumption is violated, and each component carries unique information.
This thesis investigates the changes in whether compound nouns were closed (written as one word), open (written as separate words) or hyphenated in Early New High German between 1550 and 1710. Due to the fact that there were no orthographic norms in the German of this time, graphematic phenomena in this period of the German language are very fruitful to examine. The study is based on a corpus of 249 sermons in 90 different postils. Since this thesis aims to show a diachronic development, the corpus texts originate from six time windows centred around the years 1550, 1570, 1600, 1620, 1660 and 1710. The results of the study show a general development from 1550, when around 80% of the occurrences of compound nouns were written as one word, to 1620, when this way of writing dominated almost entirely. In the texts from the last two time windows, the hyphenation spreads, and by 1710, nearly two thirds of the instances of compound nouns were written with a hyphen. The present study also shows that the geographical origin of a text is of lesser importance for the writing of compound nouns as one word, separate words or with a hyphen. However, the distinction between genuine compound nouns (a compound noun with the modifier in an unmarked case) and artificial ones (a compound noun with the modifier in an oblique case) seems to be of greater relevance. The artificial compound nouns are closed to a lesser extent in the period between 1550 and 1620 and hyphenated to a higher extent from 1660 onwards than the genuine compound nouns. In a second part of the study, the compound nouns of the different time windows are examined from a lexical point of view, showing that many compound noun lexemes were almost consistently written in the same way (either as one word, as separate words, or with a hyphen) in all occurrences within each time window.