The Orthographic Junctions of English A Reproducible Corpus-Wide Analysis of Morpheme-Boundary Statistics and ConsonantVowel Information Asymmetry (p. 1) Boicho Dimitrov Temelakiev Saxon Ventura Research Ltd 28th of May, 2026 CC BY Abstract This paper reports a reproducible, corpus-wide statistical analysis of English word structure derived entirely from a single public word list of 455,246 entries, computed in a spreadsheet with no specialized tooling (p. 1). While the distinct roles of consonants and vowels in language processing are well-recognized psycholinguistically (p. 5), and the dual-stratum organization of English morphophonology is established theoretically (p. 6), this study provides an original, datadriven quantification of these properties directly at the orthographic level. The morpheme boundary—the orthographic junction between a stem and an affix—is treated as the primary object of measurement, reading the distribution of boundary characters across the corpus (p. 1). Three results are established: 1. 2. 3. The junction carries a stable, structured filter: a consonant backbone (T, L, N, R, S, I) admitted by nearly all suffixes, an absolute floor (J, Q) admitted by none, and a distributional sparsity that scales inversely with an affix’s productivity (pp. 1-2). Affix relationships are structural: The relationship between any two affixes is quantified by the correlation of their boundary distributions, measuring their shared stem population (\(r \approx 0.99\) for etymological doublets down to \(r = 0.57\) for productivity-asymmetric near-twins) (pp. 1, 4). A massive information asymmetry partitions the lexicon: The written word decomposes into a invariant consonant skeleton carrying lexical identity (53.0% unique recoverability) and a mobile vowel tissue carrying grammatical form (1.8% unique recoverability) (pp. 1, 5). Multiple independent measures—consonant recoverability, derivational class-marking, and freestem fraction—converge on a single partition separating a transparent Germanic core from a bound Latinate superstructure (pp. 1, 6). The method, its corrections, and its limits are reported in full (p. 1). 1. Introduction and Method Traditional models of English morphophonology have long recognized that the lexicon is organized into distinct, historical strata—principally a native Germanic core and a bound Latinate superstructure (pp. 1, 6). Classic frameworks in generative phonology and lexical morphology, such as those pioneered by Chomsky and Halle (1968) and expanded by Kiparsky (1982), demonstrate that affixes of differing origins impose strict constraints on the phonetic and structural traits of the stems they recruit. Concurrently, cognitive and psycholinguistic research 1 has established a foundational "consonant-vowel functional asymmetry," demonstrating that human language processing systematically relies on consonants to preserve lexical and lexical-root identity, while vowels are dynamically manipulated to signal grammatical operations (Nespor et al., 2003; Bonatti et al., 2005). While these qualitative boundaries and cognitive patterns are deeply documented, this paper presents a mean-free, data-driven methodology to extract, quantify, and map these structural phenomena directly from corporate-scale English orthography without relying on heavy linguistic machinery. We introduce The Orthographic Junctions of English (OJE), an empirical approach that frames the morpheme boundary—the exact character interface between a stem and an affix—as an informational filter whose statistical properties reveal the historical, structural, and cognitive divisions of the vocabulary. All results derive from one corpus analysed by one elementary procedure, and the reproducibility of that procedure is treated as part of the contribution (p. 1). The corpus utilized is the opensource dwyl/english-words repository (words_alpha.txt), comprising 455,246 alphabetic entries with a total of 4,254,354 letter occurrences and a mean word length of 9.345 letters (p. 1). Each letter is assigned its ordinal value (\(A=1\) through \(Z=26\)) (p. 1). Words bearing a given suffix are isolated by end-anchored matching and aligned on their final letter, so that each suffix position returns its exact ordinal value as a safety check (p. 1); the first stem letter preceding the suffix—the linker—is then read as a full A–Z frequency distribution rather than as a mean (p. 1). The governing methodological constraint is that distributions are read in full and never collapsed to a mean prematurely, that no numerical coincidence is treated as a finding until tested across many cases, and that every claim is backed by a precise empirical count (p. 1). By avoiding any dependency on complex machine-learning libraries or external lexical databases, the framework ensures that every architectural pattern discovered can be verified using standard data operations. 2. The Orthographic Junction and Its Backbone 2 The initial phase of this investigation examined unconditioned letter bigrams across the corpus, which yielded no meaningful morphological signal. Structural regularities appeared only when character distributions were explicitly conditioned on a single morpheme boundary—the character interface linking a stem to an affix. This structural conditioning serves as the foundation of the OJE framework. A preliminary tabulation of adjacent letter pairs across the corpus—measuring which letters follow which, without regard to structural position within the word—yielded baseline frequency patterns. These patterns are entirely reducible to general English orthographic constraints and carry no isolable morphological content. The structural signal emerged only when a specific suffix was fixed and the preceding characters were read as a discrete population. Conditioning on the junction, rather than measuring adjacency as a flat sequence, renders the underlying boundary constraints visible. The set of characters that legally occupy the stem side of a morpheme boundary proves narrow, highly structured, and remarkably stable across suffixes. Reading the linker distribution across the mapped suffix inventory reveals a highly stratified, three-tier architectural filter: A structural backbone of six letters—T, L, N, R, S, and I—is admitted at high frequency by nearly every suffix in the English lexicon. Within this backbone, T serves as the single most frequent linker across the inventory and recurs as the dominant boundary letter across independent suffixes. Conversely, an absolute floor of two letters—J and Q—is admitted by no suffix at a measurable frequency. This absolute prohibition is confirmed corpus-wide and is statistically consistent with their status as the two rarest letters in English orthography overall (with J accounting for 0.18% and Q for 0.19% of all letter occurrences). Between the backbone and the floor lies a selective middle whose specific character composition varies dynamically by affix, providing the distinct orthographic footprint wherein an individual suffix’s identity resides...