The Orthographic Junctions of EnglishA Reproducible Corpus-Wide Analysis of Morpheme-Boundary Statistics and Consonant–Vowel Information AsymmetryBoicho Dimitrov Temelakiev · Saxon Ventura Research Ltd · CC BY · Version 3, June 2026Changes in this version (v3)This version corrects and strengthens the deposited record. (i) The corpus provenance is corrected: the analysed file is the dwyl `words.txt` list after cleaning, not `words_alpha.txt` as stated in v1–v2; the exact 455,246-entry corpus is deposited so the source is unambiguous. (ii) The descriptive "warm/cold" parameter (v1 §2/v2 §3) is withdrawn as a measurement-scale error: it performs arithmetic on the alphabet's ordinal position, which is a nominal code, and we show it carries no order-invariant signal. (iii) The associated "conjugation temperature" claim is restated without any appeal to alphabet position. (iv) All doublet correlations were recomputed on the deposited corpus under explicit nesting controls; the values are reported as measured, and one figure (−ANT/−ANCE) is corrected. (v) A permutation-null robustness test is added as a standing methodological control. No surviving result depends on the alphabet's ordering.AbstractThis paper reports a reproducible, corpus-wide statistical analysis of English word structure derived entirely from a single public word list of 455,246 entries, computed in a spreadsheet with no specialized tooling. The morpheme boundary — the orthographic junction between a stem and an affix — is treated as the primary object of measurement, and the distribution of the letters that may occupy the stem side of a junction is read across the corpus. Three results are established. First, the junction carries a stable, structured filter: a consonant backbone (T, L, N, R, S, I) admitted by nearly all suffixes, an absolute floor (J, Q) admitted by none, and a sparsity that scales inversely with a suffix's productivity. Second, the relationship between any two affixes is quantified by the correlation of their boundary distributions, which measures the degree to which they share a stem population; this correlation ranges from ~0.97 for etymological doublets to 0.58 for productivity-asymmetric near-twins. Third, the written word decomposes into a consonant skeleton carrying lexical identity (53.0% of the vocabulary uniquely recoverable from consonants alone) and a vowel tissue carrying grammatical form (1.8% recoverable from vowels alone), a ~29-fold information asymmetry. Multiple independent measures — consonant recoverability, derivational class-marking, and free-stem fraction — partition the lexicon at a single boundary separating a transparent Germanic core from a bound Latinate superstructure. Every quantitative claim is tested for invariance under permutation of the alphabet, so that no result depends on the arbitrary ordering of the letters; the method, its corrections, and its limits are reported in full.1. Corpus and MethodAll results derive from one corpus analysed by one elementary procedure, and the reproducibility of that procedure is treated as part of the contribution.The corpus derives from the public dwyl/english-words list (`words.txt`), from which non-alphabetic entries were removed and the remainder case-folded and deduplicated, yielding 455,246 unique alphabetic entries with 4,254,354 letter occurrences and a mean word length of 9.345 letters. Because the upstream list drifts over time, the exact 455,246-entry file analysed here is deposited with this record and is the corpus of record; a reader downloading the live upstream list today will not recover the same entry count.Words bearing a given suffix are isolated by end-anchored matching and aligned on their final letter; the first stem letter preceding the suffix — the linker — is read as a full per-letter (A–Z) frequency distribution rather than as a mean. Each letter may be referred to by its position in the alphabet purely as a label; no quantity in this paper depends on treating that position as a number (see §9). The governing methodological constraints are that distributions are read in full and never collapsed to a mean prematurely, that no numerical coincidence is treated as a finding until tested across many cases, that every claim is backed by a count, and — new in this version — that every claim is invariant under relabelling of the alphabet.Nesting among suffixes (for example −MENT within −ENT, or −ATION within −TION within −ION) is controlled by excluding longer relatives before counting. No statistical software, machine-learning library, or external lexical database is used at any stage of the core analysis; every figure can be reconstructed from the deposited corpus with a spreadsheet alone.2. The Junction and Its BackboneThe investigation began as a survey of unconditioned letter bigrams, which returned no morphological signal; structure appeared only when letter distributions were conditioned on a single morpheme boundary, and that conditioning is the method's foundation.Conditioning character statistics on a morpheme boundary places this work within the successor-variety tradition of boundary detection introduced by Harris (1955) and first implemented computationally by Hafer and Weiss (1974). That tradition uses transitional letter predictability to segment words into morphemes; the present method inverts the emphasis, holding a known boundary fixed and characterizing the distribution of stem-side letters it admits — a characterization of the junction rather than a segmentation of the word.A preliminary tabulation of adjacent letter pairs across the corpus — which letters follow which, without regard to position within the word — yielded frequency patterns reducible to general orthographic regularities and carrying no isolable morphological content. The signal emerged only when a specific suffix was fixed and the letters preceding it were read as a population. The set of letters that may legally occupy the stem side of a morpheme boundary then proves narrow, structured, and stable across suffixes.Reading the linker distribution across the mapped suffix inventory reveals a three-tier structure. A backbone of six letters — T, L, N, R, S, and I — is admitted at high frequency by nearly every suffix; T is the single most frequent linker across the inventory and recurs as the dominant boundary letter in suffix after suffix. An absolute floor of two letters — J and Q — is admitted by no suffix at measurable frequency, a prohibition confirmed corpus-wide and consistent with their status as the two rarest letters overall (J at 0.18%, Q at 0.19% of all letter occurrences). Between backbone and floor lies a selective middle whose composition varies by suffix and in which each suffix's identity resides. (The backbone, floor, and selective middle are stated as sets of letters; nothing in the three-tier description depends on the order of the alphabet, and all of it is invariant under the permutation test of §9.)The degree of selectivity is itself a measurement. The count of forbidden letters at a junction — its sparsity — scales inversely with the suffix's productivity: derivational suffixes that attach choosily to a constrained stem class forbid many letters, whereas inflectional or highly productive suffixes forbid few. Sparsity is therefore not noise but signal: the pattern of exclusion characterizes the suffix as informatively as the pattern of admission.3. The Suffix AtlasSuffixes are described by the distribution of letters that survive at their boundary — read in full, never reduced to a mean. The corrective lesson is explicit and was learned in this program: −NESS was first misjudged from a summary statistic and only described correctly once its full distribution was read (E 27%, D 15%, S 15%, I 13%). The general rule that follows is that a junction must be read as a distribution, not a single number, and in particular not as a mean of letter positions — a point developed formally in §9.Two representative distributions illustrate the contrast between a concentrated and a broad boundary: −ABLE (T-led and broad) and −IBLE (a sparse, frozen Latinate boundary). Both are reported as per-letter frequencies; the comparison between them is made by correlation (§4), which is invariant under relabelling of the letters.Table 1. Linker distribution of −ABLE (n = 4,694). Letters ≥3% shown.Linker Count PercentT 846 18.0%R 513 10.9%N 420 9.0%E 392 8.4%S 330 7.0%I 292 6.2%D 286 6.1%L 286 6.1% Table 2. Linker distribution of −IBLE (n = 738). Five letters carry ~90% of the population.Linker Count PercentS 248 33.7%T 208 28.3%C 86 11.7%D 63 8.6%G 59 8.0%N 20 2.7% Read by distribution, the suffix inventory resolves into a small set of boundary shapes, each fixed by the population of stems the suffix recruits rather than by its function, origin, or spelling.4. The Doublet Principle, QuantifiedTwo suffixes that draw on the same stem population share the same boundary distribution, and the correlation between their distributions measures the extent of that shared population directly. Because correlation is computed component-by-component over the same set of letters in both vectors, it is invariant under any relabelling of the alphabet — it is one of the order-independent quantities the program now requires (§9).All correlations below were recomputed for this version on the deposited corpus under explicit nesting controls (Pearson r over the 26-component per-letter frequency vectors). Across three classes of suffix pairs the boundary correlation forms an interpretable gradient.Etymological doublets — the same Latin stem class in two guises — correlate highly: −ENT/−ENCE at r = 0.97 (nesting-controlled, excluding −MENT from the −ENT population) and −ANT/−ANCE at r = 0.94. The −ANT/−ANCE value is corrected here: earlier versions reported r ≈ 0.99, which does not reproduc