The Junction Grammar of English A Reproducible Corpus-Wide Analysis of Morpheme-Boundary Statistics and Consonant–Vowel Information Asymmetry Boicho Dimitrov Temelakiev Saxon Ventura Research Ltd 28th of May,2026 CC BY Abstract This paper reports a reproducible, corpus-wide statistical analysis of English word structure derived entirely from a single public word list of 455,246 entries, computed in a spreadsheet with no specialized tooling. The morpheme boundary—the junction between a stem and an affix—is treated as the primary object of measurement, and the distribution of the letters that may occupy each side of a junction is read across the corpus. Three results are established. First, the junction carries a stable structure: a consonant backbone (T, L, N, R, S, I) admitted by nearly all suffixes, an absolute floor (J, Q) admitted by none, and a sparsity that scales inversely with a suffix’s productivity. Second, the relationship between any two affixes is quantified by the correlation of their boundary distributions, which measures the degree to which they share a stem population; this correlation ranges from ~0.99 for etymological doublets to 0.57 for productivity-asymmetric near-twins. Third, the written word decomposes into a consonant skeleton carrying lexical identity (53.0% of the vocabulary is uniquely recoverable from consonants alone) and a vowel tissue carrying grammatical form (1.8% recoverable from vowels alone), a ~29-fold information asymmetry. Multiple independent measures—consonant recoverability, derivational classmarking, and free-stem fraction—partition the lexicon at a single boundary separating a transparent Germanic core from a bound Latinate superstructure. The method, its corrections, and its limits are reported in full, and a program of remaining work is stated. 1. Corpus and Method All results derive from one corpus analysed by one elementary procedure, and the reproducibility of that procedure is treated as part of the contribution. The corpus is the dwyl/english-words list (words_alpha.txt), comprising 455,246 alphabetic entries with a total of 4,254,354 letter occurrences and a mean word length of 9.345 letters. Each letter is assigned its ordinal value (A=1 through Z=26). Words bearing a given suffix are isolated by end-anchored matching and aligned on their final letter, so that each suffix position returns its exact ordinal value as a sanity check; the first stem letter preceding the suffix—the linker—is then read as a full A–Z frequency distribution rather than as a mean. The governing methodological constraint is that distributions are read in full and never collapsed to a mean prematurely, that no numerical coincidence is treated as a finding until tested across many cases, and that every claim is backed by a count. Page 1 Method: a derived column applies the ordinal map; end-anchored COUNTIF and MID/CODE formulas extract suffix positions and the linker; per-letter tallies at the linker yield the boundary distribution. Nesting among suffixes (for example ‑MENT within ‑ENT, or ‑ATION within ‑TION within ‑ION) is controlled by excluding longer relatives before counting. No statistical software, machine-learning library, or external lexical database is used at any stage of the core analysis. The constraint that the analysis remain computable by elementary means is not incidental; it ensures every figure in this paper can be independently reconstructed from the public corpus with a spreadsheet alone. 2. The Junction and Its Backbone The investigation began as a survey of unconditioned letter bigrams, which returned no morphological signal; structure appeared only when letter distributions were conditioned on a single morpheme boundary, and that conditioning is the method’s foundation. A preliminary tabulation of adjacent letter pairs across the corpus—which letters follow which, without regard to position within the word—yielded frequency patterns reducible to general orthographic regularities and carrying no isolable morphological content. The signal emerged only when a specific suffix was fixed and the letters preceding it were read as a population: conditioning on the junction, rather than measuring adjacency as such, is what renders the boundary structure visible. The set of letters that may legally occupy the stem side of a morpheme boundary then proves narrow, structured, and stable across suffixes. Reading the linker distribution across the mapped suffix inventory reveals a three-tier structure. A backbone of six letters—T, L, N, R, S, and I—is admitted at high frequency by nearly every suffix; T is the single most frequent linker across the inventory and recurs as the dominant boundary letter in suffix after suffix. An absolute floor of two letters—J and Q—is admitted by no suffix at measurable frequency, a prohibition confirmed corpus-wide and consistent with their status as the two rarest letters overall (J at 0.18%, Q at 0.19% of all letter occurrences). Between backbone and floor lies a selective middle whose composition varies by suffix and in which each suffix’s identity resides. Two descriptive terms are used throughout. Because each letter carries an ordinal value (A=1 through Z=26), a linker distribution may be summarized by where its mass falls on that scale: a distribution concentrated on early-alphabet letters (low ordinal values, A through roughly M) is termed cold, and one concentrated on late-alphabet letters (high ordinal values, roughly N through Z) is termed warm. The terms refer solely to ordinal position on the A–Z scale and carry no semantic content; the backbone letters, for instance, span both ends (cold I and L against warm N, R, S, T). The degree of selectivity is itself a measurement. The count of forbidden letters at a junction—its sparsity—scales inversely with the suffix’s productivity: derivational suffixes that attach choosily to a constrained stem class forbid many letters, whereas inflectional or highly productive suffixes forbid few. Sparsity is therefore not noise but signal: the pattern of exclusion characterizes the suffix as informatively as the pattern of admission. Page 2 Method: forbidden-letter counts are read directly from the linker distribution (frequency below 0.5% taken as floor); cross-checks against COUNTIF totals confirm the populations; J/Q absence is verified against the bare ‑S population of 84,208 words and against corpus-wide letter frequencies. The junction thus behaves as a filter whose admitted and forbidden letters together encode the morphological role of the boundary, with the consonant backbone bearing the structural load and the rare letters marking its limits.