345 canon novels, one per author (166 authors), after regressing out secular prose-style drift. An unsupervised k-NN similarity graph over each novel's opening prose, clustered by community detection - no genre label is ever fed in. Colour is the community the model found; each cluster's held-out Gutenberg subject label (never used to build the graph, only to check it after) is shown alongside. Click a novel, or a genre below, to trace its cluster. Full method: design doc, README.
What this is: genres recovered purely from distinctive prose vocabulary, after controlling for three confounds in turn - corpus density, secular style drift, and prolific-author voice (one book per author). What's solid: the recovered clusters match recognizable genres and each is independently confirmed by a held-out Gutenberg subject label the model never saw. What's open: only one cluster is temporally concentrated enough to call a genuine, datable emergence (z ≤ -2); the rest are perennial modes spread across the full 1660-1928 span, and a null model (shuffled publication years) can't be distinguished from real chronology overall (z = -0.27).