Literature Mutations · S7

Marks in Books

A mark in a book is not an opinion about influence. It is a physical fact — and it turns out to measure something narrower than we hoped.

Abstract

Every influence graph in digital humanities is validated against attribution — a scholar, a crowd, or a model asserting that A influenced B. That is an interpretation of evidence, made by someone who was not there. This is the evidence: 216,553 marks that 555 writers left in other authors' books, joined from eight archives into one corpus with a licence gate and an explicit evidence ladder.

Two measurements. The passages a reader marked do not resemble his own prose more than unmarked passages of the same book do — except in the one book he was demonstrably copying from, where the effect is sharp and directional. But at the author level, reading does predict style: for three of four readers, the authors they physically read sit measurably closer to them in a similarity graph than the authors they did not. Reading shows up in a writer's overall voice; the specific lines he underlined are not the route by which it happens.

1 What is on record

Structured, transcribed marginalia exists for a handful of readers, deeply — not for many readers, thinly. Eight archives, joined:

SourceVolumesMarksOwn wordsAuthors
Archaeology of Reading — Harvey, Dee36189,98028,27531
Shakespeare and Company — 253 writers' library cards7,7418,4321,420
Beckett Digital Library8695,7772,677487
Melville's Marginalia Online524,86436719
Annotated Books Online — Luther, Erasmus, Newton213,5273,50617
Whitman Archive3622,3031,384160
Derrida's Margins2221,670211107
Nietzsche — naming sources in his own prose1,081 refs60

Table 1. “Own words” counts marks where the reader wrote something, as opposed to underlining or scoring a line. Every figure is produced by the build, never typed.

Marks are tiered, and the tiers are never summed, because they are different kinds of evidence: T1, the reader's own words in the margin, at a passage; T2, a non-verbal mark that locates a passage but says nothing; T3, attested contact with a copy and no surviving mark. A fourth tier — a scholar asserting the influence — is what everyone else validates against, and this corpus deliberately does not emit it.

Two of the four centuries here are the same handful of people read deeply. That is a fact about the field's funding rather than about the evidence: Nietzsche's 18,000 annotated pages are photographed and sitting in Weimar, untranscribed.

2 Marked passages recover borrowing, not influence

Attribution can tell you Shakespeare influenced Melville. It cannot tell you which lines. This can — so the first question is whether the lines answer.

Melville is the only reader with both a large body of transcribed marks and a large body of his own public-domain prose. For each book he marked, his marked passages are located inside a full text of the same work and compared against 2,000 random spans of that book, matched to the same lengths. The control is Thomas Beale's Natural History of the Sperm Whale (1839) — a book Melville uncontestedly lifted into Moby-Dick's cetology chapters. If the method cannot recover Beale, it cannot be trusted on Shakespeare.

Source authorMarksLocatedzpDirection zp
Thomas Beale — the control169158+1.690.042+2.850.0015
Ralph Waldo Emerson7653+0.050.44+1.290.064
William Hazlitt285145−0.590.72−1.850.98
Matthew Arnold171131−0.780.81−0.720.78
John Milton480394−1.080.99+0.040.31
Nathaniel Hawthorne320183−1.390.93+1.540.037
William Shakespeare704615−1.880.99+1.150.12

Table 2. z compares the marked passages' lexical similarity to Melville's whole corpus against random spans of the same book. Direction z is the stronger test: how much more the marked passages resemble Moby-Dick and after than they resemble Typee, Omoo and Redburn — again against random spans of the same book, so a gap that is merely a property of the book cancels.

The control fires and nothing else does. The passages Melville marked in Beale are significantly more like the prose he wrote afterwards than like the prose he wrote before — the borrowing is visible in the arithmetic. Shakespeare runs the other way: the lines he marked are less like his own writing than random lines of the same book are. Milton and Hawthorne trend the same way.

What that means, and what it does not

The honest reading is that marked-passage lexical similarity measures source-text appropriation, not influence. Melville's debt to Shakespeare is structural — soliloquy, tragic architecture, Ahab's rhetoric — and a bag-of-words instrument cannot see it. He may also have marked precisely what did not sound like him, which would produce exactly this negative.

Reporting this as a failure of the corpus would be wrong. It is a finding about the reach of the instrument, in the same spirit as this project's earlier result that there is no global rate at which genres mutate.

3 Reading does predict style — one reader at a time

The second measurement asks a coarser question, and gets a cleaner answer. It also began as a false positive, which is the more useful part.

The version that was wrong

Compared against a graph-wide null, the reading edges looked excellent: z = 4.60, beating both attribution sources this project already validates against. Then the pairs were printed instead of trusted. All twenty-one reference pairs were Nietzsche → X, and all six heavy-mark pairs were Melville → X. The entire result was one author's position in the graph. A graph-wide null is structurally blind to that.

The version that holds

Hold the reader fixed. Are the authors a reader demonstrably read more similar to him than the authors in the same graph he did not? Both sides of the comparison then share one endpoint, so the reader's own position, register and corpus density all cancel.

Reader · evidenceReadNot readStyle zpTheme zp
Friedrich Nietzsche · named in his own prose2153+4.240.0005+2.840.002
James Joyce · marked670+3.180.0015+2.420.0045
Herman Melville · marked669+2.360.019+1.140.13
Walt Whitman · attested reading1264−0.360.63+0.060.48

Table 3. 2,000 draws per null. “Style” is lexical similarity, “Theme” an embedding similarity — kept apart, never merged into one influence score.

Three of four readers, at n as small as six. And Whitman, the one null, is the reader whose evidence is thinnest per edge: most of his edges come from a bibliographical handlist — attested reading with no surviving mark, the weakest tier in the corpus. The tier ladder predicted which reader would fail.

4 Two marginal notes

Worth reading verbatim, because they are what this whole apparatus is built to keep:

“Here's a touch Shakspearean— Regan talks of ingratitude!” Herman Melville, in the margin of King Lear, Act II

“Base senses (Kant)” Jacques Derrida, in the margin of Proust's Du Côté de chez Swann, p. 73

Neither is an interpretation of a reader. Each is a reader, at a particular page, in a copy that still exists. Derrida's is the shape the corpus is best at: one writer reaching for a second while reading a third, in his own hand, at a located line.

5 What binds

Not the marginalia. The influence graph these edges are tested against has 77 authors; the corpus names 2,139. Only 24 mark edges and 21 reference edges have both ends in the graph, and four readers overlap at all. Every number above rests on n between 6 and 21, and should be read as a strong hint rather than a settled result.

So the next step is not more marks. It is widening the author set the marks are tested against — which was already deferred work, and now has a concrete target list.

Both analyses are seeded and reproducible; the corpus builds from a licence gate that refuses to run on a source whose terms nobody has read. Two of the eight sources ship structure without text, and one — Beckett's, unpublished and in copyright until roughly 2059 — ships no text at all.

Code & sources

  1. The corpus, its licence gate and its evidence ladder: github.com/2016judea/marginalia-corpus
  2. Both analyses, as a submodule of the parent project: github.com/2016judea/literature-mutationsanalyze_marginalia_prose.py, validate_graph_against_marginalia.py, docs/S7-MARGINALIA.md
  3. Melville's Marginalia Online. Steven Olsen-Smith and Peter Norberg, eds. melvillesmarginalia.org
  4. The Archaeology of Reading in Early Modern Europe. Johns Hopkins, UCL, Princeton. archaeologyofreading.org
  5. Derrida's Margins. Katie Chenoweth et al., Center for Digital Humanities, Princeton. CC BY. derridas-margins.princeton.edu
  6. The Shakespeare and Company Project. Joshua Kotin, dir., Princeton. CC BY. shakespeareandco.princeton.edu
  7. The Beckett Digital Library. Dirk Van Hulle, Mark Nixon and Vincent Neyt, eds. beckettarchive.org
  8. Annotated Books Online. Arnoud Visser and Paul Dijstelberge, eds. CC BY-NC. annotatedbooksonline.com
  9. The Walt Whitman Archive. Kenneth M. Price and Ed Folsom, dirs. whitmanarchive.org
  10. Thomas Beale, The Natural History of the Sperm Whale, London 1839 — via the Internet Archive.
Literature Mutations · S7 ← the parent project August 2026