TRANSLATE THIS ARTICLE
Integral World: Exploring Theories of Everything
An independent forum for a critical discussion of the integral philosophy of Ken Wilber
John AbramsonJohn Abramson is retired and lives in the Lake District in Cumbria, England. He obtained an MSc in Transpersonal Psychology and Consciousness Studies in 2011 when Les Lancaster and Mike Daniels ran this course at Liverpool John Moores University. In 2015, he received an MA in Buddhist Studies from the University of South Wales. He can be contacted at johnabramson@btinternet.com


Sheldrake Against Mechanism
MAIN ESSAY | Abramson 1 | Visser 1

The Question That Can Be Answered

Distinguishing genuine resonance from statistical dependence

A reply to Frank Visser on Sheldrake

John Abramson / Claude

Rupert Sheldrake and the Revolt Against Mechanism

Frank Visser's essay "Rupert Sheldrake and the Revolt Against Mechanism" is fair. It refuses both cheap dismissals—scientists reject this, therefore it is false and the equally cheap embrace—this is anomalous, therefore it must be true—and it lands on the right demand: a genuinely radical hypothesis earns its place not by exposing the limits of mechanism but by predicting, measuring, replicating and explaining. I want to take that demand seriously enough to act on it, because buried in the essay is a question that, almost uniquely among the questions one can ask about morphic resonance, has a definite answer available right now.

The essay asks it twice. In the body: what distinguishes genuine resonance from ordinary statistical dependence? In the appendix, sharpened: what quantitative regularity does morphic resonance reliably produce that would separate it from competing explanations? Frank offers these as the point at which metaphor has not yet become mechanism. He is right that they are unanswered. He is, I think, wrong that they are the hardest questions in the vicinity. They are the most tractable, because they are not really questions about morphic fields at all. They are questions about statistics.

The missing middle has an antechamber

Frank's appendix identifies a "missing middle" between similarity and causal influence—the unspecified coupling that would make resonance a mechanism rather than a name. I have no missing middle to offer him; nobody does. But there is an antechamber to that room that we can enter without knowing what furniture is inside it. Before you ask what couples these systems, you can ask a prior and far more answerable question: is there any coupling to explain at all, once ordinary shared structure has been removed?

Nearly every empirical morphic-resonance claim has the same logical shape. Some quantity—how fast rats learn a maze, how readily a compound crystallises, how well Wordle players solve today's puzzle—is measured across many nominally independent systems, and it appears to improve as more systems have done it before. The morphic reading is that the systems are in tune across space and time. The mundane reading is that they share a confound: methods diffusing between laboratories, seed crystals drifting through the air, difficulty changing, populations quietly getting better, the sample being drawn differently on different days. Sheldrake's proponents and his critics have argued about which reading is correct for forty years, mostly by trading intuitions.

But the disagreement is not, at bottom, metaphysical. It is a decomposition problem, and decomposition problems have machinery. Given repeated measurements across systems and over time, you can separate the variation that lives within a single system from the variation shared across systems—and, crucially, you can estimate that shared component from neighbouring systems and subtract it, then ask what survives. A genuine collective effect, of the kind Sheldrake predicts, would show up as structure intrinsic to each system that persists after the shared drift is removed. A mundane confound would be exactly the removable shared drift, and would vanish under the subtraction. This is an old idea in statistics—nested variance components, random effects—pointed at a new target. It does not require believing in morphic fields, or disbelieving in them. It is built to return a null.

A test with teeth, shown on physics first

A discriminator is only worth anything if it can actually debunk. So before turning it on Sheldrake I turned it on a case from physics where an apparent collective effect is real in the data and its origin is knowable: a published KCBS contextuality experiment, recorded shot by shot. Correlations between sequential measurements there scattered about ten times more than ordinary sampling noise permits—precisely the signature you would expect if separate runs were somehow "in tune" with one another. Taken at face value, it looks like exactly the kind of anomaly a collective-memory enthusiast would seize on.

It is nothing of the sort. Within any single run the measurements are perfectly well behaved; the excess lives entirely between runs. And when you estimate the shared component from runs recorded close together in time—runs measuring different quantities, so nothing but a common drift could link them—it accounts for around ninety per cent of the excess and can be subtracted away. What remains is a slow, ordinary instrumental drift common to everything measured in the same window. The decisive touch: the effect did not even replicate. Three datasets from the same apparatus, and only one showed the large anomaly. A genuine property of the measurement—a real "capacity"—would appear in all three. A drift excursion appears in one bad session and not its neighbours. The tool returned the mundane verdict on a real dataset, with numbers, and it distinguished "there is a striking correlation here" from "there is a striking correlation here that means anything."

That is the whole point of showing it on physics first. A method that can only confirm is worthless; a method that visibly dissolves an impressive-looking correlation into removable drift is one whose eventual positive verdict, if it ever returned one, would be worth something.

Sheldrake's own ground: the Wordle study

Which brings me to a piece of Sheldrake's own recent work, and the reason this is not idle. In 2026 Georgia Black, Bethany Butzer and Sheldrake published three Wordle studies in the Journal of Anomalous Experience and Cognition, on the hypothesis that as a day's puzzle is solved by more people, later solvers should do better. The pattern is a null in the first study, a lone significant effect in the second, and no replication in the third. The first study is the one to sit with. It was the only one of the three that could connect a player's own timing to their own performance—it gathered individual attempts, times and time zones, rather than pooled global snapshots—which makes it, for all that it was small and (as the authors fairly note) underpowered, the closest thing in the paper to a clean test of the claim. On the real Wordle, played by millions, it found a correlation of essentially zero. So the lone positive result is contradicted from both sides: by the cleanest of the three designs before it, and by a same-instrument failure to replicate after it—significant, then gone. That is the fingerprint of an apparent effect sitting inside a mis-estimated baseline, the same fingerprint the physics dataset wore before it was taken apart.

But before reanalysing anything one has to say what the second study actually measured, because it is easy to get wrong, and the description decides everything. It did not compare morning people with evening people. The researcher queried WordleBot—a New York Times tool that reports, for a given day, the percentage of a large global sample solving on each guess—once in her own morning and once in her own evening, across twelve days, and averaged the results. Each query is a single snapshot of the whole world's pool of finishers up to that moment. So the evening snapshot is not the same players later in their day; it is a later, larger, differently composed slice of the global pool—on average some 1.7 million players against the morning's 640,000. The morning-to-evening axis is therefore not time of day at all. It is how much of the world has already played.

Why this design cannot separate resonance from its rivals

That reframing is the fairest thing one can say about the study, and it is also what makes the result hard to read. Three quantities rise together, in lockstep, from the morning snapshot to the evening one: the cumulative number of prior solvers, which is the morphic dose itself; the composition of the pool, which shifts westward across the day to take in whole populations absent in the morning; and information leakage—the day's answer and its hints, which by the evening snapshot have been circulating on social media and the "Wordle answer today" sites for many more hours. Dose, composition and leakage cannot be separated along this axis, because the axis is the passage of global playing time, and all three are functions of it. Averaging twelve such days does not unbind them.

Of the three, leakage is the one a sceptic will press, and the study's own numbers point straight at it. The significant morning-to-evening increase in the second study was confined to the first and second guesses; the later guesses moved the other way. But the first guess is exactly where a leaked answer acts—someone who has seen the solution simply types it. A genuine resonance that eased the whole process of solving should show across the guesses, not collapse into the one slot a spoiler operates through. The authors do consider cheating, and set it aside on the grounds that success on the later guesses fell in the evening; but that addresses the wrong attempts. It leaves untouched the early-guess result, which is at once the only significant positive signal in the paper and the one most consistent with an answer that has had all day to spread.

The data the test needs

So the obstacle is not a missing column; it is that in Wordle the morphic treatment and its most stubborn confound are very nearly the same physical quantity. "How many have played" is the dose; "how many have played, and posted" is the leakage. Only one handle prises them apart, and it is the shape of the guess distribution: resonance should help across the guesses, whereas leakage piles into the first. A real test therefore has to reach inside that distribution rather than compare two pooled snapshots—and that needs data the published studies did not gather.

What it would need comes in three tiers. The ideal is individual results—or at least results resolved by the solver's timezone and place in the day's global roll-out—carrying the full guess distribution, so that one can ask whether the count of prior solvers predicts performance on the later guesses, where leakage does not reach, and net of which populations are in the sample. The minimum worth having is the raw per-day WordleBot records rather than the twelve-day averages—each morning and evening snapshot with its sample size and its six attempt-percentages—which supports a narrower question: whether the morning-to-evening shift is stable, and stays confined to the first two guesses, once those millions-strong counts are given honest error bars and the day-to-day drift is removed. And the fallback, the published averages alone, can only ask whether the reported contrast survives a properly widened baseline at all. The minimum and the fallback can refute—show the effect dissolving into noise, or into the leakage channel—but neither can confirm resonance, because dose, composition and leakage stay welded together in any pooled comparison. I would rather say that in advance than have it pointed out afterwards.

One piece of the published work bears directly on this and is worth asking for by name. The first, questionnaire-based study collected individual players' attempts, their time zone, and the two-hour window in which they played. That is the only individual-level, timezone-bearing data in the whole enterprise; it is small and it was null, but it is the right shape—the shape the reanalysis actually needs—and it is the natural place to begin. Which of these the authors can share I do not know, so the right first move is not to presume, but to ask what the files contain.

This request is specific, then, and it is fair, and it hinges on access to the data behind Sheldrake's own paper. The paper is published under an open licence; Sheldrake has a long record of posting his data and inviting independent testing—it is, as Frank rightly notes, one of the things that distinguishes him from the ordinary run of heterodoxy—and he has written for these pages himself. A pre-registered reanalysis by someone with no stake in the outcome, using a method demonstrably willing to return a null, is the kind of scrutiny a confident proponent should welcome rather than resist. If this exchange can help put the request in front of him, that would be the most useful thing it could do.

What either answer would be worth

I want to be exact about what is on offer, because overclaiming here would betray the whole point. If the apparent Wordle effect dissolves into shared drift, that is a quantified refutation—more informative than another bare failure to replicate, because it says what the effect was rather than merely that it went away. If, against my expectation, a dose effect survives that reaches the later guesses—with sample composition and information leakage built into the null rather than assumed away—that would be the most disciplined positive result this forty-year debate has produced, and it would be credible precisely because the same instrument debunks the physics case and expects to debunk this one. That qualification is doing real work: a leftover morning-to-evening shift confined to the first two guesses would be no such thing, only an anomaly with a better vocabulary.

Even that, I should say plainly to anyone hoping for more, would not be Frank's missing middle. It would establish that there is something in the room—a regularity not accounted for by the mundane confounds—without saying what couples the systems or how. The mechanism would still be missing. But "we have shown, with a pre-registered test, that a real regularity remains after the ordinary explanations are subtracted" is a categorically different claim from "we find it suggestive," and it is the claim the debate has never been in a position to make in either direction.

Frank's essay ends on the standard by which Sheldrake should be judged—predict, measure, replicate, explain—and observes, fairly, that morphic resonance has not met it. He is right. But the reason it has not is at least partly that no one has built the measuring instrument for its most basic empirical claim: that similar systems are correlated across time beyond what ordinary dependence would produce. That instrument now exists, it has been shown to work on a case where the truth is knowable, and what it waits on is data of a kind I have tried to specify precisely rather than presume. Whether the answer flatters Sheldrake or buries him—or, most likely, tells us only that the published design cannot decide it—it would at last be an answer arrived at scrupulously.

Note

The KCBS inequality is due to Klyachko, Can, Binicioğlu and Shumovsky (2008). The shot-level contextuality data reanalysed here are from a published experiment, cited in full in the working paper in preparation; the reanalysis and its figures are my own.




PLEASE NOTE: Comments containing links are not allowed, to avoid spam.


Widget is loading comments...