TL;DR
The “science has been wrong before” argument treats all revisions as equivalent, which is where it fails. Asimov’s arithmetic: a flat Earth implies a curvature of 0 per mile, a sphere implies 8 inches per mile, and the oblate spheroid refines that to between 7.973 and 8.027 — the second correction is about 300 times smaller than the first.1 Meanwhile the replication problem is entirely real: of 100 psychology studies, 97% had significant original results and 36% of replications did, with effect sizes halved.2 In social-science papers from Nature and Science, 13 of 21 replicated.3 But researchers’ own prediction markets correctly called 18 of those 21 outcomes in advance,3 the reproducibility project was itself publicly contested in print by other scientists,4 and the machinery kept running. Revisability isn’t the weakness in the argument. It’s the only place the error rate is published at all.
The instruments improve, so the corrections get smaller. That is what the whole argument turns on. Photo: Ousa Chea on Unsplash.5
The Argument, Stated Fairly
Isaac Asimov once received a letter from an English literature student who wanted to correct him. The argument was this: in every century, people have thought they understood the universe, and in every century they were shown to be wrong. Therefore the one thing we can reliably say about current scientific knowledge is that it too is wrong.1
This is worth taking seriously rather than sneering at, because it is a valid-looking induction over a genuine track record. Phlogiston, luminiferous ether, spontaneous generation, stomach ulcers as a stress disease, continental fixity — the list of confidently held, institutionally endorsed, subsequently discarded scientific positions is long and the discarding was often slow.
And the conclusion people draw from it is doing real work in public life. If science is just the current story, then the current story has no more claim on you than any other, and choosing between them becomes a matter of taste or loyalty. That’s the payload. Everything upstream is preamble.
Wrongness Has a Magnitude
Asimov’s reply is the cleanest demolition I know, and it consists mostly of arithmetic.1
Take the shape of the Earth. A flat Earth implies a curvature of zero per mile. That is wrong — but notice how wrong. The actual curvature is about 0.000126 per mile, or roughly 8 inches per mile. The flat-Earth theory is off by 8 inches per mile, which is precisely why it survived so long: over the distances early civilisations measured, it is nearly right, and the difference is hard to detect with the instruments available.1
Then the sphere is also wrong. The Earth is an oblate spheroid: the equatorial diameter is 12,755 km against a polar diameter of 12,711 km, an oblateness of about one third of one percent. Translated into the same units, curvature on the real Earth varies between 7.973 and 8.027 inches per mile.1
Line up the three corrections. Flat to sphere: 8 inches per mile. Sphere to oblate spheroid: about 0.027 inches per mile at the extreme — about 300 times smaller. And when Vanguard I measured the Earth’s gravitational field in 1958 and found the southern bulge slightly larger than the northern, the resulting “pear-shaped” correction was, in curvature terms, millionths of an inch per mile.1
Asimov’s line to his correspondent is the one worth keeping: if you think the spherical Earth is just as wrong as the flat Earth, your view is “wronger than both of them put together.”1 His summary of the pattern is that established theories tend to be refined rather than reversed — not so much wrong as incomplete.1
That’s the structural answer. “Science has been wrong before” is true and tells you almost nothing, because it collapses a distribution of correction sizes into a binary. The useful question is not has this been revised but how large were the revisions, and are they getting smaller?
And Yet the Problem Is Real
Now the part where the sceptic gets their evidence back, because the last fifteen years have handed them a great deal of it.
The Open Science Collaboration replicated 100 experimental and correlational studies from three psychology journals, using high-powered designs and original materials where available. The results, in their own reporting: 97% of original studies had statistically significant results; 36% of replications did. Replication effect sizes were half the magnitude of the originals. 47% of original effect sizes fell within the 95% confidence interval of the replication. 39% of effects were subjectively rated as having replicated.2
It is not confined to psychology’s second tier. Camerer and colleagues replicated 21 social-science experiments published in Nature and Science between 2010 and 2015, with samples roughly five times larger than the originals. Thirteen — 62% — produced a significant effect in the same direction, and replication effect sizes averaged about half the original.3
Both results say the same thing: in these fields a published, peer-reviewed, high-prestige finding has meaningfully less than a certainty of being real, and the effect you read about is on average twice the size of the effect that exists.
Any honest defence of science has to hold that alongside Asimov, not instead of him. Some scientific claims are refinements of refinements with error bars in the millionths. Others are one underpowered study away from evaporating. They are not the same epistemic object and it is not intellectually respectable to defend them with the same sentence.
What the Replication Numbers Also Show
There is a second layer in that data which almost never gets quoted, and it changes the picture.
Before running the replications, Camerer’s team ran prediction markets and surveys in which researchers bet on which findings would hold. Those markets correctly predicted 18 of the 21 outcomes, and market beliefs correlated strongly with the eventual replication effect sizes.3 The Open Science Collaboration found something compatible: replication success was better predicted by the strength of the original evidence than by characteristics of the original or replication teams.2
Read that carefully. The field already knew. Not officially, not in the citations, but distributed across working researchers there existed a reasonably accurate signal about which published results were solid. The failure was not that scientists could not tell. It was that the publication system had no mechanism for that knowledge to show up anywhere.
Which reframes the whole complaint. The replication crisis is not evidence that scientific judgement is worthless; it is evidence that scientific publication was, for a period, a poor readout of scientific judgement. Those are very different diagnoses with very different fixes, and the fixes — pre-registration, larger samples, published replications, registered reports — are the ones that got adopted.
And the self-correction went one level higher still. The reproducibility project was itself challenged in print: a group led by Daniel Gilbert argued in Science that the paper contained statistical errors and that the data were consistent with reproducibility being high, and the Open Science Collaboration replied in the same issue that both optimistic and pessimistic readings were possible and neither was yet warranted.4 A study of whether the field corrects itself, publicly corrected, in the journal that published it. Whatever else that is, it is not a closed system.
Where Self-Correction Genuinely Fails
The honest complication runs the other way, and physics supplies the best-documented case.
Richard Feynman’s 1974 Caltech address described what happened after Millikan measured the charge on the electron and got a number slightly too low, because he used an inaccurate value for the viscosity of air. Plot the subsequent measurements against time and they do not scatter around the true value. Each one is a little larger than the last, creeping upward until they settle at the correct figure years later.6
Feynman’s explanation is about how the error was allowed to persist: when researchers got a value well above Millikan’s they assumed something had gone wrong and hunted for the fault, and when they got something close to his, in his phrasing, “they didn’t look so hard.”6
This matters because it is the strongest possible version of the objection. Physics has unambiguous quantities, decisive instruments and no ideological stake in the electron’s charge — and the community still drifted rather than jumped. Self-correction happened. It was also biased, slow, and shaped by what the previous person had published.
I take two things from it. First, the correction did eventually arrive, and we know about the drift because physicists documented and published it on themselves. Second, the mechanism Feynman describes is exactly the one that makes agreement a weak signal in general: each researcher deferring slightly to the prior published value, and the collective converging on something that no individual’s data quite supported. That’s a cascade, running inside a laboratory.
What Actually Follows
The original question — does self-correction make science non-absolute? — contains a hidden assumption, which is that absolute was ever on the table.
It wasn’t, and no serious account of science has claimed otherwise for at least a century. What is on offer is a body of claims with wildly varying reliability, embedded in an institution that measures and publishes its own failure rate. The alternatives on the table — tradition, revelation, intuition, a confident stranger — are also revisable, and none of them run reproducibility projects on themselves.
So the practical upshot is not “trust science” or “distrust science,” neither of which is a coherent instruction about a body of millions of heterogeneous claims. It is that confidence should be attached to claims, not to the word “science,” and that the relevant question about any specific claim is answerable: how many independent groups have found it, how large is the effect relative to the noise, how much has the estimate moved as measurement improved, and would anyone’s career have been made by refuting it?
The boiling point of water at sea level is not going to be revised. A single priming study with sixty undergraduates might not survive the year. Both are science. Anyone treating them as equally provisional — in either direction — is making Asimov’s correspondent’s error, just with better vocabulary.
Where I’d Hold This Loosely
Three limits.
The replication figures are field-specific, and I have generalised somewhat freely. Social and behavioural psychology were selected for scrutiny precisely because they looked shaky; cancer biology and economics have their own numbers and they differ. Extrapolating a 36% replication rate to “science” would be exactly the collapsing-of-distinctions this piece is arguing against, and I don’t want to have committed it while complaining about it.
Second, the Asimov argument works beautifully for measurement sciences with a stable quantity underneath and much less well elsewhere. There is no equivalent of “8 inches per mile” for a contested claim about how minds or societies work, and in those domains the reassurance that revisions are shrinking is an assumption rather than a demonstration. The strongest version of my case is about physics and chemistry; it thins out considerably as you move.
Third — the caveat against myself — “the field already knew” is a more comfortable conclusion than it deserves to be. Prediction markets calling 18 of 21 is genuinely encouraging about the community’s tacit judgement, but it also means the findings were published, promoted, taught and cited by people who could have guessed they were fragile. Knowing and acting are different, and the gap between them is the part that self-correction, on this evidence, corrects slowest.
Footnotes
-
https://www.sas.upenn.edu/~dbalmer/eportfolio/Nature%20of%20Science_Asimov.pdf ↩ ↩2 ↩3 ↩4 ↩5 ↩6 ↩7 ↩8
-
https://www.sciencedaily.com/releases/2018/08/180827121303.htm ↩ ↩2 ↩3 ↩4
-
https://unsplash.com/photos/white-microscope-on-top-of-black-table-gKUC4TMhOiY ↩
-
https://calteches.library.caltech.edu/51/2/CargoCult.htm ↩ ↩2



