TL;DR
A medical credential is evidence about a person. It is not evidence about a claim, and the two get conflated constantly. The best measurement of the gap is a BMJ prospective study that took 80 randomly selected recommendations from each of The Dr Oz Show and The Doctors and had a team of experienced evidence reviewers go looking for support. For The Dr Oz Show: evidence supported 46% of recommendations, contradicted 15%, and could not be found at all for 39%. Believable or somewhat believable evidence backed just 33%. A specific benefit was described for 43% of recommendations, the magnitude of that benefit for 17%, and potential conflicts of interest were disclosed alongside 0.4%.1 The show averaged twelve recommendations an episode. The authors’ conclusion is unusually direct for a journal: “The public should be skeptical about recommendations made on medical talk shows.”1 The structural problem is that correction is expensive and assertion is cheap — and that asymmetry, not ignorance, is what the attention economy rewards.
Two instruments. Only one of them is subject to review. Photo: Negative Space on Pexels.2
The Category Error
Credentials solve a real problem. You cannot personally evaluate the evidence behind every medical claim you encounter, so you delegate — to someone who has been trained, examined and licensed. That delegation is rational and mostly works.
It contains a specific failure mode. A licence certifies that a person has met a standard of training. It says nothing about whether the particular sentence coming out of their mouth is supported. Those two things are related but not the same, and the distance between them widens the further a claim gets from the speaker’s actual field, the closer it gets to a product, and the more the medium rewards confidence over calibration.
Clinical medicine speaks in probabilities, effect sizes and conditions. Marketing speaks in certainties. When a physician moves onto a platform whose currency is attention, the grammar of the second starts to displace the grammar of the first — not because anyone decided to lie, but because hedged statements perform worse.
Somebody Actually Measured It
This is usually argued anecdotally. It has been measured properly once, and the design is clean.
Investigators randomly selected 40 episodes each of The Dr Oz Show and The Doctors from early 2013, identified every recommendation made, and then had a group of experienced evidence reviewers independently search for and collectively evaluate the evidence behind 80 randomly chosen recommendations from each show — 160 in total.1
The results, for The Dr Oz Show:
- Evidence supported 46% of recommendations
- Evidence contradicted 15%
- No evidence could be found for 39%
- Only 33% were backed by evidence the reviewers judged believable or somewhat believable
The Doctors did better: 63% supported, 14% contradicted, 24% not found.1 Across both shows, at least a case study or better could be found for 54% of recommendations (95% CI 47% to 62%).
Two details matter more than the headline percentages.
Specificity was rare. A specific benefit was described for 43% of The Dr Oz Show’s recommendations, and the magnitude of that benefit for 17%.1 Advice without an effect size is not clinically actionable; it is a mood.
Disclosure was almost absent. Potential conflicts of interest accompanied 0.4% of recommendations across the study.1 That is roughly one recommendation in 250.
At twelve recommendations per episode, daily, syndicated internationally, the arithmetic gets large quickly.
What the Senate Hearing Actually Established
The Dr Oz case is worth dwelling on because it is documented in public record rather than in argument.
In June 2014 Mehmet Oz appeared before the US Senate Subcommittee on Consumer Protection, Product Safety and Insurance, chaired by Senator Claire McCaskill, at a hearing on false advertising for weight-loss products.3 Members took issue with claims he had made about products with thin evidence behind them, notably green coffee bean extract. McCaskill’s summary of the science was blunt: “The scientific community is almost monolithic against you in terms of the efficacy of the three” products at issue.3
The mechanism the hearing identified is the interesting part, and it is not about one man’s beliefs. The Federal Trade Commission’s Mary Koelbel Engle testified that within weeks of an April 2012 episode promoting green coffee bean extract, marketers of a supplement were advertising claims like “lose 20 pounds in four weeks.”4 McCaskill described the effect directly: featuring a product on the show “creates what has become known as the ‘Dr. Oz Effect’ — dramatically boosting sales.”4 The FTC had brought 82 enforcement actions on deceptive weight-loss claims over the preceding decade.4
Oz’s own account is the most useful thing in the transcript, because he did not dispute the framing. He acknowledged using “flowery language” about certain products, and explained his role this way: “My job, I feel, on the show is to be a cheerleader for the audience, and when they don’t think they have hope, when they don’t think they can make it happen, I want to look, and I do look everywhere, including in alternative healing traditions, for any evidence that might be helpful.”3
That is an honest statement of an incentive, and it is incompatible with the job of evaluating evidence. A cheerleader’s task is to sustain hope. An evidence reviewer’s task is to report that three small short-term studies averaging a five-pound loss led their own meta-analysts to conclude that “more rigorous trials are needed.”3 Both roles can be occupied by the same qualified person. They cannot be performed in the same sentence.
Why Correction Loses
The asymmetry has a name — Alberto Brandolini’s aphorism that refuting nonsense takes an order of magnitude more energy than producing it. It is a saying rather than a finding, but the underlying dynamic has been measured.
Vosoughi, Roy and Aral examined every verified true and false story distributed on Twitter from 2006 to 2017 — roughly 126,000 stories tweeted by about 3 million people — and found falsehood diffused significantly farther, faster, deeper and more broadly than the truth in every category.5
Their explanation is frequently misreported, including in an earlier version of this post, so it is worth stating correctly. The mechanism they identify is novelty: false news was more novel than true news, and people are likelier to share novel information.5 Not fear, not disgust, not outrage. Novelty. And the finding that most undercuts the usual account of the problem: robots accelerated the spread of true and false news at the same rate, which means humans rather than bots are why false news travels.5
Apply that to medical claims and the structural disadvantage is obvious. “This single food is damaging you” is novel, actionable and clean. “The dose-response relationship is unclear, the studies are small, and occasional consumption in an otherwise adequate diet is unlikely to matter much” is none of those things. It is also, usually, the accurate statement. I have written separately about how small and concentrated the resulting exposure actually turns out to be — the failure is real without being universal.
The Correction Problem Has Its Own Failure Mode
Because polite correction travels badly, the people who do it often escalate — attacking the credibility of the source rather than the content of the claim, on the reasonable theory that a claim’s protection comes from the authority behind it.
In December 2025 two physicians with large followings had a public disagreement over how to communicate about ultra-processed foods, prompted by a social-media reel. Mainstream coverage described it as exactly that: a public exchange between doctors that “brought attention to how medical professionals communicate nutrition recommendations online,” and which raised, in one outlet’s framing, “bigger questions about nutrition science and online misinformation.”6 The substantive question underneath — whether an occasional tub of popcorn or portion of chicken nuggets meaningfully harms an otherwise adequate diet — is a genuine dose-response question on which reasonable clinicians differ.7
I am deliberately not adjudicating that dispute or characterising either participant’s claims, because the sources I can actually read describe the episode in neutral terms and do not establish who was right. An earlier version of this post did characterise it, at length, on the strength of a since-unreadable social media post. That was not defensible and I have removed it.
What the episode does illustrate, without anyone needing to be the villain, is the shape of the problem: the simplified claim travels, the qualified correction requires explaining toxicology and frequency-versus-intensity, and the exchange gets absorbed as a personality conflict rather than an evidentiary one. That absorption is the real cost. Once an epistemic disagreement becomes a tribal one, the question of what the evidence says stops being the thing anyone is arguing about.
What Actually Helps
Three things follow, none of which require identifying a villain.
Ask for the effect size, not the direction. “Ultra-processed food is bad for you” and “substituting X for Y at this frequency changes this outcome by this much” are different kinds of statement. Only the second can be checked. The BMJ study found magnitude specified 17% of the time, which tells you how often you are being given something checkable.
Treat a disclosed interest as information, not disqualification. The problem in the data is not that physicians sell things; it is that potential conflicts accompanied 0.4% of recommendations. A disclosed commercial relationship lets you discount appropriately. An undisclosed one removes the option.
Separate the credential question from the claim question. “Is this person qualified?” and “is this specific assertion supported?” have different answers, and a yes to the first does very little work on the second. This is the same move as noticing that agreement among experts is weak evidence about a proposition — the authority is real, and it still is not the thing you wanted to know.
Where I’d Hold This Loosely
Five limits, and the first two are corrections to this post rather than caveats about the world.
This post previously made specific allegations about a named physician on the strength of a social-media post that can no longer be retrieved, and republished pejorative epithets about him. Those claims were not supported by any source I can read, and the two mainstream reports of the episode describe it neutrally and do not support the characterisation. I have removed all of it. The failure was mine, and it is the reason the post now argues from a peer-reviewed study rather than from a feud.
Second, it also asserted that a named company bypasses clinical trials, sourced to an Instagram post, and made claims about two other named doctors sourced to an unrelated PubMed record. Both removed for the same reason.
Third, the BMJ study is a 2014 publication analysing 2013 episodes of two American television programmes. Television talk shows are not social media, the era predates the current creator economy, and I am extending the finding by analogy when I apply it to physician influencers. The analogy is reasonable — same credential, same commercial pressure, same absence of review — but it is an analogy, not a measurement of the thing I am actually writing about. Nobody has run that study on Instagram.
Fourth, “no evidence found” is not the same as “false.” A third of The Dr Oz Show’s unsupported recommendations may well have been correct and simply unstudied, which is true of a great deal of clinical practice. The finding is about the gap between confidence and support, not about the truth value of each claim.
Fifth, and against my own framing: the alternative to imperfect popularisation is not perfect information but silence, and silence has its own costs. Physicians who can hold an audience are genuinely valuable, and a standard strict enough to exclude every unstudied recommendation would exclude most useful health communication along with the bad. The critique here is narrow and I want to keep it narrow: state effect sizes where they exist, disclose interests, and do not let a licence do argumentative work it cannot do.
Footnotes
-
https://www.pexels.com/photo/silver-iphone-6-near-blue-and-silver-stethoscope-48603/ ↩
-
https://www.cnn.com/2014/06/17/health/senate-grills-dr-oz ↩ ↩2 ↩3 ↩4
-
https://www.nbcnews.com/better/diet-fitness/dr-oz-effect-senators-scold-mehmet-oz-diet-scams-n133226 ↩ ↩2 ↩3
-
https://www.medboundtimes.com/daily-pulse/doctor-influencers-processed-foods-dispute ↩
-
https://www.indiatoday.in/health/story/popcorn-chicken-nuggets-ultraprocessed-foods-real-villains-harmless-rare-indulgences-2833685-2025-12-10 ↩



