A man in an X discussion challenged me to produce a vaccine-safety study.
I pointed him to a nationwide Danish cohort study of 657,461 children. The researchers used linked health registries and more than five million person-years of follow-up to examine a specific question: Did MMR vaccination increase the risk of an autism diagnosis?
The study did not answer every conceivable question about every vaccine. It was not a randomized trial, and its follow-up came from registries rather than repeated contact with each family. Those are fair distinctions. A serious argument should never pretend that one study establishes more than it does.
But it should not pretend that a study establishes nothing because it does not establish everything.
The paper reported its main estimate, confidence interval, methods, covariates, crude cumulative-incidence graph, subgroup analyses, sensitivity analyses, and limitations. The supplementary material even included the SAS input and output for the main model.
His first response was that he could not find the results.
I pointed him to the Results section.
Then the objection became the use of adjusted analysis.
I pointed him to the unadjusted cumulative-incidence graph and explained why age and other differences between groups matter in an observational study.
Then he said he wanted “the raw data.” That phrase can mean two different things: unadjusted summary results, which the paper and supplement provide, or the underlying individual-level registry records, which are restricted by privacy rules. Those are not the same request.
Every question in that sequence can be legitimate. The issue is what happens after an answer is supplied. Does the answer change our confidence, even slightly? Or does a new requirement immediately replace the old one?
That is where skepticism can stop functioning as a method of inquiry and start functioning as a shield.
I cannot know another person’s private motives from a social-media exchange. Neither can anyone else. What we can examine is the visible reasoning pattern: Were the standards defined before the evidence appeared? Did they remain stable afterward? Were they applied just as strictly to evidence supporting the preferred conclusion?
Those questions get us much closer to the difference between critical thinking and conclusion-protection.
A hypothesis must risk something
The broad claim in this discussion was that children who receive more vaccines experience “more harm than benefit on balance.”
That sounds testable until we try to define the test.
What counts as “more vaccines”: injections, doses, antigens, products or visits? Which benefits count: infections prevented, hospitalizations avoided, disability prevented, deaths prevented or reduced transmission? Which harms count, over what period, and how should events of radically different severity be compared? Who belongs in the comparison group? Most important, what result would count against the claim?
Those are not evasions. They are the work required to turn a conviction into a hypothesis.
No single study can settle the net balance of every vaccine, every outcome and every population. A study can ask a narrower question and answer it well or badly. The Danish study asked about MMR and autism. Its value depends on how well it answered that question—not on whether it resolved the entire childhood schedule in one paper.
Until a broad claim is broken into measurable questions, it can absorb almost any result. If one outcome shows no association, another can be substituted. If one time window shows nothing, a different window can be chosen. If a large cohort is unfavorable, it can be rejected for using adjustments. If crude results are supplied, the demand can shift to individual records.
A hypothesis that cannot lose is not strong. It is protected.
A National Research Council report on scientific inquiry describes empirically testable and refutable questions, methods capable of challenging competing explanations, explicit reasoning, professional scrutiny and replication as guiding principles. It also makes an important qualification: there is no single mechanical recipe called “the scientific method.” Observation, measurement, exploration, model building and hypothesis testing may occur in different sequences.
What holds those activities together is exposure to correction. Scientific reasoning makes it possible, at least in principle, to discover that an explanation is incomplete or wrong.
Two different jobs for reasoning
Psychologist Ziva Kunda distinguished reasoning aimed at accuracy from reasoning aimed at a preferred conclusion.
Under an accuracy goal, we try to use the rules and evidence most likely to produce a correct answer. Under a directional goal, we search memory, evidence and argument for a defensible route to the answer we want. That does not necessarily feel dishonest. We still believe we are reasoning. We still require an explanation that sounds plausible to us.
That is what makes the problem difficult to see from the inside.
Confirmation bias is often described too simply as looking only for information that agrees with us. The more revealing form is unequal scrutiny. Supporting evidence is accepted at face value. Contrary evidence is inspected for every conceivable weakness.
In a classic 1979 experiment, Charles Lord, Lee Ross and Mark Lepper showed people with opposing views two studies that appeared to support different conclusions about capital punishment. Participants judged the study supporting their existing position more favorably and found more faults in the study pointing the other way.
Later researchers examined related patterns under labels such as disconfirmation bias and motivated skepticism. We do not always stop thinking when evidence threatens a belief. Sometimes we think harder—but mainly to defeat it.
This does not mean every disagreement is psychological defensiveness. People can have different background knowledge, assumptions or legitimate concerns about source credibility. Researchers have warned against inferring a person’s motivation merely from a directionally biased judgment. The safer test is behavioral:
Are the standards stable?
Are they applied symmetrically?
Does contrary evidence ever reduce confidence?
Does answering an objection move the discussion forward?

The moving standard
Consider the sequence in the MMR discussion:
Show me a safety study.
Where are the results?
Those results are adjusted.
Show me the raw data.
This sequence alone cannot prove anyone’s motive. Nor is asking a follow-up question the same as moving the goalposts. Good scrutiny often produces new questions.
The warning sign is more specific: a stated requirement is met, but meeting it never changes the person’s confidence because another requirement instantly takes its place. The conclusion remains fixed while the standard needed to challenge it keeps changing.
There is also an important fairness point here. The Danish paper used registry follow-up, not active clinical contact. If active follow-up is an essential criterion, that rule should be stated clearly and applied consistently to every registry study—whether its result supports or challenges the preferred position. Rejecting this paper on that ground would still not make it irrelevant to the narrower MMR-autism question. It would identify a limitation to weigh.
The asymmetry becomes easier to see when favorable evidence receives gentler treatment. In this discussion, a self-selected survey promoted by Steve Kirsch was defended as support for a vaccine-autism hypothesis. Such a survey may generate questions, but without representative recruitment, verified records and a suitable comparison group, it cannot provide a reliable population estimate of comparative risk.
Meanwhile, a nationwide registry study was treated as unusable even though it reported crude incidence patterns, detailed methods, sensitivity analyses and the model input and output. If “raw data” meant the individual registry records, public release was not a realistic requirement under Danish privacy rules.
That is a result-dependent hierarchy of evidence:
A weak method supports the conclusion, so its weaknesses are treated as tolerable.
A stronger method challenges the conclusion, so every limitation is treated as fatal.
This is the point at which “question everything” becomes “question only what threatens my position.”
Adjustment is not evidence tampering
The objection to adjusted analysis deserves attention because it sounds reasonable: Why not look at the raw comparison first?
We should. The Danish study did. Figure 2 shows the crude cumulative incidence of autism by age and MMR-vaccination status. The supplement provides additional unadjusted incidence graphs and crude associations for the variables used in the analysis.
But crude does not mean pure, and adjusted does not mean corrupted.
In an observational study, the groups being compared may differ in ways that affect the outcome. Here, MMR status changed over time: a child contributed unvaccinated follow-up before receiving MMR and vaccinated follow-up afterward. Autism diagnosis also varies strongly with age. Ignoring age would not remove researcher judgment; it would build a known distortion into the comparison.
The researchers therefore used Cox regression with attained age as the underlying time scale. In the fully adjusted model, they stratified the baseline hazard by birth year, sex, other early-childhood vaccinations, sibling history of autism and decile of an autism-risk score based on measured factors.
The main estimate comparing MMR-vaccinated with MMR-unvaccinated follow-up was an adjusted hazard ratio of 0.93, with a 95% confidence interval from 0.85 to 1.02. In plain language, the study did not find an increased autism hazard associated with MMR. Because the confidence interval includes 1.00, it also does not establish that MMR reduced autism risk.
The investigators tested the result several ways. They changed the age at which follow-up ended, required two autism registrations instead of one, examined autism categories separately, accounted for the second MMR dose, replaced stratification with covariate adjustment and substituted the individual risk factors for the combined risk score. The estimates remained similar.
The study was not flawless. It was observational, residual confounding remained possible, and the researchers did not review individual medical charts to validate every diagnosis. Those limits belong in the assessment. They do not erase the design, the data, or the result.
Could adjustment be done badly? Certainly. Researchers can choose inappropriate variables, omit important confounders or select models after seeing the outcome. That is why we inspect the rationale for the variables, the stated analysis, crude patterns, alternative models, and sensitivity analyses.
“Adjusted” is not a synonym for “manipulated.” It is an instruction to inspect the adjustment.
The table repeatedly posted in the exchange did not contain the MMR estimate. It was Supplement Table 2, titled “Decile cutpoints and hazard ratios from the autism risk score model.” The crude hazard ratios in that table describe the autism-risk score used in the analysis—not the association between MMR and autism.
That is the danger of searching a paper for something that looks suspicious before establishing what the table actually measures.
Raw data matter, but they are not a universal veto
Requests for data can be entirely appropriate. Reanalysis can uncover errors, test alternative assumptions, and increase confidence. The National Academies distinguishes reproducibility—obtaining consistent computational results with the same data and methods—from replication, in which a new study addresses the same question with new data.
Open data make reproduction easier. Medical records, however, cannot always be posted for public download.
Statistics Denmark provides authorized institutions with secure access to pseudonymized microdata. Its transfer rules prohibit removing individual-level microdata and require disclosure control before analysis results leave its research systems. Those protections limit what an ordinary reader can reproduce at home. They do not turn the underlying records into nonexistent evidence.
The Hviid supplement did provide the full SAS input and output for the main Cox model. That is not the same as releasing the microdata, but it makes the model specification and reported output inspectable.
Restricted data are a real limitation. The proper response is to ask whether qualified researchers can obtain controlled access, whether the variables and methods are described, whether code or model specifications are available, whether sensitivity analyses are reported and whether independent studies using other populations reach compatible results.
Demanding transparency is scientific. Declaring every study that protects confidential medical data invalid is not.
Exploration is not the enemy
There is nothing wrong with noticing a pattern and developing a hypothesis afterward. Much of science begins that way.
The problem comes when we forget which stage we are in.
Brian Nosek and colleagues describe preregistration as one way to preserve the distinction between prediction and postdiction. Researchers specify the primary question, measures, exclusions and analysis before seeing the relevant results. Preregistration does not guarantee a good study. It makes later changes visible.
Exploratory analysis asks, What might be happening here? Confirmatory analysis asks whether a specific prediction survives a test capable of failing. Both matter. Trouble begins when an exploratory pattern is presented as though it had been predicted in advance, or when outcomes, time windows and evidentiary standards are repeatedly revised after the results are known.
This is not a problem unique to vaccine critics or to one political side. Researchers are vulnerable too. Joseph Simmons, Leif Nelson and Uri Simonsohn showed how undisclosed flexibility in sample size, outcomes, exclusions and statistical models can make false-positive findings easier to produce. Scientific safeguards exist because scientists are human.
The same habits should guide public reasoning:
Define the claim before selecting the evidence.
State what observation would weaken it.
Separate exploration from confirmation.
Apply the same evidentiary rules regardless of the result.
Confront the strongest contrary evidence, not merely the evidence easiest to dismiss.
Revise confidence instead of switching standards.
The mirror test
There is a simple way to check whether skepticism has become selective:
If this study had produced the opposite result, would I make the same methodological criticism with the same intensity?
Charles Lord, Mark Lepper and Elizabeth Preston tested a version of this strategy in 1984. Simply telling people to be objective did little. Asking them to consider how they would evaluate the same evidence if it had produced the opposite result reduced biased assimilation.
The question works because it separates the method from the result. It forces us to decide whether a design is unacceptable or merely inconvenient.
Apply that test here. If a nationwide registry analysis of 657,461 children had reported a credible increase in autism after MMR, would the same critic dismiss it because the main estimate was adjusted? Would restricted access to the individual records make the finding unmentionable?
Now reverse the test. If a self-selected parent survey had found no relationship between vaccination and reported autism onset, would he still defend its recruitment, verification and comparison methods?
The answers tell us whether the methodological standard is real or result-dependent.
Can your belief lose?
Before claiming that the evidence supports what you already believe, ask:
What exactly is my hypothesis?
What result would make me reduce my confidence in it?
Did I choose that standard before or after seeing the result?
Would I accept this method if its conclusion favored the other side?
Am I treating a limitation as something to weigh—or as permission to discard the entire study?
Have I demanded stronger evidence from opponents than from my own side?
Am I describing an exploratory clue as though it were a confirmatory test?
When an objection is answered, does my confidence change at all?
If no possible answer changes the conclusion, the discussion is no longer about evidence.
What critical thinking actually requires
Critical thinking is not the ability to find a flaw. Every study has limitations. Every dataset has boundaries. With enough effort, an intelligent person can construct an objection to almost anything.
The harder skill is deciding how much a flaw matters—and applying that judgment consistently.
Real skepticism does not mean distrusting institutions by default or trusting them by default. It does not require accepting one large study as final. It asks whether the study addressed the question it claimed to address, whether its methods handled plausible alternative explanations, and whether its result should change our confidence when considered with the rest of the evidence.
Science advances because its conclusions remain exposed. Better measurements, stronger designs, failed replications and new evidence can challenge them. A conclusion protected by endlessly changing requirements may survive every argument, but that survival tells us nothing about whether it is true.
The most important question in critical thinking is not, Can I find a reason to reject this?
It is this:
What evidence would I allow to change my mind?
If the honest answer is “none,” skepticism is no longer serving the search for truth. It is guarding the door.
Follow the evidence with us
Critical thinking is easy when it challenges someone else. The harder test is applying it to our own conclusions.
If you value careful, evidence-based examinations of the claims spreading across social media, subscribe to A Mind Less Wasted. It’s free, and every new article will be delivered directly to your inbox.
Resources and further reading
Ziva Kunda, “The Case for Motivated Reasoning”, Psychological Bulletin (1990).
Charles G. Lord, Lee Ross and Mark R. Lepper, “Biased Assimilation and Attitude Polarization”, Journal of Personality and Social Psychology (1979).
Kari Edwards and Edward E. Smith, “A Disconfirmation Bias in the Evaluation of Arguments”, Journal of Personality and Social Psychology (1996).
Charles S. Taber and Milton Lodge, “Motivated Skepticism in the Evaluation of Political Beliefs”, American Journal of Political Science (2006).
Ben M. Tappin, Gordon Pennycook and David G. Rand, “Thinking Clearly About Causal Inferences of Politically Motivated Reasoning”, Current Opinion in Behavioral Sciences (2020).
Brian A. Nosek and colleagues, “The Preregistration Revolution”, Proceedings of the National Academy of Sciences (2018).
Joseph P. Simmons, Leif D. Nelson and Uri Simonsohn, “False-Positive Psychology”, Psychological Science (2011).
National Research Council, Scientific Research in Education: Guiding Principles for Scientific Inquiry (2002).
National Academies of Sciences, Engineering, and Medicine, Reproducibility and Replicability in Science (2019).
Anders Hviid and colleagues, “Measles, Mumps, Rubella Vaccination and Autism: A Nationwide Cohort Study”, Annals of Internal Medicine (2019).
Statistics Denmark, Data for research and analysis and Rules on transfer of analysis results.
Charles G. Lord, Mark R. Lepper and Elizabeth Preston, “Considering the Opposite: A Corrective Strategy for Social Judgment” (1984).





