All topics › Evidence gaps controversies
Why do parenting studies so often disagree, and how should a parent weigh conflicting findings?
One week breastfeeding "makes babies smarter", the next week it "makes no difference". Sleep training is "harmless" in one study and "floods babies with cortisol" in another. Why do parenting studies so often disagree — and when they do, how should you weigh the conflicting findings?
Parenting studies disagree for ordinary, knowable reasons: parents who make different choices differ in other ways too (confounding); only the exciting results tend to get published (publication bias); researchers have many analytical choices and usually report the ones that "worked" (p-hacking); small studies produce noisy, exaggerated results; and studies differ in who was studied, what was measured, and when. Disagreement is usually science converging, not science failing — later, bigger, preregistered studies tend to settle toward smaller, more honest estimates. When studies conflict, trust the body of evidence over any single study, randomised trials over observational ones, and preregistered findings over post-hoc ones [1][2][5][8].
Most parenting findings start life as observations: parents who do X have children with more Y. But parents who breastfeed longer, for example, also tend to be more educated and wealthier — and education and wealth predict child outcomes on their own. This is confounding: a third variable driving both the choice and the outcome. This book's breastfeeding topic shows the mechanism in action: a large cluster-randomised trial of a breastfeeding promotion programme (PROBIT) found +7.5 verbal IQ points, while sibling studies — which compare breastfed and non-breastfed children of the same mother, holding family background constant — found no difference after adjusting for maternal IQ [9]. Different designs, different answers; the likely explanation is that maternal characteristics, not the milk alone, drive much of the observed association. Whenever you see an observational parenting claim, ask: what kind of parent makes this choice, and could that — rather than the choice — explain the result?
Confounding also runs backwards. Reverse causation is rife in parenting research: the "difficult" baby gets more of the intervention, the anxious parent seeks more advice, the struggling feeder tries more products. An association between a parenting behaviour and a child outcome can easily run child → parent rather than parent → child. Sibling designs and randomised trials are the main defences; ordinary cohort studies mostly can't separate the directions [9].
Journals prefer exciting, positive results, so negative studies quietly never get published — the "file drawer" problem. The cleanest demonstration comes from antidepressant trials, where regulators hold the complete dataset: of 74 FDA-registered trials, 31% were never published, and whether a trial saw print depended on its outcome. The published literature made it look as though 94% of trials were positive; the FDA's complete data showed 51%. Pooled effect sizes were inflated by about a third [2]. Parenting research has no FDA forcing disclosure, so the file drawer is likely deeper: small null trials of parenting interventions simply vanish, and the published record looks more positive than reality. This is one reason early findings in a field almost always look bigger than later ones — the first published study is often the luckiest draw, not the truest [5].
A researcher analysing a dataset faces dozens of small decisions: which outcomes to report, which subgroups to examine, when to stop collecting data, which covariates to adjust for. Simmons and colleagues showed, with simulations and real experiments, that this undisclosed flexibility dramatically inflates false-positive rates — it can become easier to "find" a false effect than to correctly find none [3]. In parenting research the degrees of freedom are enormous: dozens of plausible child outcomes, many ways to define "breastfed" or "sleep trained", many subgroups. A lone significant result in a sea of unreported analyses should be treated as a rumour, not a finding.
The important nuance — and the reason this topic isn't nihilistic: when Head and colleagues text-mined thousands of articles across disciplines, they found p-hacking is indeed widespread, but its distortion of meta-analytic conclusions appears weak relative to real effect sizes [4]. Single studies are fragile; bodies of evidence are sturdier. P-hacking corrupts the bricks more than the building.
Small studies are noisy, and noise plus publication bias produces a specific distortion: only the small studies that got "lucky" with big effects get published, so the literature fills with exaggerated early estimates that later shrink. Button and colleagues showed that low statistical power not only makes false positives more likely but inflates the effect sizes of the "significant" findings that do appear — the "winner's curse" [6]. This book has a textbook case: the viral claim that sleep training "floods babies with cortisol" traces to Middlemiss et al. 2012 — 25 infants, no control group, saliva sampled twice a night, in a hospital where nurses did the settling. The only randomised trial to measure cortisol with a control group (Gradisar et al. 2016) found cortisol declined slightly in the sleep-training groups [9]. The small, uncontrolled study made the headlines; the bigger, controlled one made the correction.
Sometimes studies disagree because they're answering different questions. This book's home-birth topic is the exemplar: the Netherlands, England, and the US get different results because their maternity systems are genuinely different — midwifery integration, transfer pathways, and data quality all vary [9]. The sleep-training literature agrees behavioural methods have small effects under 6 months, but two reviews "disagree" because they measured different outcomes (Douglas & Hill 2013 vs Kempler et al. 2016) [9]. And a 2010 meta-analysis claiming home birth doubled or tripled neonatal mortality (Wax et al.) collapsed under scrutiny: it mixed planned and unplanned home births, used non-comparable denominators, and its finding vanished when poor-quality studies were excluded [9]. Before concluding "science can't make up its mind", check whether the studies were ever measuring the same thing in the same kind of people.
The vitamin D in pregnancy story from this book's prenatal topic is the perfect illustration of convergence: an earlier review suggested vitamin D might prevent complications; the 2024 update ran a trustworthiness check, threw out 21 studies as unreliable, and concluded the evidence is now very uncertain [9]. The "contradiction" between the old and new reviews isn't science flip-flopping — it's science taking out the rubbish. Similarly, the iron story: a 2012 review found iron significantly reduced low birthweight; the updated review found a nearly identical estimate (risk ratio 0.84 vs 0.81) that was no longer statistically significant with more data [9]. The effect didn't disappear; the precision changed. Disagreement between an old review and a new one is often just the evidence base growing up.
Ioannidis's famous 2005 essay argued, via simulations, that most claimed research findings are probably false — especially when studies are small, effects are small, many hypotheses are tested, and analyses are flexible [1]. The assumptions behind that model are disputed [10], but the empirical follow-ups rhymed with it: a massive effort replicating 100 psychology studies found only 36% of replications were statistically significant (versus 97% of originals), and average effect sizes roughly halved [5]. Even the replication project itself was disputed — critics argued its statistics were flawed and reproducibility was actually high [11] — which is itself the point: science argues, in public, with data, and converges slowly.
The structural fix is preregistration: researchers publicly declare their hypotheses, outcomes, and analysis plan before collecting data, removing the freedom to report only what "worked". The evidence that it matters is striking: among large US heart-disease trials, 57% reported positive results before trial registration became the norm around 2000, versus 8% after — and registration was strongly associated with the shift toward null findings [7]. Preregistration doesn't make science pessimistic; it makes the published record honest.
And the practical tool is the hierarchy of evidence: systematic reviews of randomised trials outrank single trials, which outrank observational studies, which outrank expert opinion — with study quality mattering within each level [8]. This book's own A–D evidence ratings are built on the same idea. When two parenting studies conflict, don't average them: weight them. A big preregistered RCT beats five small observational studies; a systematic review beats a single trial; and a single observational study reported as a breakthrough deserves a shrug until it's replicated.
Permission, not paralysis. The most useful parental takeaway: when the evidence genuinely conflicts, either choice is defensible. If breastfeeding-vs-formula studies disagree on IQ, or home-vs-hospital studies disagree on rare outcomes, that disagreement is itself information — it means the effect, if it exists, is small enough to hide. Small effects don't deserve large guilt. "The science is mixed" is not a failure of the science; it's a licence to decide on other grounds: your values, your circumstances, your clinician's advice.
Don't renovate your life on one paper. The single most actionable rule: no parenting decision should ever turn on a single study. Wait for the review, the replication, the second design. The Middlemiss cortisol scare, the Wax home-birth scare — both were single fragile studies that later evidence corrected [9]. Headlines report the first study; wisdom waits for the second.
A better way to argue about parenting. Couples, grandparents, and internet forums fight proxy wars over single studies ("but this study says…"). The hierarchy gives you a calmer move: "that's one small observational study; here's what the review of the trials says." It depersonalises disagreement — you're not dismissing their concern, you're weighting the evidence.
An honest gap: whether understanding these concepts actually reduces parental anxiety or changes decisions hasn't been tested in trials — the evidence base is meta-research in psychology and medicine, not parenting interventions [1][4][5]. The claim here is modest: these are the questions researchers themselves ask when studies conflict, so they're the right questions for you too. If learning why studies disagree makes the noise feel less threatening, that's a reasonable expectation, not a proven outcome.
(Adapted for this topic: four disagreements from this book, what likely explains each, and what to trust.)
| Disagreement | Study A | Study B | Likely explanation | What to weight most |
|---|---|---|---|---|
| Breastfeeding and IQ | PROBIT (cluster RCT of promotion programme): +7.5 verbal IQ points | Sibling studies: no difference after maternal-IQ adjustment | Confounding by maternal characteristics; PROBIT randomised hospitals, not infants [9] | The sibling studies for the causal question — but note the honest residual uncertainty |
| Sleep training and infant cortisol | Middlemiss 2012 (n=25, no control): "elevated" cortisol | Gradisar 2016 (RCT with control): cortisol declined slightly | Small, uncontrolled, unusual-setting study vs controlled trial [9] | The RCT — overwhelmingly |
| Home birth safety | Wax 2010 meta-analysis: neonatal mortality doubled/tripled | Birthplace, de Jonge, Hutton: little/no difference in integrated systems | Flawed synthesis (mixed planned/unplanned births, bad denominators); different systems/populations [9] | The large prospective studies and pooled analyses of good-quality data |
| Vitamin D in pregnancy | Pre-2020 reviews: possible benefits | 2024 Cochrane: 21 studies excluded as unreliable → very uncertain | Unreliable trials polluting the record, then cleaned out [9] | The trustworthiness-checked review |
| Iron and low birthweight | 2012 review: RR 0.81, significant | 2015 update: RR 0.84 (0.69–1.03), not significant | Growing evidence base; estimate stable, precision changed — not a contradiction [9] | The updated review (and note: "not significant" ≠ "no effect") |
Weighing two conflicting studies — the checklist: