All topicsEvidence gaps controversies

Why parenting studies contradict each other

Why do parenting studies so often disagree, and how should a parent weigh conflicting findings?

Evidence: C — Weak / mixed evidence Last reviewed: 2026-09-09 Discussion ↓

The question

One week breastfeeding "makes babies smarter", the next week it "makes no difference". Sleep training is "harmless" in one study and "floods babies with cortisol" in another. Why do parenting studies so often disagree — and when they do, how should you weigh the conflicting findings?

Short answer

Parenting studies disagree for ordinary, knowable reasons: parents who make different choices differ in other ways too (confounding); only the exciting results tend to get published (publication bias); researchers have many analytical choices and usually report the ones that "worked" (p-hacking); small studies produce noisy, exaggerated results; and studies differ in who was studied, what was measured, and when. Disagreement is usually science converging, not science failing — later, bigger, preregistered studies tend to settle toward smaller, more honest estimates. When studies conflict, trust the body of evidence over any single study, randomised trials over observational ones, and preregistered findings over post-hoc ones [1][2][5][8].

What the strongest evidence says

1. Confounding: the hidden third variable

Most parenting findings start life as observations: parents who do X have children with more Y. But parents who breastfeed longer, for example, also tend to be more educated and wealthier — and education and wealth predict child outcomes on their own. This is confounding: a third variable driving both the choice and the outcome. This book's breastfeeding topic shows the mechanism in action: a large cluster-randomised trial of a breastfeeding promotion programme (PROBIT) found +7.5 verbal IQ points, while sibling studies — which compare breastfed and non-breastfed children of the same mother, holding family background constant — found no difference after adjusting for maternal IQ [9]. Different designs, different answers; the likely explanation is that maternal characteristics, not the milk alone, drive much of the observed association. Whenever you see an observational parenting claim, ask: what kind of parent makes this choice, and could that — rather than the choice — explain the result?

Confounding also runs backwards. Reverse causation is rife in parenting research: the "difficult" baby gets more of the intervention, the anxious parent seeks more advice, the struggling feeder tries more products. An association between a parenting behaviour and a child outcome can easily run child → parent rather than parent → child. Sibling designs and randomised trials are the main defences; ordinary cohort studies mostly can't separate the directions [9].

2. Publication bias: the file drawer

Journals prefer exciting, positive results, so negative studies quietly never get published — the "file drawer" problem. The cleanest demonstration comes from antidepressant trials, where regulators hold the complete dataset: of 74 FDA-registered trials, 31% were never published, and whether a trial saw print depended on its outcome. The published literature made it look as though 94% of trials were positive; the FDA's complete data showed 51%. Pooled effect sizes were inflated by about a third [2]. Parenting research has no FDA forcing disclosure, so the file drawer is likely deeper: small null trials of parenting interventions simply vanish, and the published record looks more positive than reality. This is one reason early findings in a field almost always look bigger than later ones — the first published study is often the luckiest draw, not the truest [5].

3. P-hacking: many choices, one reported

A researcher analysing a dataset faces dozens of small decisions: which outcomes to report, which subgroups to examine, when to stop collecting data, which covariates to adjust for. Simmons and colleagues showed, with simulations and real experiments, that this undisclosed flexibility dramatically inflates false-positive rates — it can become easier to "find" a false effect than to correctly find none [3]. In parenting research the degrees of freedom are enormous: dozens of plausible child outcomes, many ways to define "breastfed" or "sleep trained", many subgroups. A lone significant result in a sea of unreported analyses should be treated as a rumour, not a finding.

The important nuance — and the reason this topic isn't nihilistic: when Head and colleagues text-mined thousands of articles across disciplines, they found p-hacking is indeed widespread, but its distortion of meta-analytic conclusions appears weak relative to real effect sizes [4]. Single studies are fragile; bodies of evidence are sturdier. P-hacking corrupts the bricks more than the building.

4. Small studies, exaggerated effects

Small studies are noisy, and noise plus publication bias produces a specific distortion: only the small studies that got "lucky" with big effects get published, so the literature fills with exaggerated early estimates that later shrink. Button and colleagues showed that low statistical power not only makes false positives more likely but inflates the effect sizes of the "significant" findings that do appear — the "winner's curse" [6]. This book has a textbook case: the viral claim that sleep training "floods babies with cortisol" traces to Middlemiss et al. 2012 — 25 infants, no control group, saliva sampled twice a night, in a hospital where nurses did the settling. The only randomised trial to measure cortisol with a control group (Gradisar et al. 2016) found cortisol declined slightly in the sleep-training groups [9]. The small, uncontrolled study made the headlines; the bigger, controlled one made the correction.

5. Different studies, different answers: populations, measures, eras

Sometimes studies disagree because they're answering different questions. This book's home-birth topic is the exemplar: the Netherlands, England, and the US get different results because their maternity systems are genuinely different — midwifery integration, transfer pathways, and data quality all vary [9]. The sleep-training literature agrees behavioural methods have small effects under 6 months, but two reviews "disagree" because they measured different outcomes (Douglas & Hill 2013 vs Kempler et al. 2016) [9]. And a 2010 meta-analysis claiming home birth doubled or tripled neonatal mortality (Wax et al.) collapsed under scrutiny: it mixed planned and unplanned home births, used non-comparable denominators, and its finding vanished when poor-quality studies were excluded [9]. Before concluding "science can't make up its mind", check whether the studies were ever measuring the same thing in the same kind of people.

6. Unreliable studies pollute the record — then get cleaned out

The vitamin D in pregnancy story from this book's prenatal topic is the perfect illustration of convergence: an earlier review suggested vitamin D might prevent complications; the 2024 update ran a trustworthiness check, threw out 21 studies as unreliable, and concluded the evidence is now very uncertain [9]. The "contradiction" between the old and new reviews isn't science flip-flopping — it's science taking out the rubbish. Similarly, the iron story: a 2012 review found iron significantly reduced low birthweight; the updated review found a nearly identical estimate (risk ratio 0.84 vs 0.81) that was no longer statistically significant with more data [9]. The effect didn't disappear; the precision changed. Disagreement between an old review and a new one is often just the evidence base growing up.

7. The replication crisis, and why it's actually reassuring

Ioannidis's famous 2005 essay argued, via simulations, that most claimed research findings are probably false — especially when studies are small, effects are small, many hypotheses are tested, and analyses are flexible [1]. The assumptions behind that model are disputed [10], but the empirical follow-ups rhymed with it: a massive effort replicating 100 psychology studies found only 36% of replications were statistically significant (versus 97% of originals), and average effect sizes roughly halved [5]. Even the replication project itself was disputed — critics argued its statistics were flawed and reproducibility was actually high [11] — which is itself the point: science argues, in public, with data, and converges slowly.

8. The fixes: preregistration and the hierarchy of evidence

The structural fix is preregistration: researchers publicly declare their hypotheses, outcomes, and analysis plan before collecting data, removing the freedom to report only what "worked". The evidence that it matters is striking: among large US heart-disease trials, 57% reported positive results before trial registration became the norm around 2000, versus 8% after — and registration was strongly associated with the shift toward null findings [7]. Preregistration doesn't make science pessimistic; it makes the published record honest.

And the practical tool is the hierarchy of evidence: systematic reviews of randomised trials outrank single trials, which outrank observational studies, which outrank expert opinion — with study quality mattering within each level [8]. This book's own A–D evidence ratings are built on the same idea. When two parenting studies conflict, don't average them: weight them. A big preregistered RCT beats five small observational studies; a systematic review beats a single trial; and a single observational study reported as a breakthrough deserves a shrug until it's replicated.

What it means for the parents

Permission, not paralysis. The most useful parental takeaway: when the evidence genuinely conflicts, either choice is defensible. If breastfeeding-vs-formula studies disagree on IQ, or home-vs-hospital studies disagree on rare outcomes, that disagreement is itself information — it means the effect, if it exists, is small enough to hide. Small effects don't deserve large guilt. "The science is mixed" is not a failure of the science; it's a licence to decide on other grounds: your values, your circumstances, your clinician's advice.

Don't renovate your life on one paper. The single most actionable rule: no parenting decision should ever turn on a single study. Wait for the review, the replication, the second design. The Middlemiss cortisol scare, the Wax home-birth scare — both were single fragile studies that later evidence corrected [9]. Headlines report the first study; wisdom waits for the second.

A better way to argue about parenting. Couples, grandparents, and internet forums fight proxy wars over single studies ("but this study says…"). The hierarchy gives you a calmer move: "that's one small observational study; here's what the review of the trials says." It depersonalises disagreement — you're not dismissing their concern, you're weighting the evidence.

An honest gap: whether understanding these concepts actually reduces parental anxiety or changes decisions hasn't been tested in trials — the evidence base is meta-research in psychology and medicine, not parenting interventions [1][4][5]. The claim here is modest: these are the questions researchers themselves ask when studies conflict, so they're the right questions for you too. If learning why studies disagree makes the noise feel less threatening, that's a reasonable expectation, not a proven outcome.

What remains uncertain

Benefits and risks in absolute terms

(Adapted for this topic: four disagreements from this book, what likely explains each, and what to trust.)

DisagreementStudy AStudy BLikely explanationWhat to weight most
Breastfeeding and IQPROBIT (cluster RCT of promotion programme): +7.5 verbal IQ pointsSibling studies: no difference after maternal-IQ adjustmentConfounding by maternal characteristics; PROBIT randomised hospitals, not infants [9]The sibling studies for the causal question — but note the honest residual uncertainty
Sleep training and infant cortisolMiddlemiss 2012 (n=25, no control): "elevated" cortisolGradisar 2016 (RCT with control): cortisol declined slightlySmall, uncontrolled, unusual-setting study vs controlled trial [9]The RCT — overwhelmingly
Home birth safetyWax 2010 meta-analysis: neonatal mortality doubled/tripledBirthplace, de Jonge, Hutton: little/no difference in integrated systemsFlawed synthesis (mixed planned/unplanned births, bad denominators); different systems/populations [9]The large prospective studies and pooled analyses of good-quality data
Vitamin D in pregnancyPre-2020 reviews: possible benefits2024 Cochrane: 21 studies excluded as unreliable → very uncertainUnreliable trials polluting the record, then cleaned out [9]The trustworthiness-checked review
Iron and low birthweight2012 review: RR 0.81, significant2015 update: RR 0.84 (0.69–1.03), not significantGrowing evidence base; estimate stable, precision changed — not a contradiction [9]The updated review (and note: "not significant" ≠ "no effect")

Practical considerations

Weighing two conflicting studies — the checklist:

  1. Design first. RCT > prospective cohort > retrospective/cross-sectional > expert opinion [8]. A randomised trial beats five observational studies on the same question.
  2. Size and precision. Small studies are noisy and their published effects are usually exaggerated [6]. Look for the confidence interval, not just the headline result.
  3. Was it preregistered? A preregistered trial or analysis plan is worth more than a post-hoc one — registration is associated with far fewer "positive" findings [7].
  4. Who was studied, when, where? Different populations, measures, and eras explain many "contradictions" that aren't contradictions [9].
  5. One study or a review? Trust the systematic review over the single study; trust the updated review over the old one [9].
  6. Who funded it, and what did they measure? Conflicts of interest and outcome-switching are worth a glance — preregistration is the defence [3][7].
  7. Does the "contradiction" survive the checklist? Often it dissolves: different designs, different populations, or one fragile study versus a solid one.

When to talk to your doctor, midwife, or pediatrician

References

  1. Ioannidis JPA. Why most published research findings are false. PLoS Medicine. 2005;2(8):e124. — [C] https://doi.org/10.1371/journal.pmed.0020124
  2. Turner EH, Matthews AM, Linardatos E, Tell RA, Rosenthal R. Selective publication of antidepressant trials and its influence on apparent efficacy. New England Journal of Medicine. 2008;358(3):252–260. — [B] https://doi.org/10.1056/NEJMsa065779
  3. Simmons JP, Nelson LD, Simonsohn U. False-positive psychology: undisclosed flexibility in data collection and analysis allows presenting anything as significant. Psychological Science. 2011;22(11):1359–1366. — [C] https://doi.org/10.1177/0956797611417632
  4. Head ML, Holman L, Lanfear R, Kahn AT, Jennions MD. The extent and consequences of p-hacking in science. PLoS Biology. 2015;13(3):e1002106. — [C] https://doi.org/10.1371/journal.pbio.1002106
  5. Open Science Collaboration. Estimating the reproducibility of psychological science. Science. 2015;349(6251):aac4716. — [B] https://doi.org/10.1126/science.aac4716
  6. Button KS, Ioannidis JPA, Mokrysz C, Nosek BA, Flint J, Robinson ESJ, Munafò MR. Power failure: why small sample size undermines the reliability of neuroscience. Nature Reviews Neuroscience. 2013;14(5):365–376. — [C] https://doi.org/10.1038/nrn3475
  7. Kaplan RM, Irvin VL. Likelihood of null effects of large NHLBI clinical trials has increased over time. PLoS One. 2015;10(8):e0132382. — [C] https://doi.org/10.1371/journal.pone.0132382
  8. Evans D. Hierarchy of evidence: a framework for ranking evidence evaluating healthcare interventions. Journal of Clinical Nursing. 2003;12(1):77–84. — [framework] https://doi.org/10.1046/j.1365-2702.2003.00662.x
  9. Case studies from this book's topics (full primary citations; case-study analysis lives in each topic): - Breastfeeding/IQ: Kramer MS et al. Promotion of Breastfeeding Intervention Trial (PROBIT): a randomized trial in the Republic of Belarus. JAMA. 2001;285(4):413–420; and Kramer MS et al. Breastfeeding and child cognitive development: new evidence from a large randomized trial. Arch Gen Psychiatry. 2008;65(5):578–584 — [B] (cluster RCT of promotion programme; +7.5 verbal IQ points at age 6.5; maternal IQ not directly measured). Vs sibling studies: Colen CG, Ramey DM. Is breast truly best? Soc Sci Med. 2014;109:55–65 — [B]; Der G, Batty GD, Deary IJ. Effect of breast feeding on intelligence in children: prospective study, sibling pairs analysis, and meta-analysis. BMJ. 2006;333(7575):945 — [B] (no difference after maternal-IQ adjustment). Analysis in this book's breastmilk-vs-formula-outcomes topic. - Sleep-training cortisol: Middlemiss W et al. Asynchrony of mother–infant HPA axis activity following extinction of infant crying responses. Early Hum Dev. 2012;88(4):227–232 — [C] (n=25, no control group, two saliva samples per night, residential hospital with nurses settling). Vs Gradisar M et al. Behavioral interventions for infant sleep problems: a randomized controlled trial. Pediatrics. 2016;137(6):e20151486 — [B] (RCT with control; cortisol declined slightly in intervention groups). Analysis in this book's sleep-training-attachment topic. - Iron and low birthweight: Peña-Rosas JP et al. Daily oral iron supplementation during pregnancy. Cochrane Database Syst Rev. 2015 (updated from the 2012 version) — [B] (2012: RR 0.81, significant; 2015: RR 0.84, 95% CI 0.69–1.03, not significant). Vitamin D: Palacios C et al. Vitamin D supplementation for women during pregnancy. Cochrane Database Syst Rev. 2024 — [C] (21 studies excluded after trustworthiness assessment; evidence very uncertain). Calcium funnel-plot asymmetry: Hofmeyr GJ et al. Calcium supplementation during pregnancy for preventing hypertensive disorders and related problems. Cochrane Database Syst Rev. 2014 — [B]. Analysis in this book's prenatal-vitamins-folic-acid topic. - Home-birth disagreements: Wax JR et al. Maternal and newborn outcomes in planned home birth vs planned hospital births: a metaanalysis. Am J Obstet Gynecol. 2010;203:243.e1–8 — [C] (critiqued in Michal CA et al. Planned home vs hospital birth: a meta-analysis gone wrong. Medscape Ob/Gyn. 2011); vs Brocklehurst P et al. BMJ. 2011;343:d7400, de Jonge A et al. BJOG. 2009;116:1177–1184, and Hutton EK et al. EClinicalMedicine. 2019;14:59–70 — [B] (integrated systems: little/no difference). Analysis in this book's home-birth-vs-hospital-safety topic. - Sleep training under 6 months: Douglas PS, Hill PS. J Dev Behav Pediatr. 2013;34(7):497–507 vs Kempler L et al. Sleep Med Rev. 2016;29:15–22 — [B] (different outcomes measured; both agree effects are small at best). Analysis in this book's sleep-training-attachment topic.
  10. Goodman S, Greenland S. Assessing the unreliability of the medical literature: a response to "Why most published research findings are false". Johns Hopkins University, Department of Biostatistics. 2007 (Working Paper 135). — [working paper] https://biostats.bepress.com/jhubiostat/paper135/ (disputes Ioannidis's assumptions — low prior probabilities, dichotomised p-values, bias factor — and his "hot fields" claim; see evidence log.)
  11. Gilbert DT, King G, Pettigrew S, Wilson TD. Comment on "Estimating the reproducibility of psychological science". Science. 2016;351(6277):1037. — [C] https://doi.org/10.1126/science.aad7243

Changelog

← All topics