Head Start at Sixty: Six Decades of Evidence on America's Largest Early-Childhood Experiment
Head Start, launched in the summer of 1965 as the centerpiece of the War on Poverty's investment in young children, is the largest and most intensively evaluated early-childhood program in American history. This review traces its research record from the pessimistic Westinghouse report of 1969, through the randomized Head Start Impact Study, to the modern quasi-experimental and census-linked literature. The pattern that emerges is consistent and, at first glance, paradoxical: initial test-score advantages often shrink or vanish within a few years of school entry, while studies that follow participants into adulthood find higher educational attainment, better health and behavior, greater economic self-sufficiency, and measurable benefits for participants' own children. The review examines why short-run and long-run findings diverge, where the evidence remains genuinely contested, what distinguishes more effective Head Start centers from less effective ones, and what the program's comprehensive, two-generation design implies for parents and educators weighing any early-childhood setting.
In January 1964, President Lyndon B. Johnson used his first State of the Union address to declare a national War on Poverty, and the Economic Opportunity Act passed that summer created a new federal agency, the Office of Economic Opportunity, under Sargent Shriver. Among the facts confronting the new agency was that poverty in America was disproportionately a condition of childhood, and that by the time poor children reached first grade many were already far behind their classmates. In early 1965 Shriver convened a planning committee of pediatricians, psychologists, social workers, and educators — chaired by the Johns Hopkins pediatrician Robert Cooke and including the young Yale psychologist Edward Zigler — and gave it roughly six weeks to design a national program for preschool-aged children living in poverty (Vinovskis, 2005). The pace was so frantic that insiders nicknamed the effort "Project Rush-Rush" (Zigler & Muenchow, 1992).
The committee envisioned a cautious pilot serving perhaps 25,000 children. Shriver, convinced that a small demonstration program would be politically expendable, insisted on launching at national scale. When President Johnson announced Project Head Start in May 1965, it opened that same summer as an eight-week program serving more than 560,000 children, staffed heavily by volunteers (Zigler & Muenchow, 1992; First Five Years Fund, n.d.). Zigler's own insider history records both the exhilaration of that launch and its scientific recklessness: a program built in weeks, evaluated for decades.
What the committee designed, however, proved more durable than the circumstances of its birth. From the beginning, Head Start was more than a preschool. It was conceived as a comprehensive child-development program: early learning, yes, but also medical and dental screening, immunizations, nutritious meals, mental-health and social services, and — unusually for its era — a formal role for parents in classrooms and in program governance. The unit of intervention was the family, not the child alone. In the vocabulary of this Research Series, then, Head Start is not a pedagogy in the sense that Montessori or Waldorf is; local programs choose their own curricula within federal standards. It is a delivery model and a standard of care. And that is precisely what makes its research base distinctive: Head Start is the largest sustained test of whether comprehensive early-childhood investment, delivered at public scale rather than in boutique demonstration projects, changes the arc of poor children's lives.
From demonstration project to national institution
Head Start's evolution has been substantially a policy story. In 1969 the program moved from the Office of Economic Opportunity to the new Office of Child Development in the Department of Health, Education, and Welfare; today it is administered by the Office of Head Start within the Administration for Children and Families, which awards grants directly to more than 1,700 local public and private agencies operating in every state (First Five Years Fund, n.d.). Head Start Program Performance Standards, first issued in the 1970s and revised repeatedly since, set uniform requirements for health, safety, staffing, and family services, and have come to function as a de facto national quality benchmark referenced well beyond the program itself. Congress created Early Head Start for infants, toddlers, and pregnant women in the program's 1994 reauthorization, and the bipartisan Improving Head Start for School Readiness Act of 2007 raised teacher-qualification requirements and strengthened coordination with state early-childhood systems. The delivery system has diversified as well: services are provided not only in centers but also in family child care homes and through home-based options built around regular home visits. In fiscal year 2023 the program was funded to serve 778,420 children and pregnant women in these settings (Office of Head Start, n.d.); since 1965 it has served more than 37 million children and their families (First Five Years Fund, n.d.). Even at that scale, access is far from universal: appropriations have never sufficed to reach all who qualify, and the program serves only a fraction of income-eligible children.
The first verdict: Westinghouse and the birth of "fade-out"
Head Start was barely four years old when it received its first national evaluation, and the result nearly ended it. The Westinghouse Learning Corporation, working with Ohio University under contract to the Office of Economic Opportunity, compared first-, second-, and third-graders who had attended Head Start with matched classmates who had not, drawing data from 104 centers across the country (Westinghouse Learning Corporation & Ohio University, 1969). The conclusions were blunt: the summer programs produced no lasting cognitive or affective gains, and the full-year programs appeared only marginally effective, with initial cognitive advantages fading by the second and third grades.
The study was criticized immediately and has been criticized ever since. Its ex post facto design could not guarantee that comparison children were genuinely comparable, and children served in the program's improvised first summers were hardly a fair test of a mature program. But its influence was immense. Summer-only programs were eventually phased out in favor of full-year services, and the report installed at the center of American education debate an idea that has framed Head Start research for more than five decades: "fade-out," the observation that early test-score gains diminish after children enter elementary school. Much of the subsequent literature is, in one way or another, an argument about what fade-out actually means. Zigler's insider account describes how the program survived the report's aftermath — and several later political near-death experiences — less on the strength of its early evidence than on the loyalty of the families and communities it served (Zigler & Muenchow, 1992).
The experimental test: the Head Start Impact Study
The methodological gold standard arrived with the Head Start Impact Study, mandated by Congress in the program's 1998 reauthorization. Beginning in fall 2002, roughly 5,000 newly entering three- and four-year-olds across a nationally representative sample of 84 grantee and delegate agencies were randomly assigned either to a group with access to Head Start or to a control group without it (Puma, Bell, Cook, & Heid, 2010). Because assignment was decided by lottery among comparable applicants, differences between the groups could be read as causal effects of the offer of Head Start — the kind of inference the Westinghouse design could never support. At the end of the Head Start year, the study found positive effects in cognitive, health, and parenting domains — gains in pre-literacy skills, improved access to dental care, more reading at home — with social-emotional benefits concentrated in the younger cohort.
Then came the finding that dominated headlines: by the end of first grade, most of those advantages were no longer statistically detectable, and the third-grade follow-up reported few significant differences between the Head Start and control groups on cognitive, social-emotional, health, or parenting outcomes, apart from scattered favorable effects within particular subgroups of children (Puma et al., 2012). For the program's critics, this settled the question. For most researchers, it sharpened a different and more interesting one: how can a rigorous randomized study find vanishing test-score effects while studies tracking earlier cohorts into adulthood — reviewed below — find large and durable life benefits? Answering that question has produced some of the best social science of the past two decades.
Reading fade-out correctly: counterfactuals and substitution
Part of the answer is that "the effect of Head Start" depends enormously on what the comparison children are doing instead. Shager and colleagues (2013), meta-analyzing Head Start evaluations of cognitive and achievement outcomes, found an average immediate effect of 0.27 standard deviations — a respectable figure by education-research standards — but also found that about 41 percent of the variation in estimates across evaluations could be explained by research-design features, above all the extent to which control-group children were receiving other early care and education.
Kline and Walters (2016) made the point precise using the Impact Study's own data: roughly a third of Head Start participants were drawn from competing preschool programs, many of them also publicly funded, and many control-group children enrolled in alternative preschools when denied Head Start. A modern experiment therefore does not measure Head Start against no preschool; it measures Head Start against the surrounding ecology of early care. Accounting for this substitution, Kline and Walters estimated that effects for children who would otherwise have remained at home were substantially larger, and that the program's benefit–cost profile is considerably more favorable than the raw experimental contrast suggests. In 1965, when the counterfactual for a poor four-year-old was almost always home care, the same program would have looked far more powerful on the same tests.
The other part of the answer concerns what happens after Head Start. Currie and Thomas (1995), in the study that revived serious econometric interest in the program, found that Head Start significantly reduced grade repetition among white children, with gains persisting into adolescence, while test-score gains for African-American children faded during the elementary years — a divergence that the authors and subsequent researchers have connected, in part, to the lower quality of the schools many Black Head Start graduates went on to attend. Notably, the same study found that children of both groups who attended Head Start or other preschools gained greater access to preventive health services, an effect of the program's non-academic mission that no achievement test would register. Fade-out, on this reading, is not only a property of preschools; it is also a property of the schools that follow them.
The sleeper effects: what adult follow-ups show
The deepest challenge to test-score pessimism comes from studies that simply waited longer. Garces, Thomas, and Currie (2002), using sibling comparisons in the Panel Study of Income Dynamics, found that white adults who had attended Head Start were significantly more likely to complete high school and attend college than their own non-participating siblings, while African-American participants were significantly less likely to have been booked or charged with a crime; the authors also found positive spillovers to participants' younger siblings.
Ludwig and Miller (2007) exploited a natural experiment from the program's founding: in 1965 the federal government provided grant-writing assistance to the nation's 300 poorest counties, creating a lasting discontinuity in Head Start funding at the eligibility cutoff. Counties just rich enough to miss the assistance provide a compelling comparison group, and at the cutoff the authors found a large drop in mortality among children from causes that Head Start's health services could plausibly address, together with suggestive evidence of higher educational attainment. It is worth pausing on what kind of evidence this is: not classroom evidence at all, but evidence for the comprehensive-services model — screenings, immunizations, nutrition — that has distinguished Head Start from ordinary preschool since the Cooke committee.
Deming (2009) supplied the cleanest statement of the paradox. Using sibling comparisons in the National Longitudinal Survey of Youth, he confirmed that Head Start's test-score gains largely faded — and then showed that participants nonetheless gained about 0.23 standard deviations on a summary index of young-adult outcomes spanning high-school graduation, college attendance, idleness, crime, teen parenthood, and health. That gain closes roughly a third of the gap between children from median-income and bottom-quartile families, and is about 80 percent as large as the effect of the famous Perry Preschool demonstration, a far more intensive model program. Skills that surface in adult life, he argued, need not show up on a second-grade test.
Health and behavior tell a similar story. Carneiro and Ginja (2014), identifying effects from discontinuities in the program's eligibility rules, found that Head Start participation reduced behavioral problems, health problems, and obesity among boys at ages 12 and 13, lowered depression and obesity in adolescence, and reduced criminal activity and idleness in young adulthood.
The most statistically powerful evidence arrived with Bailey, Sun, and Timpe (2021), who linked restricted census and American Community Survey records to exact dates and places of birth drawn from Social Security records — yielding an analysis sample roughly four orders of magnitude larger than the longitudinal surveys on which earlier studies relied — and exploited the county-by-county rollout of the program in its first years. Children with access to Head Start went on to complete 0.65 more years of schooling; high-school completion rose by 2.7 percent, college enrollment by 8.5 percent, and college completion by 39 percent, with corresponding gains in adult economic self-sufficiency. For a program whose per-child cost has always been a fraction of the model demonstrations, these are remarkable numbers.
Intellectual honesty requires two caveats. First, this long-run literature is quasi-experimental — built on sibling comparisons, funding discontinuities, eligibility rules, and rollout timing — and it necessarily describes cohorts served decades ago, when the alternative to Head Start was usually no preschool at all. Second, not every reanalysis cooperates. Pages, Lukes, Bailey, and Duncan (2020), replicating and extending Deming's sibling design with an additional decade of data, found no statistically significant impact on earnings, mixed evidence on other adult outcomes, and mostly null or even negative estimates for more recent birth cohorts. Whether that pattern reflects a genuine decline in Head Start's relative advantage as other preschool options expanded, the fragility of sibling-comparison designs, or statistical noise remains unresolved; it is currently the field's most active argument, and honest summaries of Head Start research must include it.
Across generations, and in concert with schools
Two newer strands of research widen the lens. Barr and Gibbs (2022) provided the first evidence that a scaled early-childhood program's effects can cross generations. Comparing children whose mothers were just young enough to have access to early Head Start programs with children whose mothers just missed it, they found that the second generation — children of participants — showed substantially higher high-school graduation (about 11 percentage points) and college enrollment (about 18 percentage points), together with reductions in teen parenthood (about 8 percentage points) and criminal involvement (about 13 percentage points), gains the authors project would translate into wage increases of roughly 6 to 11 percent across the second generation's working lives. They trace these effects partly to improved home environments and to the second generation's own higher preschool participation: parenting itself appears to have changed.
Johnson and Jackson (2019) examined how early investment interacts with what follows. Comparing cohorts differentially exposed to Head Start spending and to court-ordered increases in K–12 school spending, they found the two investments are dynamic complements: Head Start's long-run benefits were larger for children who went on to better-funded schools, and school-spending increases accomplished more for children who had attended Head Start. Early education, on this evidence, is neither an inoculation nor a waste; it is a foundation whose value depends on what is built upon it.
Not all centers are alike
Averages conceal as much as they reveal. Walters (2015), analyzing variation within the randomized Impact Study, estimated that the standard deviation of short-run cognitive effects across Head Start centers is 0.18 test-score standard deviations — larger than typical estimates of variation in teacher or school effectiveness. Some observable practices predicted effectiveness: centers offering full-day service and home visiting produced larger gains, while centers drawing more children away from other preschools produced smaller measured ones. Strikingly, several inputs that dominate policy debate — curriculum choice, teacher educational credentials, class size — were essentially uncorrelated with center effectiveness. The lesson generalizes far beyond Head Start: program labels and structural markers are weak proxies for what actually happens between adults and children in a particular classroom.
Practical implications
For families and educators, several implications follow, and most of them extend well past Head Start itself. First, early test scores are a poor summary of what early education accomplishes; the outcomes that ultimately matter — graduation, health, self-sufficiency, staying out of trouble — have repeatedly diverged from what a kindergarten-readiness assessment predicts. Parents evaluating any preschool should weight the conditions research links to long-run flourishing: language-rich interaction, stable and warm relationships, attention to health, and genuine partnership with families. Second, Head Start's whole-child, two-generation design has been vindicated by the very studies that questioned its test-score effects: the mortality findings of Ludwig and Miller and the intergenerational findings of Barr and Gibbs are effects of a program that screens teeth and coaches parents, not merely one that teaches letters. It is reasonable to ask of any early-childhood setting, whatever its pedagogy, how it engages the family. Third, variation within a model dwarfs the label on the door: a visit to the actual classroom tells parents more than the brand, and features like full-day programming and home visiting carried real weight in the data. Fourth, early gains compound only when followed by good schooling; choosing a preschool is the first move in a longer sequence, not a one-time purchase of advantage. Finally, for income-eligible families, the weight of the evidence clearly favors enrollment: even the most skeptical modern findings concern Head Start's advantage over today's other preschool options, not over staying home, and the program's health and family services have no ready substitute. For educators, meanwhile, Head Start's history is a standing caution about evaluation itself: judging a program four years into its existence, on the narrowest available measures, produced a verdict that six decades of better evidence has substantially overturned.
Sixty years after its improvised first summer, Head Start's scientific legacy may prove as consequential as its social one. It taught researchers that the value of early-childhood investment cannot be read off the next year's test scores, and it taught the field to look for effects where children actually live their lives: in graduation rates, health records, earnings, and eventually in the fortunes of their own children. The record is not a triumph over every doubt — the fade-out of measured cognitive gains is real, quality varies markedly from center to center, and the program's edge over an increasingly rich preschool landscape is genuinely contested. But the accumulated evidence supports a conclusion the Cooke committee could only assert in 1965: that a program which treats the whole child and the whole family, delivered imperfectly and at scale, can still bend the trajectory of a life — and, it now appears, of the life after that one.
References
- Bailey, M. J., Sun, S., & Timpe, B. (2021). Prep school for poor kids: The long-run impacts of Head Start on human capital and economic self-sufficiency. American Economic Review, 111(12), 3963–4001. https://doi.org/10.1257/aer.20181801
- Barr, A., & Gibbs, C. R. (2022). Breaking the cycle? Intergenerational effects of an antipoverty program in early childhood. Journal of Political Economy, 130(12). https://doi.org/10.1086/720764
- Carneiro, P., & Ginja, R. (2014). Long-term impacts of compensatory preschool on health and behavior: Evidence from Head Start. American Economic Journal: Economic Policy, 6(4). https://doi.org/10.1257/pol.6.4.135
- Currie, J., & Thomas, D. (1995). Does Head Start make a difference? American Economic Review, 85(3), 341–364. https://ideas.repec.org/a/aea/aecrev/v85y1995i3p341-64.html
- Deming, D. (2009). Early childhood intervention and life-cycle skill development: Evidence from Head Start. American Economic Journal: Applied Economics, 1(3), 111–134. https://doi.org/10.1257/app.1.3.111
- First Five Years Fund. (n.d.). A brief history and overview of the Head Start program. https://www.ffyf.org/a-brief-history-and-overview-of-the-head-start-program/
- Garces, E., Thomas, D., & Currie, J. (2002). Longer-term effects of Head Start. American Economic Review, 92(4), 999–1012. https://doi.org/10.1257/00028280260344560
- Johnson, R. C., & Jackson, C. K. (2019). Reducing inequality through dynamic complementarity: Evidence from Head Start and public school spending. American Economic Journal: Economic Policy, 11(4), 310–349. https://doi.org/10.1257/pol.20180510
- Kline, P., & Walters, C. R. (2016). Evaluating public programs with close substitutes: The case of Head Start. Quarterly Journal of Economics, 131(4), 1795–1848. https://doi.org/10.1093/qje/qjw027
- Ludwig, J., & Miller, D. L. (2007). Does Head Start improve children's life chances? Evidence from a regression discontinuity design. Quarterly Journal of Economics, 122(1), 159–208. https://academic.oup.com/qje/article-abstract/122/1/159/1924719
- Office of Head Start. (n.d.). Head Start program facts: Fiscal year 2023. U.S. Department of Health and Human Services, Administration for Children and Families. https://headstart.gov/program-data/article/head-start-program-facts-fiscal-year-2023
- Pages, R., Lukes, D. J., Bailey, D. H., & Duncan, G. J. (2020). Elusive longer-run impacts of Head Start: Replications within and across cohorts. Educational Evaluation and Policy Analysis, 42(4), 471–492. https://doi.org/10.3102/0162373720948884
- Puma, M., Bell, S., Cook, R., & Heid, C. (2010). Head Start Impact Study: Final report. Washington, DC: U.S. Department of Health and Human Services, Administration for Children and Families, Office of Planning, Research and Evaluation. ERIC ED507845. https://eric.ed.gov/?id=ED507845
- Puma, M., Bell, S., Cook, R., Heid, C., Broene, P., Jenkins, F., Mashburn, A., & Downer, J. (2012). Third grade follow-up to the Head Start Impact Study: Final report. OPRE Report 2012-45. Washington, DC: Office of Planning, Research and Evaluation, Administration for Children and Families, U.S. Department of Health and Human Services. https://eric.ed.gov/?id=ED539264
- Shager, H. M., Schindler, H. S., Magnuson, K. A., Duncan, G. J., Yoshikawa, H., & Hart, C. M. D. (2013). Can research design explain variation in Head Start research results? A meta-analysis of cognitive and achievement outcomes. Educational Evaluation and Policy Analysis, 35(1), 76–95. https://doi.org/10.3102/0162373712462453
- Vinovskis, M. A. (2005). The birth of Head Start: Preschool education policies in the Kennedy and Johnson administrations. Chicago: University of Chicago Press. https://press.uchicago.edu/ucp/books/book/chicago/B/bo3533726.html
- Walters, C. R. (2015). Inputs in the production of early childhood human capital: Evidence from Head Start. American Economic Journal: Applied Economics, 7(4), 76–102. https://doi.org/10.1257/app.20140184
- Westinghouse Learning Corporation & Ohio University. (1969). The impact of Head Start: An evaluation of the effects of Head Start on children's cognitive and affective development. Report to the Office of Economic Opportunity. ERIC ED036321. https://eric.ed.gov/?id=ED036321
- Zigler, E., & Muenchow, S. (1992). Head Start: The inside story of America's most successful educational experiment. New York: Basic Books. https://archive.org/details/headstart00edwa
