Skip to main content
NIHPA Author Manuscripts logoLink to NIHPA Author Manuscripts
. Author manuscript; available in PMC: 2024 Aug 26.
Published in final edited form as: J Public Econ. 2022 Apr 4;209:104617. doi: 10.1016/j.jpubeco.2022.104617

When Scale and Replication Work: Learning from Summer Youth Employment Experiments

Sara B Heller 1
PMCID: PMC11346829  NIHMSID: NIHMS1957656  PMID: 39188417

Abstract

Two sources of treatment heterogeneity can undermine the scale-up and replication of successful human capital interventions: variation in the treatment itself and changes to the population served. This paper combines two new summer youth employment experiments in Chicago and Philadelphia with previously published evidence to show how repeated study of an intervention as it scales and changes contexts can guide decisions about public investment. Results show that these programs generate consistently large proportional decreases in criminal justice involvement, even as administrators recruit additional youth, hire new local providers, find more job placements, and vary the content of their programs. Using both endogeneous stratification within cities and variation in 62 new and existing point estimates across cities uncovers a key pattern of individual responsiveness: impacts grow linearly with the risk of socially costly behavior each person faces. Identifying more interventions that combine this pattern of robustness to treatment variation with bigger effects for the most disconnected could aid efforts to reduce social inequality efficiently.

1. Introduction

As policymakers consider how to reduce poverty, provide effective alternatives to policing, and find cost-effective ways to support residents’ well-being, they must judge what evidence is promising enough to merit expanded investment. The existence of multiple randomized studies with similar findings is a common bar for considering an approach “evidenced-based.” Research aggregators like Blueprints for Healthy Youth Development and the What Works Clearinghouse give their top ratings to programs with one or two high-quality randomized controlled trials (RCTs), with Blueprints calling programs in its top categories “ready for scale.” Yet “ready for scale” and “scalable” are not the same thing. There are many examples of approaches that succeeded in one or two settings but had different effects when scaled up or moved elsewhere.1 Anticipating which interventions will successfully replicate and scale requires understanding whether two main aspects of program growth or context generate heterogeneous treatment effects: 1) variation in what the treatment is (e.g., program structure, staff quality, and counterfactual opportunities), and 2) variation in who is served (see Al-Ubaydli et al., 2019; Davis et al., 2017, for broader discussions of the challenges of scale).

This paper shows how repeated study of an intervention as it scales and changes contexts, in this case summer youth employment programs (SYEPs), can guide decisions about public investment. SYEPs are a useful demonstration of the uncertainty involved with expanding and replicating “evidence-based” programs. Experiments in Chicago, New York, and Boston have found generally similar patterns of SYEPs’ effects: large declines in criminal justice involvement and violence, despite little improvement in future employment on average (Davis and Heller, 2020; Gelber et al., 2016; Heller, 2014; Kessler et al., 2021; Modestino, 2019). Education impacts are more mixed, with most studies finding small or no improvements in high school or college outcomes (Davis and Heller, 2020; Gelber et al., 2016; Heller, 2014; Leos-Urbel, 2014; Schwartz et al., 2015), and one showing larger benefits (Modestino and Paulsen, 2019). But the Chicago and Boston studies focus on a relatively small subset of the cities’ summer programs, with service providers selected into each evaluation. Given the potential for unrepresentative provider selection and the difficulty other human capital programs have faced with staffing and implementation quality as programs scale (e.g., Allcott, 2015; Bhatt et al., 2021; Jepsen and Rivkin, 2009), there is reason to worry that efforts to expand on SYEPs’ success might diminish their effectiveness. NYC’s SYEP is city-wide and at scale, but existing evaluations focus on how the program was implemented 15 years ago in a different economic and criminal justice context. So it is unclear whether shifts in residential population or local context have changed the impact the SYEP would have today. To make informed decisions about where future investments are likely to succeed, policymakers require more than past successes; they need a better understanding of how net effects respond to variation in what the treatment is, who provides it, and whom it serves.

To help provide this understanding, this paper combines two new randomized controlled trials with evidence from prior SYEP experiments. The first new experiment tests a program called One Summer Chicago Plus (OSC+) in summer 2015, which tripled in size (n = 5,405) relative to the 2012 and 2013 versions of the program studied elsewhere (Davis and Heller, 2020; Heller, 2014). The second new experiment studies Philadelphia’s WorkReady program, which has not been evaluated before, in the summers of 2017 and 2018 (n = 4,497, a subset of the larger program).2OSC+, which targets youth at high risk of violence, scaled up without generating much change in its applicant or complier population, while growing from 5 to 19 providers and changing the youth development programming it provided over time. WorkReady’s universal, city-wide scale means it served a less targeted, less criminally active population than OSC+, and it involved much broader variation in program models across 50–60 providers, with little scope for providers to select into or out of the evaluation. So in addition to providing a valuable replication exercise, the two settings demonstrate what changes as program administrators recruit additional youth, hire new local providers, find more job placements, and vary the content of their programs.

In many settings, this kind of variation in programming, providers, and populations has undermined programs’ ability to replicate originally positive results. But for these SYEPs, results show that both new experiments generated large reductions in criminal justice contact in the first year after random assignment, with declines in some types of arrests and incarceration on the order of 50–80 percent. For cohorts where enough time has passed to measure longer-term effects, there were also indications of arrest declines during years 2 and 3.3 Both experimental designs involved randomization to either different program models or different providers. These tests find no significant variation by model or provider, which could help to explain why effects replicate across different scales and contexts. Neither experiment is fully powered to test this question, however, in part because non-compliance resulted in first stages that are both around 0.3. Still, all existing SYEP experiments show relatively similar proportional declines in criminal justice involvement across considerable variation in the number of providers and jobs, as well as the type of enrichment activities they offer. And all show crime declines with similar time paths that rule out a simple incapacitation mechanism, with effects continuing to accrue after the end of the program. This replicability, at least across the large urban settings where SYEPs have been tested, is somewhat unusual in human capital development programs. It suggests that something about the core program structure, not a specific implementation, setting, or population drives behavioral change.

The fact that proportional declines are relatively constant across different populations implies that the absolute number of crimes prevented is larger for groups with higher control means. With this in mind, the second part of the paper turns to how treatment effects β vary with the population served. I lay out a framework motivating why heterogeneous responses across youth who face different counterfactual risk levels Y0 is of particular interest for questions of scale and replication. In addition to informing optimal targeting, the distribution of gains by counterfactual risk level also informs questions of equity and social justice (see e.g., Heckman et al., 1997). If those whose behavior changes the most in response to intervention are not those at the lowest part of an outcome distribution, there may be tradeoffs between equity and efficiency in deciding whom to serve. More broadly, the shape of the relationship between Y0 and β may also reflect something deeper about the nature of heterogeneity. Debates about the merits of prevention versus remediation are, in part, a question of whether some level of Y0 is so high as to prevent responsiveness to the treatment.

I use two strategies—endogenous stratification within the two new experiments and an analysis of variation in youths’ risk level across all published SYEP experiments—to estimate the relationship between Y0 and β. First, within each new experiment, I generate an index of all the socially costly outcomes with significant treatment effects. I then use endogenous stratification to look at heterogeneity over that risk (Abadie et al., 2018). Point estimates suggest that the highest-risk group has a treatment effect 13 to 26 times as large as the lowest-risk group. But as is often the case with subgroup analysis in a single study, standard errors make it hard to pin down the shape of the relationship between risk group and response.

The second strategy brings cross-study variation to the estimation of the risk-responsiveness relationship. I collect 62 statistically significant impact estimates across the two experiments here and four other peer-reviewed experimental studies of SYEPs. Across locations, outcomes, and subgroups, I plot each group’s control mean for a socially-costly outcome—a measure of baseline risk—against its estimated LATE. The relationship between how common a costly outcome is in a subpopulation and the size of an SYEP’s impact is strikingly linear; SYEPs generate bigger declines for populations where the outcome is more common.4

Whether the correlation between the risk of harmful outcomes individuals face and their responsiveness is causal or a result of other factors that regularly vary with Y0 across settings, its consistency provides guidance for how effects may change in new contexts or expanded programs: magnitudes are likely to scale with the risk level of the population served. The fact that effects grow with Y0, at least among the populations who have been part of prior studies, provides some evidence against the idea that there is a point when it is “too late” to change behavior. And the fact that initially worse-off youth benefit the most suggests a virtuous complementarity between efficiency and equity in targeting decisions.

In practice, policymakers may want to prioritize program goals not measured here, such as providing income transfers or developing labor market skills for a broader population. And if peer interaction is an important mechanism, major shifts in participant populations beyond what has already been tested could diminish treatment effects. Nonetheless, the findings suggest that shifting SYEPs’ focus towards populations at elevated risk of crime and related negative outcomes could help maximize program benefits on those outcomes, and would likely do so across different contexts and program designs. Identifying other interventions that, like SYEPs, seem to be robust to substantial variation in implementation while also generating the largest benefits among those facing the most challenges, could help efforts to reduce social inequality efficiently. More broadly, the paper demonstrates how assessing the different elements required for a program to scale and replicate can provide crucial input into decisions about how and where to invest in human capital interventions.

2. Program Descriptions and Experimental Design

This section provides a brief overview of each program and experimental design, focusing on the comparisons within and across studies that contribute to our understanding of scale and treatment heterogeneity. All programs serve teenagers and young adults, and all share the basic elements of job readiness training, a 6–8 week part-time job placement at or near minimum wage, and some type of professional or personal development activities. Table 1 provides additional details about program elements for the two new experiments reported here, as well as other studies that contribute to the cross-study comparisons in Section 7. For more discussion of program and study design, see Appendix A.

Table 1:

Comparison of SYEP Programs

Program WorkReady
2017/2018
OSC+
2015
OSC+
2012
OSC+
2013
Boston
2015
New York City
2005–2008
Approx. Slots City-Wide 8300/9700 24,000 17,000 20,000 10,000 54,000
Approx. Slots In Study 1100/375 2000 700 1000 1186 54,000
Num. Providers In Study 59/45 19 3 7 1 59
Length 6 weeks 7 weeks 8 weeks 6 weeks 6 weeks Up to 7 weeks
Hours Per Week 20 25 25 25 25 Up to 25
Hourly Wage (Nominal) $7.25 to $10.00 $8.25 $8.25 $8.25 $9.00 $6.00 to $7.15
Job type Government, nonprofits, private sector Government, nonprofits, private sector, infrastructure Government, nonprofits Government, nonprofits, private sector Government, nonprofits, private sector Government, nonprofits, private sector
Eligible Population • 14–21 year olds
• All city residents
• 16–21 year olds
• 49 high-violence
CPS high schools
• 14–21 year olds
• 13 high-violence
 CPS high schools
• 16–22 year olds
• Male only
• Justice agencies and general OSC applicants
• 14–24 year olds
• All city residents
• 14–21 year olds
• All city residents
Separate Adult Mentor No Randomly assigned to 50% of participants Yes Yes No No
Training and Enrichment • Professional development sessions throughout summer • 1 week job readiness training
• 5 hours/week civic leadership curriculum randomly assigned to 50% of participants
• 1 day job readiness training
• 2 hours/day social-emotional curriculum randomly assigned to 50% of participants
• 1 day job readiness training
• 2 hours/day social-emotional curriculum
• Some post-summer activities
• 20 hours job readiness and professional development training • 17.5 hours job readiness, career exploration, financial literacy training

2.1. One Summer Chicago Plus

OSC+ is the 2015 version of the program that was evaluated in Heller (2014) and Davis and Heller (2020), run by Chicago’s Department of Family and Support Services. Two key sets of changes to the 2015 program facilitate the study of what happens as programs scale and adapt over time.5 First, the program almost tripled in size relative to the first 2012 cohort, from 700 to 2,000 program slots. In addition to requiring program staff to identify 3 times as many job placements, the scale-up also required hiring almost 4 times as many local providers to implement the program (19 in 2015 compared to 5 in 2012) and recruiting students at almost 4 times as many schools (49 compared to 13). This growth helps to identify whether a program that purposefully targets young people at elevated risk of violence can maintain the same youth population as it scales, and whether the challenges of recruiting new providers and finding new jobs at a much larger scale diminish treatment effects.

The second set of changes helps to identify whether treatment effects are driven by specific elements of programming. Unlike in the original studies, which included a social-emotional learning curriculum and a separate, paid adult mentor, program operators tested two versions of the 2015 program. Half the participants worked in their jobs for 25 hours per week, with no separate adult mentor and no additional youth development curriculum.6 The other half worked for 20 hours per week, and for 5 hours on Fridays engaged in a Civic Leadership Foundation curriculum focused on civic leadership.7

A total of 5,405 youth applied to OSC+. The research team grouped them by geography to help minimize commute times, then randomly assigned each individual to the control group n=2,911, the job-only group n=1,252, or the job+mentor group n=1,242 within strata. All individuals were also randomly assigned to a provider within strata.

2.2. WorkReady

In contrast to the more targeted OSC+, WorkReady is a universal summer jobs program open to all city youth, run by the Philadelphia Youth Network (PYN). Its operation at broader scale—between 8 and 10 thousand slots during the study summers—changes the population served relative to OSC+, generates the need for more job placements, creates considerably more variation in both local provider agencies and program details, and reduces the scope for providers to select into the evaluation based on their expected treatment effect. In the 2017 and 2018 study years, PYN contracted with 50–60 local agencies to provide one of three program models: service learning to address a community problem, work experience with skill development and ongoing adult interaction, or an internship that included professional development and less intensive adult mentoring.8 Professional development activities were left up to providers, so they varied considerably in structure and content across agencies, ranging from developing business models to sexual health education.9

In summers 2017 and 2018, a subset of WorkReady applicants entered a randomized lottery to allocate the limited number of slots. To minimize disruption to the city-wide program, only a small fraction of slots were assigned by lottery across both summers—about 12 percent in 2017 and 5 percent in 2018. A handful of providers were exempt from the lottery by preference or for logistical reasons, but the lotteried slots were generally representative of the program’s at-scale operation; there was limited scope for contracted agencies to select into or out of the evaluation. The choice margin providers faced was at the point of take-up. To facilitate cooperation among providers, PYN did not force providers to serve treatment youth, and they discouraged but did not prohibit providers from serving control youth (if, for example, a control youth established a relationship with a provider after the lottery).

The experimental design differed depending on how youth applied to the program and the study year. In 2017, youth who applied directly to a local provider were pre-screened for eligibility, then randomly assigned within provider ( N=1,554 across 39 providers). Youth with no pre-existing provider relationships could submit online applications directly to PYN N=1,838. These latter applicants were stratified by geography and age, then individually randomly assigned within strata to both treatment or control groups and to 20 different providers (again generating random variation in provider assignment within strata). In 2018, there was only one lottery consisting of youth who applied to PYN without connections to a provider N=1,105. In this cohort, there was no random variation in provider; PYN matched youth to one of 45 providers. In both study years, I randomized about 20 percent more applicants to treatment than requested slots to ensure that all slots could be filled, even when youth were hard to find, failed to complete paperwork, or were no longer interested. More detail on the experimental designs is in Appendix A.

3. Data

The data come from administrative databases capturing youth contact with various government agencies. I use application and program participation data from the organizations in charge of administering each program. In both cities, I use administrative police records to measure arrests, and in Philadelphia, service records from the City’s integrated data system, known as CARES, to measure juvenile incarceration (including both detention and prison, but not any adult incarceration) and related court-ordered services. The main text focuses on criminal justice outcomes, which are the best measured and most comparable across studies. Appendix B discusses the details of other available education and social service data, which include measures of child protective service receipt, mental health and substance abuse treatment, homeless shelter use, and fertility; it also explains the data linkage process.

Arrest records cover lifetime histories for the Chicago sample but capture fewer years of data for the WorkReady sample (4.5 pre-randomization years for the 2017 cohort and 5.5 for the 2018 cohort). Both cover only arrests made by each city’s police department. I categorize each arrest as violent (a crime against a person), property (all theft, burglary, or larceny), drug (sale or possession), or other (everything else including vandalism, trespassing, illegal use of a weapon, warrant arrests, and other minor offenses) based on the offense’s description. Youth who have never been arrested do not appear in the data, so I assign zero arrests for unmatched youth.10

4. Analytical Methods

I estimate the intent-to-treat effect with the following ordinary least squares regression:

Yist=β0+β1Tis+B2Xist1+γs+εist

where Yist is the outcome of interest for individual i in randomization strata s in period t. Tis is an indicator for individual i being randomly assigned to be offered a program slot. Xist1 is a set of individual i ’s pre-randomization characteristics, and γs is a vector of randomization strata fixed effects (see Appendix F for a list of baseline covariates). For any missing baseline covariates I impute 0s and include an indicator for missingness. I show results with no baseline covariates other than the strata fixed effects (and for OSC+, duplicate application indicators) required for identification as a robustness check in Appendix F. To ease interpretation, the main analysis uses ordinary least squares. Since outcomes are either indicators or counts, Appendix F.2 reports average marginal effects from logistic or Poisson regression (with robust standard errors to relax the assumption that the mean and variance are equal); substantive conclusions are unchanged.

The ITT estimates the effect of receiving an offer to participate in a summer jobs program. Since not all treatment youth and some control youth participated in the program, the ITT will understate the effect of actually receiving program services. To estimate the effect of actually participating, I use random assignment as an instrument for ever participating, defined as having more than 0 hours recorded in program records. Given the two-sided noncompliance, this estimator is a local average treatment effect (LATE) for compliers rather than a treatment-on-the-treated effect. To assess the magnitude of these effects, I report estimates of control complier means (CCMs) as a baseline.11

In addition to reporting heteroskedasticity-robust standard errors, clustered on person in WorkReady where 132 applicants appear in both cohorts, I also conduct randomization inference to test the sharp null of no program effects for anyone in the sample (Athey and Imbens, 2017; Fisher, 1935). Appendix F.3 reports these adjustments, as well as inference adjusting for multiple testing by controlling either the family-wise error rate or the false discovery rate (Anderson, 2008; Benjamini and Hochberg, 1995; Westfall and Young, 1993). I adjust within families of outcome types: the different overall measures of criminal justice involvement (incarceration, juvenile justice services, and total arrests, available in Philadelphia only) and the type of arrests (violent, property, drug, and other in both studies). The appendix also reports impacts and multiple testing adjustments for family outcomes (child protection services, shelter use, and fertility) and behavioral health (substance abuse and mental health services), which are only available in Philadelphia.

To test whether there is significant treatment variation across providers, I use the subset of the data with random variation in provider assignment (all of Chicago and part of the 2017 sample in Philadelphia). I include baseline covariates, strata fixed effects, and provider-specific random effects on the treatment indicator, then use a likelihood ratio test to assess whether the model allowing variation by provider statistically differs from the fixed treatment effect model.12 Since provider assignment is random within strata, the regression does not require the inclusion of provider fixed effects; however, their inclusion does not change the results. Each stratum with random provider variation contains 2–4 providers.

To test for heterogeneity by risk, I implement both the leave-one-out and repeated-split-sample procedures in Abadie, Chingos and West (2018). Details are in Section 7.2.

5. Descriptive Statistics and Compliance

5.1. Sample Composition as Programs Scale

Table 2 shows baseline characteristics for both the WorkReady and OSC+ study populations prior to random assignment. The table demonstrates three key points, with the third discussed in the next section. First, randomization worked: no more of the differences are significant than would be expected by chance, and as shown in the last two rows, the tests of joint significance in both studies confirm that treatment and control groups are balanced.13

Table 2:

WorkReady and OSC+ Descriptive Statistics and Baseline Balance

WorkReady (Philadelphia)
One Summer Chicago Plus
Treatment
Mean
Control
Mean
Complier
Mean
Treatment
Mean
Control
Mean
Complier
Mean
N 1,786 2,711 2,494 2,911
Demographics
 Age 15.7 15.6 15.5 17.4 17.4 17.3
 Male 0.40 0.39 0.34 0.39 0.41 0.43
 Black 0.77 0.79 0.82 0.75 0.74 0.82
 Hispanic 0.12 0.12 0.09 0.22 0.23 0.16
 White 0.05 0.04 0.02 0.01 0.01 0.01
 Other Race 0.06 0.05 0.07 0.02 0.02 0.01
 Is a Parent 0.017 0.015 0.008
Contact with Justice System
 Ever Incarcerated as a Juvenile 0.018 0.023 0.013
 Ever Received Juvenile Justice Services 0.019 0.024 0.024
 Ever Arrested 0.039 0.046 0.039 0.225 0.223 0.217
 Total Number of Prior Arrests 0.050 0.060 0.052 0.628 0.617 0.477
  Violent 0.019** 0.033 0.039 0.186 0.186 0.136
  Property 0.018 0.017 0.008 0.067 0.085 0.065
  Drug 0.002 0.003 0.000 0.097 0.095 0.093
  Other 0.010 0.007 0.006 0.278 0.251 0.183
Education
 Enrolled in School 0.87 0.87 0.93 0.75 0.74 0.76
 Graduated 0.08 0.06 0.04 0.24 0.24 0.24
 Grade 9.6 9.5 9.4 10.7 10.7 10.7
 Days Absent 17.6 18.1 12.8 24.2 23.9 24.0
 Grade Point Average 2.46 2.41 2.44 2.40 2.39 2.35
 Ever Suspended 0.14 0.16 0.16 0.19 0.20 0.22
Receipt of Social Services
 Ever Received Child Protective Services 0.13* 0.15 0.15
 Ever Stayed in a Shelter 0.03*** 0.05 0.05
 Any Behavioral Health Service 0.25 0.27 0.26
  Any Substance Abuse Services 0.01 0.01 0.01
  Any Mental Health Services 0.25 0.27 0.25
P-value on F-test of treatment-control comparison for all non-behavioral health baseline characteristics 0.55 0.61
P-value on F-test of treatment-control comparison for all behavioral health baseline characteristics 0.84

Note: WorkReady N=4497, OSC+ N=5405. Stars on treatment column indicate p-value from test that treatment and control means are equal, adjusting for randomization block (* p < 0.1, ** p < 0.05, *** p < 0.01). Compiler means calculated with Abadie (2003) kappa weights, adjusting for stratified randomization. Enrollment and graduation reported for those with non-missing school records (WorkReady N=4144, OSC+ N=5380). Grade level from application data in WorkReady (N = 4482). Other education measures reported for non-missing data on non-graduates only, excluding charters in Philadelphia (for days absent, WorkReady N=2336, OSC+ N=5308; for suspensions, WorkReady N=2337, OSC+ N=5308; for GPA, WorkReady N=2228, OSC+ N=5158). The Philadelphia school year has 180 days; Chicago has 178. Behavioral health services, including substance abuse and mental health services, are held in a separate data set to maintain confidentiality of HIPAA-covered data and are thus a separate F-test from other baseline characteristics.

Second, by comparing the populations across studies, we can investigate how the applicant population changes with the scale of an SYEP. Relative to the initial 2012 OSC+ study, the 2015 program roughly tripled the number of slots and quadrupled the number of schools in which recruiting occurred. Yet many of the characteristics of the 2015 population—about 40 percent male, 22 percent with an arrest record, GPAs of 2.4, and 24 days absent—are fairly similar to the initial 2012 cohort. The 2012 applicants were a little younger (16.3 rather than 17.4 on average, because 14- and 15-year-olds were eligible in 2012), but also about 40 percent male, 20 percent with an arrest record, 2.4 GPAs, and 33 days absent (Davis and Heller, 2020; Heller, 2014). The one major difference between samples is that the current study participants were about 23 percent Hispanic, but only 3 percent of the initial study. Given the residential segregation in Chicago, this change is likely due to the expansion of the program into more Hispanic areas. The relative similarity of observable characteristics across study cohorts demonstrates that even the tripling of the applicant pool between the 2012 and 2015 studies did not dramatically shift the make-up of the applicant pool. Maintaining similar youth populations as programs grow may not be feasible in every setting. But in a large city like Chicago, providers were able to identify and recruit a broader group of youth without much change in school engagement or criminal involvement.14

By contrast, WorkReady is a larger, less-targeted program open to every young person in Philadelphia. It therefore serves a less disadvantaged population. Applicants were considerably less involved in the criminal justice system (only 4 percent had an arrest record) and slightly more engaged in school (18 days absent and 15 percent with a prior suspension, compared to 24 days absent and 20 percent with suspensions in Chicago).15 The additional social service data available in Philadelphia highlights that despite the differences that come with universal eligibility, SYEP applicants still face more challenges than the average City youth. About 14 percent of applicants were in families that previously received child protective services (some when they were very young); 4 percent had stayed in a homeless shelter; 1.6 percent were parents themselves; and about 26 percent had received behavioral health services, mostly mental health care. These rates are about 50 percent higher than the population of youth in the City’s service database who did not apply to WorkReady.

5.2. Compliance

Both studies faced compliance challenges. In Chicago, this was due to an error in a new online system that the City implemented to transmit lists of lottery winners to program providers. During the initial recruitment period, the city’s contracted programmer unintentionally allowed unrestricted access to all applicants, listing the control youth after the treatment youth rather than withholding the waitlist from view. Because of this error, agencies could initially click through to view all of their control group applicants, though not all of them did so prior to the research team catching the error. Additionally, as had been the case in the past, not all treatment youth could be reached by their assigned agency or were still interested in participating. The resulting first stage for OSC+ is 0.26 (F-stat = 454), with 46.5 of the treatment group and 20.4 of the control group participating.

In Philadelphia, non-compliance was the result of a strategic choice by the program administrator, PYN. To minimize provider resistance to the new lottery system and increase broader outreach to new youth populations, PYN encouraged but did not force providers to adhere to random assignment. In the first study year, 44.5 percent of treatment youth and 18.5 percent of control youth worked at least one day, for a first stage of 0.26. In the second year, PYN worked hard to reduce some of the barriers providers faced in serving youth with whom they had no pre-existing relationships. Combined with the fact that the 2018 study focused solely on an applicant pool that did not apply directly to providers (so providers did not know the identity of control applicants), this strategy successfully increased take-up to 67 percent among the treatment group and reduced it to 9 percent among controls, for a first stage of 0.58. Pooling the two cohorts, the first stage is 0.34 (F-stat = 603).

Given these recruitment processes, complying with randomization is a function of both provider and individual decisions. The third column of Table 2 shows the resulting average complier characteristics for each study.16 On most characteristics, compliers are not all that different from the applicant population. In both cities, compliers are somewhat more Black than the sample as a whole. In Philadelphia, they are slightly less likely to have been incarcerated and more likely to have been enrolled in school in the prior year. In Chicago, compliers were less involved in the criminal justice system than the full applicant pool. But for the most part, there is not a lot of observable selection into compliance in either setting; the main differences across studies come from the differences in who applied to each program.

6. Results

This section presents SYEP impact estimates for the crime outcomes that are available across both cities. Education and other family and health outcomes, which are less consistently available across the two cities and so less conducive to an analysis of treatment heterogeneity across contexts, are reported in Appendix C.17 The similarity of proportional changes across the crime estimates reported here, as well as in previous studies, is useful on its own as a replication exercise. It also suggests that variation in what the treatment is—program model, staff capacity and experience, and local context—is less important to generating treatment effects than the basic intervention approach itself.

6.1. Main Effects

Table 3 shows the estimated ITT and LATE for criminal justice involvement in the first year after randomization. In both cities, the SYEP programs generate proportionally large decreases in contact with the criminal justice system. In Philadelphia, being offered the program results in 1 fewer arrest per 100 youth, a statistically significant 36 percent decline. Due to a first stage far less than 1, the effect on compliers is much larger: 3 fewer arrests per 100 participants, a 65 percent decline relative to the CCM. The decline also translates to a decrease in juvenile incarceration. Although the incarceration result is marginally significant p = 0.08 , it is also proportionally huge: Incarceration drops by 1.5 percentage points, or almost 80 percent among WorkReady participants.

Table 3:

Program Impacts in the First Year After Randomization

ITT CM LATE CCM
WorkReady (Philadelphia)

Any Juvenile Incarceration −0.005* (0.003) 0.014 −0.015* (0.008) 0.019
Any Receipt of Juvenile Justice Services −0.002 (0.003) 0.013 −0.006 (0.009) 0.018
Total Number of Arrests −0.010** (0.005) 0.028 −0.030** (0.015) 0.046
 Number of Violent Arrests −0.002 (0.003) 0.011 −0.006 (0.010) 0.016
 Number of Property Arrests −0.002 (0.003) 0.008 −0.006 (0.008) 0.012
 Number of Drug Arrests −0.003 (0.002) 0.004 −0.007 (0.004) 0.007
 Number of Other Arrests −0.004*** (0.001) 0.004 −0.011*** (0.004) 0.010

One Summer Chicago Plus

Total Number of Arrests −0.023 (0.015) 0.176 −0.087 (0.057) 0.166
 Number of Violent Arrests 0.005 (0.006) 0.037 0.021 (0.022) 0.012
 Number of Property Arrests 0.000 (0.005) 0.020 0.002 (0.018) 0.019
 Number of Drug Arrests −0.012** (0.006) 0.034 −0.046** (0.022) 0.052
 Number of Other Arrests −0.017* (0.009) 0.084 −0.063* (0.035) 0.082

Note: WorkReady N=4497, OSC+ N=5405. Table shows estimated intent-to-treat (ITT) and local average treatment effects (LATE), controlling for baseline covariates and randomization block. CM is control mean; CCM is control complier mean, rounded to 0 when estimate is negative. Robust standard errors in parentheses, clustered by person for WorkReady, where the same person can appear in both cohorts.

*

p<0.1

**

p<0.05

***

p<0.01

The point estimate on total arrests is larger in levels but proportionally similar in Chicago. OSC+ participants have almost 9 fewer arrests per 100 participants than control compliers (a 52 percent decline), although the result is not quite statistically significant p=0.125 . Across both cities, there are statistically significant and substantively large declines in drug and other arrests, ranging from 20 to 100 percent drops relative to baseline means.18 The point estimates on violent and property crime are similar in magnitude, negative, and proportionally large in Philadelphia, but not statistically different from zero. Where data are available for longer-term follow up (see Appendix G), WorkReady significantly decreases property crime by about two-thirds in the second year post-random assignment (0.6 fewer arrests per 100 youth offered the program relative to a control mean of 0.9). There is a similarly large and statistically significant decline in property and other crime arrests in year 3 for OSC+ (ITT of −0.6 per 100 youth relative to a control mean of 1.4 for property crimes, and −1.3 per 100 relative to a control mean of 6.2 for other crimes), generating a drop in total year 3 arrests of 2.2 per 100, a 20 percent decline. The pattern of effects here previews the risk-responsiveness relationship I investigate below: Outcomes with larger control means also show larger point estimates.

A number of the main results cross traditional significant thresholds after adjustments for multiple testing (see Appendix F.3). Yet this seems more likely due to the power limitations that come with non-compliance than a serious risk of Type I errors. The probability that both cities’ results are Type I errors is considerably lower than the probability either one is in isolation, so the built-in replication across sites should increase confidence in the results. And the fact that criminal justice involvement has fallen in all previous SYEP studies also strengthens confidence in the result, despite the individual p-values rising slightly above the 0.1 cutoff after multiple testing adjustments.

It is worth noting that the type of crimes that respond to SYEPs here differs somewhat from prior studies of OSC+. In prior work (Davis and Heller, 2020; Heller, 2014), OSC+ crime declines were driven by decreases in violent-crime arrests (which is why violence was the primary pre-specified outcome in the pre-analysis plan). While it is possible the shift has to do with substantive changes in programming, it is also the case that the overall level of violent-crime arrests in the control group is much smaller in the current studies, despite very similar baseline rates of arrest (1.1 and 3.7 violent-crime arrests per 100 control youth in year 1 here, relative to 7.4 per 100 in the 2012 study and 10.8 per 100 in the 2013 study).19 Looking at proportional changes, the 30–40 percent declines in violence found in prior studies are within the confidence intervals of both the WorkReady and OSC+ studies here. So although there is not enough precision to confirm a similarly-sized decline in violence in these studies, I can not rule it out. The time path of effects is also similar across studies. Appendix Section D presents evidence that as in prior work, being busy over the summer is not the key mechanism; program effects continue to accrue long after the program ends.

6.2. Variation across program structure and delivery

Because variation in what the treatment actually is can contribute to difficulty with scale and replication, the research design included experimental variation in both program model and local provider. In Chicago, I randomly varied the structure of mentorship and the enrichment curriculum across two treatment arms, and in both cities I generated random variation in provider assignment. In neither case can I reject the null that treatment effects are the same (neither across treatment arms nor across providers). However, in part due to the relatively high non-compliance, these tests are quite underpowered (see discussion and results in Appendix Section E).

Across-study comparisons are also informative about whether mentors, particular enrichment curricula, or variation in provider experience and quality are crucial to program success. As discussed above and shown in Table 1, from 2012 to 2015, OSC+ hired almost 4 times as many providers, introducing more variation in provider experience and characteristics; tripled in size, requiring a big expansion in the number of jobs provided; and changed from mentorship with a social-emotional learning curriculum to either a civics curriculum or no curriculum and no mentor. The Philadelphia study expands the scale even further, with the city-wide programming encompassing broader variation in program models, a wide range of enrichment curricula, no separate paid mentors, and about 3 times as many providers as in even the expanded OSC+.

The consistency of criminal justice involvement declines suggest that neither scaling up nor adding variation in program models and delivery undermines the programs’ ability to reduce criminal justice involvement. Of course, this does not rule out the possibility that changes in scale, providers, and program details can generate substantively important treatment heterogeneity; well-powered tests of specific kinds of variation would be useful in future work, and one particular pattern of heterogeneity is discussed more in the next section. But the replications here do demonstrate that none of the factors that varied across studies constitute the “key ingredient” for having a positive impact. Rather, something about the basic program structure that all these SYEPs share is enough to reduce crime.

7. Variation by individual risk level

Given that efforts to replicate or scale SYEPs seem relatively robust to variation in implementation, a key remaining question for future investments in these programs is whether the population served affects program impacts. In the main results, the proportional change in crime outcomes relative to the control group is similar across contexts; that implies the absolute magnitude of the point estimates is larger when control means are larger. Motivated by this pattern, which also holds across a number of subgroup results in the appendix on other socially costly outcomes, this section focuses on estimating individual heterogeneity by risk, defined as the level of some counterfactual costly outcome Y0 .

A typical approach to heterogeneity is to search for how variation in X , rather than variation in Y0 , affects β . This can be effective if there are a small number of observables that drive large differences in treatment heterogeneity, but it has limitations. Most studies are powered to detect main effects but not subgroup effects; searches over multiple interactions further reduce power by generating the need for multiple testing adjustments; and more flexible machine learning approaches require additional sample splitting to get inference right. So subgroup analyses are often quite under-powered. Plus, within any single study, it is difficult to tell whether a particular characteristic drives bigger treatment effects because something about the group causes a differential treatment response, or whether that X just happens to be correlated with something else about the setting that matters, such that targeting the group in a different setting would not be as effective. By focusing instead on the direct relationship between Y0 and β across settings, I aim to draw some broader lessons about patterns of treatment heterogeneity.

To elucidate how the relationship between Y0 and β matters for decisions about investing in a new or expanding program, this section begins with a conceptual framework for thinking about the risk-responsiveness relationship and targeting decisions. It then uses impact estimates across multiple outcomes and studies to estimate the shape of that relationship and discusses what we learn from the results.

7.1. Framework relating heterogeneity, risk level, and targeting

Consider the set of outcomes, a vector Y , that SYEPs affect. Across existing studies, these are frequently counts or indicators for crime or harmful health and welfare outcomes. So define Y such that all elements Yi0 , and decreases in each Y are socially beneficial. Different types of people, indexed by θ , have different Y0s and may respond differently to treatment. So in a potential outcomes framework, each row of the treatment effect vector, βθ,Y=EY1θEY0θ , may vary by group and by outcome.

Each occurrence of each Y has an associated social cost, CY . All else equal, policymakers considering how to generate the most social benefits with SYEPs (or any program) would like to target the θ groups who make βθ,Y'CY as negative as possible.20 Because each occurrence of Y is socially costly, a social planner would want to maximize the number of events prevented, weighted by cost. In other words, the absolute magnitude of the treatment-driven decreases in counts matters more than the size of the proportional changes; 3 fewer events from a baseline of 12 is more beneficial than 1 fewer event from a baseline of 1 (i.e., 3CY>CY ), even though the former is a 25 percent change and the latter is a 100 percent decline. A detailed dive into the social costs of the different behaviors SYEP affect—crime by type, child abuse and neglect, mortality, etc.—involves specifying a social welfare function, requiring normative judgments beyond the scope of this paper. So here I focus on what the data can tell us about the type of youth with the most negative βθ,Y , which is a key input, if not a final answer, to optimal targeting decisions.

To make the implications of the risk-responsiveness relationship for targeting concrete, consider treatment heterogeneity across variation in one counterfactual outcome, Y0 . Figure 1 shows three stylized examples of how responsiveness to treatment could vary by risk of this outcome, plotting theoretical variation in β0 across a population with different levels of risk of the outcome in the absence of the program, Y0θ . Each panel represents a different structure of treatment heterogeneity. Panel A shows a case where treatment shifts the outcome down by some constant amount α for everyone, regardless of their risk level. Across most of the distribution, βθ is a constant, α , for any choice of θ and the corresponding Yθ . The exception is for very low risk individuals whose Y0<α . Since Y1 can not be negative, there is a floor effect, with βθ getting smaller when the population served is at almost no risk of the negative outcome. Here, policymakers would generate equivalent social gains regardless of which population they served, as long as they chose θ such that Y0θ>α . Panel B, by contrast, shows a case where the treatment effect is a proportional shift in the outcome, βθ=αY0θ , 0<α<1 . Here, policymakers should want to serve individuals as far to the right of the graph as possible; bigger Y0s correspond to bigger social gains.

Figure 1:

Figure 1:

Stylized Treatment Effects Relative to Counterfactual Outcomes with Different Types of Heterogeneity

Note: Figure shows theoretical shape of risk-responsiveness relationship under different types of treatment heterogeneity. See Section 7.1 for discussion.

Panel C shows a more complicated case, motivated by the idea that behavior may only change on the margin. Suppose, for example, that those very deeply involved in crime are committed enough to their behavior that a summer intervention would have little effect. And suppose that those barely involved in crime commit offenses rarely enough that an intervention is unlikely to matter much. In this case, serving types of people with high levels of the outcome in the absence of the program is not the same as serving people with big changes in outcome due to an intervention. It is only participants in the middle whose behavior might be shifted, those who are close enough to the margin of crime for a time-limited intervention to change their decision-making.21 Here, policymakers should try to identify those for whom Y0θ is in the responsive region, ideally close to the peak of the βθ function. Of course, these are stylized examples. There are many possible forms the relationship between Y0 and β could take. The point is that the shape of the relationship matters, both for targeting choices and for understanding the nature of the behavioral response more broadly.

7.2. Estimating effect heterogeneity by baseline risk level

Estimating this relationship empirically is challenging, because Y0 is not observed for the treatment group. I take two different approaches to understanding the relationship between Y0 and βθ in the SYEP setting.22 First, I perform an endogenous stratification exercise. Abadie et al. (2018) show how to use the relationship between the Xs and Y0 in the control group to predict Y0 for the treatment group, then estimate heterogeneous effects across the predicted risk groups. The procedure involves regressing the outcome on all the baseline covariates using just the control group, predicting Y0 for the treatment group using those regression coefficients, separating the observations into groups by their level of Y^0 , then estimating separate treatment effects for each group. To avoid the finite sample bias that comes from fitting a prediction regression within sample, the authors suggest using both leave-one-out regression and repeated split samples.

I note upfront that predicting a single Y0 and estimating treatment heterogeneity by Y^0 has practical limitations. The adjustments required to avoid bias in finite samples reduce power, even more than a typical subgroup test. The approach typically estimates average treatment effects for two or three parts of the Y^0 distribution. It is difficult to extrapolate the full shape of the risk-response distribution from two or three points. And given how many different outcomes SYEPs seem to affect, we may re-introduce multiple testing concerns by repeating this exercise as many times as available outcomes.

To address the latter issue, I perform the heterogeneity test with an index that combines information about the underlying risk of socially costly behavior across measures. Although this risks masking heterogeneity that varies by outcome, it has the benefit of increasing power by combining information on outcomes that tend to move in the same direction, and reducing the number of hypothesis tests (see, e.g., Kling et al., 2007). To do this, I standardize each outcome with a statistically significant main effect and generate an unweighted average of these outcomes by city.23 In Chicago this includes drug and other arrests; in Philadelphia it includes incarceration, total arrests, and child protective services.24 Each variable is standardized on the control group, averaged together, then re-standardized so that the standard deviation of the index is 1.

Table 4 reports the overall ITT effect on the index, a 0.065 standard deviation decline in Philadelphia and 0.051 decline in Chicago, as well as the results of the endogenous stratification exercise.25 Consistent with the pattern in the main crime results, treatment point estimates are considerably more negative for the groups with higher predicted baseline index values. In the Philadelphia repeated split sample estimation, the effect for the high-risk group is more than 24 times as large as for the low-risk group; for the leave-one-out estimation, the change is even starker as the low-risk coefficient flips sign. This is an ITT analysis, so the treatment effect differences capture both differences in responses and take-up rates. Interestingly, the first stage declines considerably as risk rises, from 0.41 in the low-risk group to 0.28 in the high risk group. This suggests that those who benefit the most are least likely to take-up the program on their own, and that the actual differences in responsiveness conditional on take-up are even larger than the differences in ITT point estimates would suggest.26 A similar pattern occurs in Chicago, with point estimates 13 – 26 times larger for the highest risk group than the lowest. The decline in take-up across risk groups is more muted here, a difference that may be linked to the cities’ different recruiting strategies.27

Table 4:

Treatment Effect on Combined Index by Predicted Risk Level

Panel A: Intent to Treat Effect, Index
WorkReady OSC+
Full Sample −0.065*** (0.025) −0.051** (0.021)
Panel B: WorkReady Intent to Treat Effect by Predicted Risk Level
Predicted Risk Level Repeated Split Sample Leave One Out CM First Stage
Low −0.007 (0.016) 0.018 (0.020) −0.169 0.405
Medium −0.020 (0.022) −0.078* (0.041) −0.059 0.328
High −0.169** (0.066) −0.190* (0.081) 0.207 0.277
Panel C: OSC+ Intent to Treat Effect by Predicted Risk Level
Predicted Risk Level Repeated Split Sample Leave One Out CM First Stage
Low −0.008 (0.008) −0.005 (0.011) −0.185 0.281
Medium −0.018 (0.013) −0.012 (0.019) −0.136 0.258
High −0.115** (0.054) −0.132** (0.055) 0.321 0.262

Note: WorkReady N=4497, OSC+ N=5405. Estimation from the Abadie, Chingos, and West (2018) procedure. CM is the control mean for each group, assigned in the leave one out estimation. Table shows the effect of each program on a standardized index of the outcomes that have significant changes in the main estimates: incarceration, total arrests, and receipt of child protective services (WorkReady) and drug and other arrests (OSC+). Each variable is standardized on the control group, averaged together, then re-standardized so that the standard deviation of the index is 1.

*

p<0.1,

**

p<0.05,

***

p<0.01

Yet the practical limitations of the endogenous stratification approach are clearly reflected here. Despite being quite large changes relative to the control means, the size of the standard errors makes it difficult to differentiate the groups from each other. And although it looks like treatment effects might be growing across groups, the pattern varies somewhat across the two estimation strategies, especially in Philadelphia. So it is not clear whether effects grow proportionally across groups or are just concentrated among the highest-risk group.

The lack of power to clearly understand the pattern of treatment heterogeneity is a common problem within any single study.28 And it leads to the second approach, which is to consider what we can learn from variation in risk level and heterogeneity patterns across groups in different SYEP studies. To do so requires stepping back from the individual-level variation in a given Y0 , and instead focusing on group variation across different EY0s . That is, in the spirit of meta-analysis, I compare treatment effect variation across groups that have varying average outcomes in the control group. SYEPs provide an unusual opportunity to do so, because there are now so many experimental treatment effects available across the two separate experiments in this paper, plus published main and subgroup effects from two prior OSC+ cohorts, NYC, and Boston experiments. I use 62 different significant point estimates, across all socially-costly outcomes, study locations, and subgroups, to look at how treatment effects vary across existing, measurable variation in EY0 , estimated by Y¯0 .29

This approach has limitations. By combining information across different outcome measures, this kind of synthesis makes it difficult to draw specific conclusions about mechanisms from the heterogeneity patterns. The numerical slope of the function relating Y¯0 and βY is not directly interpretable, since a one-unit increase in Y¯0 represents different kinds of increases in the prevalence or count of different socially costly outcomes.30 And in some ways, it is a weak test of the relationship between the elements of βY and Y0θ ; if there is no clear relationship, it may just be because the various outcomes making up the Y0θ vector have different risk-responsiveness relationships, such that aggregating them masks patterns in the individual outcomes. On the other hand, if there is a clear pattern in how treatment effects vary when a set of outcomes is more or less common in a group, then using so many data points could be a productive way to gain insight into the shape of the relationship between the risk of socially costly outcomes in a population and the magnitude of the change an SYEP generates in that population.

Figure 2 plots each group’s control mean on the x-axis against the corresponding LATE estimate on the y-axis, using estimates that are individually significant at the p0.1 level. One might worry about selecting estimates based on statistical significance; since with these outcomes, variances grow as control means rise, true effects would have to be bigger to reach statistical significance on the right side of the graph, were sample size and take-up rates equal. Appendix Section I.4 discusses this issue in detail and shows the result holds both when accounting for the variance in each estimate and when including all main effects regardless of statistical significance.

Figure 2:

Figure 2:

Size of LATEs Relative to Control Means

Note: Point estimates and control means taken from this paper, Davis & Heller (2020), Gelber et al. (2016), and Modestino (2019). See text and Appendix I for details.

Though the control mean in the figure is not the most relevant baseline comparison for the compliers who drive the LATE, most other SYEP studies do not report control complier means. So using the control mean rather than the control complier mean allows us to show the same set of relationships across studies, and it still captures the basic relationship between the risk level of a group and its response to treatment, inclusive of take-up decisions. Note that if take-up decisions are correlated with treatment effects, the variation in the displayed LATEs may be a function of who the compliers are in each group; it will not necessarily correspond to variation in average treatment effects by group (i.e., if everyone were forced to participate). Appendix Figure A.2 shows a very similar relationship using the ITTs from each study, indicating that the pattern in Figure 2 is not solely due to the differences in take-up rates across groups, though it may still reflect who decides to participate.

Panel A of Figure 2 starts with the significant main effects in this paper across outcomes and cities. This focuses on the variation in Y¯0 that comes from the different city populations, as well as the prevalence of the different outcomes. The panel plots each LATE point estimate against the corresponding control mean.31 The pattern is strikingly linear; larger control means are consistently associated with larger SYEP-driven declines in the outcome. This suggests that among the outcomes responsive to the program, SYEPs have a bigger effect for groups that are more likely to be at risk of those outcomes.

To further explore the robustness of this pattern, Panel B adds point estimates and control means from the full study populations in previously-published studies of SYEPs. This adds more variation in the prevalence of each outcome across independent populations. Despite differences in programming, time periods, and local context, the relationship is quite consistent. Increases in control means are roughly linearly associated with more beneficial treatment effects. The same also appears to be true for the few adverse effects (the positive green diamonds are both increases in later property crime from the initial OSC+ study), with bigger control means linearly associated with positive effects as well.

Panel C adds significant effects by subgroups across studies, reflecting variation within each outcome and study driven by a single division on one observable characteristic at a time. Appendix H presents and discusses the substantive subgroup results for WorkReady and OSC+ 2015.32 Here I focus on the overall pattern between subgroup differences in baseline rates and subgroup responsiveness.

Across subgroups that vary in their risk of these outcomes, the size of treatment effects still seems to scale proportionally with the size of the control means. There is perhaps a bit of flattening in the middle of the graph, but overall, the declines in costly outcomes clearly grow with the size of the control mean. This pattern has been seen in one-way interactions within individual studies; Boston and Chicago had significantly bigger violent-crime effects for those with prior records than those without (Davis and Heller, 2020; Modestino, 2019), and those with prior arrests have larger point estimates for arrests and convictions than those without (Kessler et al., 2021). The analysis here shows that a similar pattern holds across outcomes in the same place (Panel A of Figure 2), across outcomes in different places and times (Panel B), and across subgroups in different places and times (Panel C).

Since these estimates were selected based on statistical significance, the pattern does not mean that all youth in groups with higher prevalence of negative outcomes respond more to the treatment; there are other subgroups and outcomes with high control means where there were no significant program impacts. But it does mean that when SYEPs change outcomes, those changes are bigger when the outcomes are more prevalent. As discussed in Appendix I.4, the negative relationship between control means and treatment effects is robust to adjusting for each estimate’s variance and to including null main effects.

There is a mirror image of the pattern for the few adverse effects as well; youth in groups where the outcome is more common have a bigger adverse response as well. These outcomes are generally much less socially costly than the outcomes that are falling (property and drug crimes sometimes increase, and it is not clear whether the increase in male fertility is from sexual behavior or increased willingness for fathers to be listed on the birth certificate, see discussion in Appendix H). So while it is worth considering how careful SYEP targeting and program adjustments may help to minimize those increases, the overall declines in outcomes like mortality and violent crime are likely dominate any cost-benefit calculation.

7.3. Interpreting effect heterogeneity by risk level

Overall, the data seem strikingly consistent with Panel B of Figure 1, suggesting we might extrapolate the absolute magnitude of SYEPs’ effects in new settings as a proportional function of the anticipated control mean. The consistency of this relationship across outcomes could have a number of explanations. First, it could be that a similar mechanism is driving SYEPs’ behavioral effects across all these outcomes – a range of crime types, measures of individual behavioral health and mortality, and measures of family stability. Income, changes in beliefs about the future, shifts in time use, or the development of social and self-regulation skills could be similar inputs into the production of these outcomes, generating this kind of proportional shift in multiple outcomes.

But there are also alternative explanations. Suppose, for example, that the key behavioral mechanism stems from the interactions between youth and adult program providers, and providers allocate time and attention to youth who are struggling the most. Or suppose there are diminishing marginal returns to adult interaction, such that youth with lower Y0s have already benefitted from other adult attention and thus have lower benefits from program-driven investment. Either explanation could generate the pattern of results, and both emphasize that this relationship is descriptive, not causal. Anything correlated with higher Y0s , including program design and implementation details that differ for higher-risk groups, could be driving the relationship.

Note it is also possible that the overall prevalence of these outcomes is low enough in all these groups that a floor effect is still binding. Even in the point farthest to the right of the graph — the decline in arrests for other crime among those with a prior arrest in the OSC+ 2015 sample — there are 30 arrests per 100 youth in the control group. Even if each of those arrests belonged to a different person, that still leaves 70 control youth for whom the program can not move other-crime arrests below 0. So even at the extreme of the data, it is possible that the graph reflects the sloped portion of Panel A in Figure 1.

Regardless of the reason, the linear relationship between risk and responsiveness holds a useful lesson for targeting SYEPs. Across all experimental studies of these programs, groups with higher baseline rates of affected outcomes have larger beneficial program effects. There does not seem to be a margin past which youth respond less. To the extent policymakers want to generate declines in the kinds of outcomes measured here, finding ways to recruit and retain SYEP participants at elevated risk of any of the focal outcomes—ideally while finding ways to minimize any adverse effects—is likely to maximize the net social benefits from the program (though it may also require additional costs to serve more disconnected populations). The risk-response relationship is also good news for policymakers concerned with equity. Since those at the highest risk of harmful outcomes seem to benefit the most, targeting the program to have the biggest impact is equivalent to serving the population that would otherwise be the most disadvantaged.

One concern about this targeting strategy would be if program composition plays a key role in behavior change. In a world where peer interactions are a key input into program effects, a targeting strategy that dramatically alters to whom youth are exposed during the program could change the program’s impact. It is perhaps informative, though, that the studies included in Panel B of Figure 2 vary quite a bit in the composition of peer groups within the program. For example, OSC+ 2013 purposefully focused on recruiting a large number of youth at elevated risk of criminal justice involvement, with almost half entering the program with an arrest record. In NYC, on the other hand, only about 3 percent had been arrested at baseline. So at least within the variants of SYEPs that have been experimentally evaluated, being careful not to extrapolate too far out of sample, the results here suggest that targeting populations at higher risk of bad outcomes will increase program benefits. The more likely applicants are to engage in the kinds of risky behavior the programs reduce, the bigger the social benefits from reducing those outcomes are likely to be.

Policymakers may have multiple goals when deciding whom to target with SYEPs, some of which are better served by enrolling youth at lower risk of socially costly outcomes. Providing widespread income transfers, for example, or developing the kinds of skills and connections that help in the labor force may imply different targeting goals (e.g., Davis and Heller (2020) find evidence that employment effects are larger for younger youth more attached to school and less involved in the criminal justice system33). But given the high social costs of the type of program effects documented in this paper, increasing effects on these outcomes should help SYEPs generate benefits that exceed program costs. And since the subgroups that seem most responsive are also marginalized in other ways, such targeting may help advance social justice and equity concerns as well.

8. Conclusion

Variation in what treatment looks like across scales and setting—program structure, staff training or experience, counterfactual opportunities, and so forth—as well as variation in who participates can sometimes dramatically change an intervention’s effects across settings. Understanding that variation should be an important input into decisions about public investments, since it is crucial to predicting whether an expanded investment in one place is likely to replicate the success of any given intervention strategy. This paper assesses and unpacks scale and replicability for one promising type of intervention, SYEPs. These programs consistently reduce criminal justice involvement in the first year after random assignment, and may have some lasting effects as well. They may also help reduce the need for child protective and behavioral health services, although these results are less precise and concentrated among some subgroups (see Appendix C).

The insensitivity of the decline in criminal justice involvement across time, location, and program implementation suggests that the basic structure of the program is more important than the details. Although treatment heterogeneity does not seem related to program structure or delivery, it is related to participant risk level. When SYEPs improve socially costly outcomes, they do so more for youth at higher risk of those outcomes. This is relevant for scaling: If programs get so big that the risk level of the population served drops, the program effects are likely to persist but get smaller in absolute magnitude, as in Philadelphia. But when programs are not universal and make purposeful targeting choices, growing while remaining smaller than the population of those who could feasibly benefit, scaling up without major differences in youth populations is feasible, as seen in OSC+.

The heterogeneity results also suggest something important about the structure of the underlying behavioral response—that at least in the contexts that have been tested so far, there is no such thing as “too late” to generate change. It is important not to generalize too far out of sample; for example, it seems unlikely that a 6–8 week program would do much to reduce severe gun violence for those at extremely elevated risk of shooting involvement. But for the outcomes that respond to treatment among populations where SYEPs have been tested, there does not appear to be a margin past which youth fail to respond; rather, making eligibility and targeting decisions that encourage youth at higher risk of crime, family instability, and health problems to participate is likely to generate bigger social gains.

There are limitations to this targeting recommendation. It is possible that massive shifts in program populations, beyond what has already been tried, could change peer exposure in a way that diminishes program impacts. There are also a range of other issues policymakers need to consider. For example, the increased costs of serving more disconnected youth could get high enough to outweigh the increased benefits. At the same time, the benefits of serving youth at very low risk of the outcomes measured here could be low enough that they do not justify program costs; that depends on the size of other benefits not captured by the SYEP studies. Alongside the lesson from this paper that the magnitude of benefits is likely to grow with the prevalence of the outcome in a particular group, detailed consideration of the different social costs across outcomes should inform final decisions about targeting.

Other limitations of the analyses here generate directions for future work. The compliance issues significantly limited statistical power, such that further research on how different program elements matter and how family and health outcomes respond would be valuable. All of the major experiments on SYEPs have occurred in large cities, where the programs are quite widespread. But as Ross and Kazis (2016) point out, these results may not generalize to smaller cities or rural areas that lack the infrastructure for program administration. A better understanding of how the basic program approach changes when implemented outside large cities would help assess the potential for broader replicability.

Despite their limitations, the overall message of the experiments reported here is fairly optimistic. The evidence suggests that SYEPs are not just promising in a way that is “ready for scale,” but that they are actually scalable—at least up to the point where there are too few people at high enough risk of crime and violence to generate social benefits that outweigh program costs. For policymakers who wish to reduce social inequality efficiently, identifying other approaches that both replicate across contexts and generate the biggest benefits for the people facing the most challenges should be a priority. More broadly, the paper demonstrates how studying treatment variation and individual heterogeneity across multiple contexts has the potential to inform public spending decisions about future investments more effectively than just establishing which novel interventions “work.”

9. Acknowledgements

This project was supported by Award No. 2016-R2-CX-0049, awarded by the National Institute of Justice, Office of Justice Programs, U.S. Department of Justice, a State and Local Innovation Initiative grant (Philadelphia) and Social Policy Research Initiative grant (Chicago) from J-PAL North America, the Robert R. McCormick Foundation, and Project Development Grant Program funding from Poverty Solutions at the University of Michigan. Louise Geraghty, Brenda Mathias, Matt Repka, Misuzu Schexnider, and Lauren Shaw provided highly-accomplished project management; Kalen Flynn managed the Philadelphia qualitative data collection and analysis; Raquel Chavez, Kenny Hofmeister, Angela Hsu, Owen McCarthy, and Mary Clair Turner provided excellent research assistance; Marianne Bertrand provided invaluable support to the Chicago experiment, as did Greg Ridgeway for the Philadelphia study. The author thanks Jon Davis, Brian Jacob, Michael Ricks, and Basit Zafar for extremely helpful comments. I am grateful to the Chicago Department of Family and Support Services, the Philadelphia Youth Network, Inc., the Philadelphia mayor’s office, and the University of Chicago Urban Labs for their partnership on these projects. I also thank the City of Philadelphia, the Philadelphia Police Department, the School District of Philadelphia, the Chicago Police Department, and the Chicago Public Schools for graciously allowing the use of their administrative data. Any further use of the data is subject to approval of each agency. The opinions, findings, and conclusions or recommendations expressed in this publication are those of the author and do not necessarily reflect those of these organizations. The studies are registered in the American Economic Association Registry under trial numbers 2451 (WorkReady) and 805 (OSC+).

A. Study Design

A.1. One Summer Chicago Plus 2015

OSC+ has been studied in two previous cohorts, 2012 and 2013. The program itself has changed somewhat over time. In the initial 2012 cohort, 14- to 21-year-old high school students spent a total of 8 weeks working in a mix of non-profit and government jobs after a 1-day training. They worked up to 15 hours a week and participated in a cognitive behavioral therapy-based social-emotional learning (SEL) curriculum for another 10 hours per week. Everyone was assigned an adult mentor to provide support and help them deal with barriers to employment throughout the summer. The second year of the program (2013), the program was only open to boys 16- to 22-years old. Half were recruited from high-violence high schools, and half were recruited through agencies connected to the criminal justice system. The program was shortened to 6 weeks, and private sector jobs became part of the program. Several of the shifts in the 2015 program are described in the main text: the population expanded; recruiting was again school-based with all 16- to 21-year-old students (including girls) in those schools eligible to apply; the SEL curriculum changed to a civics curriculum and was shortened; and the wages dropped relative to the labor market.

As in prior years, 2015 youth were placed in non-profit, government, and private sector jobs, as well as participating in personal development programming, for 25 hours per week. They received a full week of work readiness training at the beginning of the program rather than just a day. And the program obtained a waiver against the city’s new minimum wage increase, so youth were still paid $8.25 per hour (compared to a concurrent shift to a $10 minimum wage). The largest substantive difference in the 2015 program was the removal of the original social-emotional learning curriculum. Instead, program operators tested two versions of the program to isolate the effects of some of the more expensive program elements. Half the program youth had the opportunity to work in their jobs for the full 25 hours per week, with no separate adult mentor and no additional youth development curriculum. The other half worked for 20 hours per week, and on Fridays engaged in a Civic Leadership Foundation curriculum focused on civic leadership. The initial purpose was to assess whether the additional costs associated with adult mentors and professional development were a key mechanism driving the previously-documented effects, though as discussed in the main text, the reality of implementation makes this difficult to answer with certainty.34

In spring 2015, OSC+ providers advertised the program and solicited applications from youth at 49 high schools in high-violence community areas. Initially, 5,444 applications were collected. I grouped applicants into geographic strata based on the address listed on their application, because one of the main concerns about prior lotteries had been that some youth were required to make long commutes to work. I then randomly ordered applicants within strata and assigned them to both a provider serving that area and one of three treatment groups: job only, job + mentor (including a Civic Leadership Foundation curriculum delivered on Fridays), and a control group.

Through the process of matching these applications to administrative data, the research team identified 39 cases where the same youth submitted more than one application, with identifying information different enough that we did not catch the second application in our initial de-duplication process. I collapse duplicate applications into a single observation, where treatment is defined as the maximum of all random assignment indicators (since the max of two random variables is also random). To ensure the analysis accounts for the higher probability of treatment among these applicants, I include dummy variables for having submitted two applications in all analyses.

A.2. WorkReady Experimental Design 2017

To fill in a few more details relative to the program description in the main text: For over 15 years, the Philadelphia Youth Network (PYN) has offered a summer jobs program for youth ages 14 to 21 called WorkReady. During the study years, using a blend of government and private funding, PYN contracted with 50–60 local agencies to implement the six-week WorkReady summer program.35 Youth accepted to the program were assigned to a local provider, which placed them in one of three program models: service learning to address a community problem, work experience with skill development and ongoing adult interaction, or an internship that included professional development and less intensive adult mentoring. All three models focused on developing “21st-Century Workforce Skills” and offered an hourly wage, but they varied in how like a private-sector summer job they were. Youth were not randomly assigned to the different program models, so the effects estimated here are the average of the models when providers match participants to the model based on their existing work readiness.

During the study years, WorkReady’s36 3 program models were:

  1. Service learning : Youth worked in groups to address a community problem. They researched the problem, developed solutions, and provided direct service and advocacy to address the problem. This model was designed for youth with minimal or no prior work experience.

  2. Work experience: Youth participated in a structured work experience, with an emphasis on skill development and ongoing adult interaction. Youth were placed in a professional work environment appropriate for those with little or no work experience. They received training in workplace skills prior to the start of their job and completed work-based learning projects.

  3. Internship: Youth who were already prepared for the workplace were placed in summer jobs, with a focus on placements that would not otherwise be accessible to young people. They also participated in professional development workshops and interacted with a trained adult supervisor over the course of the summer.

Typically, providers assigned youth to a program model based on their age and experience. Providers were all required to offer professional development sessions, but they varied in focus, content, and structure across agencies. Some sites integrated professional development throughout the work days, while others used Friday as a mandatory professional development day. Topics varied widely, ranging from developing business models to sexual health education.

Youth interested in WorkReady applied in one of two ways: through a program provider or through PYN’s online system without any attachment to a particular provider. Youth applying through providers have often already been screened to meet a given provider’s program requirements, which vary across providers. Youth applying without a connection to a particular provider entered a group called the “general pool.” In the absence of a lottery, PYN tries to place these general pool youth in positions that open up as providers request more youth to fill their slots. Because this group is often more disadvantaged and less likely to get a program slot, PYN wanted to prioritize increasing access among this group. At the same time, many providers have strong preferences about whom they are willing to serve (sometimes driven by requirements of the jobs, e.g., passing criminal background checks), and so do not accept youth they have not screened themselves. To achieve the dual goals of expanding access to the program to youth who might not otherwise be served and allowing providers to screen youth according to their needs, I allowed providers to choose one of the two types of lotteries described below. A handful of providers were exempted from the lottery by preference or for logistical reasons, but 59 providers participated. I conducted two waves of randomization. The first occurred on May 1, 2017, and the second occurred on May 12, 2017.

Over-Recruitment

In the over-recruitment lotteries, individual program providers collected applications from more youth than they could serve. They were able to pre-screen applicants to meet their own operational requirements, but were asked to collect more applications of acceptable candidates than they had program slots. Part of the rationale from their perspective was that the City of Philadelphia was trying to scale up the program, so reaching out to new youth would inform these youth about the program and establish relationships for the future, even if they did not win a program slot in this program year.

Providers were given targets for the number of applications to submit relative to the number of available slots, but the exact number of applicants relative to those targets varied somewhat by provider. They were able to choose which youth to place directly in their non-lotteried slots, and which to submit to be randomized for the remaining open slots. Providers submitted their list of applicants to PYN, and the researchers randomized within provider. Since there are a number of steps, including paperwork completion, that generate drop-off between the program offer and actually participating, we aimed to randomly select about 120 percent times the number of available slots within each provider as treatment youth. The remainder were assigned to the control group. As a result, treatment probabilities varied across providers. Overall, 1,554 applicants took part in an over-recruitment lottery across 39 providers, with 765 assigned to treatment.

General Pool

Some program providers did not need to prescreen applicants and were willing to accept youth unknown to them who applied through the general pool. To ensure that program providers could logistically serve the general pool youth, providers submitted general location and age preferences (region of the city by zip code and whether they could serve 14- to 16-year olds, 17- to 18-year olds, or both). I individually randomly assigned general pool applicants within these age and location strata to providers based on how many program slots those providers reported needing to fill. As in the over-recruitment lottery, I randomized about 120 percent more applicants to treatment than requested slots to ensure that all slots could be filled, even when youth were hard to find, failed to complete paperwork, or were no longer interested. Within age and geography strata, provider assignment was random. Although the providers who were in these strata are not necessarily representative of all providers, this does induce some random variation in what type of program youth were assigned. Overall, 1,838 general pool applicants were randomly assigned across 20 providers, 571 to treatment and 1267 to control.

A.3. WorkReady Experimental Design 2018

To simplify the experimental design in 2018, there was only one lottery consisting of youth who applied to the general pool. PYN learned which providers were willing to accept general pool youth and assessed how many slots these providers would need to fill after having submitted all their paperwork on the youth they recruited directly. PYN also worked to collect all the necessary paperwork from the general pool youth, so that the youth would be ready to place (therefore reducing barriers to providers serving treatment youth).

As of the beginning of June, there were 1,105 general pool applications that had not been placed in other jobs, and providers anticipated needing an additional 344 youth to fill their program slots. Because we knew from prior years that not all treatment youth would participate, I randomly assigned 450 youth from the 1,105 applicants to be in the treatment group. PYN then worked to place these treatment youth at about 40 different providers according to provider needs. As a result, in this cohort, there was no random variation in provider.

B. Data

All data come from probabilistic matches to government agency data, accessed by legal agreements with the agencies. Researchers performed the match to Chicago Police Department (CPD) and Chicago Public Schools (CPS) data, and staff at each agency performend the match at the Philadelphia Police Department and School District of Philadelphia.

For the researcher-performed match in Chicago, the research team used information on name, date of birth, gender, race, and home address across program records, CPD, and CPS records to identify which rows of each database belong to the same person. Each individual database includes an individual identifier that links records on the same person, but the identifier can include some error (e.g., if a student transferred schools and was assigned a new student ID number rather than correctly transferring the existing number). The probabilistic algorithm allows for typographical or categorical errors (e.g., mismatches on race or gender) both within data sources and when linking across databases.

To minimize false positive links, we impose logical constraints during the record linkage process. If a potential link between two records would violate a prespecified constraint, this link is not permitted. Specifically, we do not link records if the resulting individual would contain school records such that the total number of attendance days in a given year exceeds 200 days, or such that a student would have appeared to take the same standardized test twice within the same year.

Because treatment may affect how many records an individual has during the post-randomizaton period, it is theoretically possible that treatment could be correlated with how likely observations are to be linked. For example, if the program changes the probability of arrests, we are more likely to observe arrest data for individuals in one group compared to the other, which means we potentially have more identifying information about individuals in that group. If we run the record linkage algorithm with pre- and post-randomization data combined, it is possible that a pre-randomization arrest record may only be linked to a person through a post-randomization arrest record.

To avoid injecting treatment effects into linkage probabilities, we run an initial match using only pre-randomization records. After the pre-randomization record linkage is complete, we then re-run the record linkage process with post-randomization school and arrest records included. Records linked in the first step cannot be broken up; that is, post-randomization information can not change the decision about whether pre-randomization records belong to a given study individual. But new records can be added to the existing matched individual, and non-matched individuals can be newly matched to their post-randomization data. This approach helps ensure that treatment does not influence the probability that given records are correctly linked to the right person.

B.1. Chicago

Chicago data were used via data agreements with the University of Chicago Urban Labs, where all data are stored and analyzed. Variables were defined in a similar way to prior work on OSC+, so the data appendix in Davis and Sara B. Heller (2020) largely applies here as well. The primary data sources for Chicago are school records from Chicago Public Schools (CPS), arrest records from Chicago Police Department (CPD), and the OSC+ study participant file. Records in all files include information about name, date of birth, gender, race, and home address.

In the Chicago Public Schools data, I observe historical enrollment records, including reasons for leaving, that capture everyone who ever enrolled in the public school district since the early 2000s. But because our data agreement with the School District of Philadelphia for this research only included records beginning in 2016–17, I define baseline measures from the pre-randomization school year only, which I observe for everyone in both studies. I have CPS data through the 2016–17 academic year, or 2 post-program school years. These enrollment records can be easily deterministically linked via a student identifier to other schooling data, such as grades, attendance, and standardized test scores. Of the initial 5,444 application records (prior to deduplication, see above), 99.5 percent matched to CPS data. The remaining 26 applications may not have matched because they were enrolled in private or parochial schools or schools outside of Chicago, or because the fields used to match were not sufficiently similar between the OSC+ and Chicago Public School data. I exclude these youth from the main education results, since I have no information on their school outcomes. They receive different imputations below in the section that assesses how missing data affects schooling results.

In the Chicago Police Department data, I observe historical arrest records, including the statute under which the arrest was made, that capture everyone who has been arrested by Chicago Police Department since 1999. Anyone not matched to the arrest data receives a 0 for all arrest counts. Crime categories are the same as in my prior work on OSC+.

B.2. Philadelphia

B.2.1. CARES Data

The City of Philadelphia’s Data Management Office maintains an integrated data system (called CARES) that collects and integrates service records across multiple city service agencies. They received our personal identifiers from PYN application information, including name, date of birth, and gender. The Data Management Office then used probablistic matching to find each youth in City service data; the Philadelphia Police Department also provided them with full arrest databases to which they could match our data. In practice, the threshold the City used as a cutoff for probablistic matching was high enough that when using only name, date of birth, and gender as matching variables, the match had to be exact to exceed the threshold. However, the process did still allow for typographical errors in records in two ways. First, they match their own records across agencies using more identifiers than I had available in my data. If two records had similar addresses and parent information but different spelling variants or DOB typos, for example, they could still match to each other. Our WorkReady record would then only have to match to one of the records to be considered a match; the record with other variants would still be linked to our study individual.

Second, as described below, the City sent identifiers and their own study indentifying numbers to the School District, so that I could merge education data back into City data despite the records being stripped of identifiers. The School District used a different probabilistic algorithm that caught some additional matches. Before the City pulled the final data, they used the School District match to reconcile the conflicting match cases. So in practice, our identifiers could match to City data despite not being exact matches on name, date of birth, and gender. As with any matching, it is possible that the matching process for these outcome measures missed some name or date of birth variants and therefore understate true service receipt (since I assign 0s for those not matched to the data). Importantly, I used personal identifiers that they youth provided on the application, prior to random assignment. This means that there is no reason to think that typographical errors or nicknames preventing exact matching would be imbalanced across treatment and control groups.

The City data team stripped the identifying information off of the data and returned it to us identified only by a study ID number. I was required to securely delete all our application records prior to receiving the data, so that there was no way to re-identify the records. The City is keeping our initial match file as a record of study participants. They also provided a separate file with HIPAA-covered records (mental health and substance abuse services). For these matches, I provided a file with identifiers attached to a smaller number of covariates, so that no individual was uniquely identified by any combination of covariates. These data have different study identifying numbers and so can not be merged with other records.

To define the non-arrest criminal justice outcomes in the main text, I use information from CARES. I code indicators for whether someone has been incarcerated in either juvenile detention or juvenile prison; I do not observe adult incarceration, so this variable is 0 for incidents occuring over the age of majority (18) or those charged as an adult. I also measure whether someone received any court-ordered services within the Department of Human Services (DHS), which is a less severe punishment than incarceration but can be offered to incarcerated youth as well. This variable understates the amount of total court-ordered services, since some services are provided outside of DHS; however, there is no reason to believe that treatment should affect which provider a youth is assigned.

The CARES database provides information on a range of other outcomes are substantively important but less often available in individual-level administrative data. Using service records from the Child and Youth Division, I create an indicator for whether a study member’s family received any services from the city’s child protection agency. This includes any service in response to a substantiated call, which is one that the Division determines merits further involvement. It counts residential placements like foster care or kinship care, as well as non-placement services such as safety checks and other in-home services.

As with any service data, a change in services has multiple interpretations. It could represent a change in the number of incidents requiring child protective services; it could also reflect more or fewer incidents being called in or reported. Given that WorkReady involves more adult interaction with youth, it seems likely that conditional on a situation requiring intervention, treatment would increase the probability the incident is reported to the Child and Youth Division. If reporting increases for treatment youth, estimated program impacts on child endangerment would be understated.

To measure behavioral health, I have records of all substance abuse and mental health services that are either covered under Medicaid or funded in part by Philadelphia County. I generate indicators for receiving any of these behavioral health services, as well as separately for mental health and substance abuse treatment. As with child protection, program effects on this measure could indicate changes in the underlying issue or changes in willingness to seek out treatment. Since the most likely treatment effect of interacting with additional caring adults would be to encourage the identification of personal issues and willingness to seek help, I expect that this measure may also understate any treatment-driven declines in the true underlying health conditions.

Lastly, I create indicators for whether a youth ever used a city homeless shelter and whether they had a child. Fertility is measured as being listed on a birth certificate in Pennsylvania vital statistics records, so there is more scope for missing data and treatment effects on reporting for fathers than mothers. All other measures in the paper are separated by year since random assignment. But to protect confidentiality, the City only provided information on whether births were pre- or post-random assignment. As a result, I do not separate the parenthood outcome by year; it is measured across all observed post-randomization years.

Together, the social services data provides a much more complete picture of youth and family welfare than has been available in prior summer jobs studies. The data capture, for the first time in administrative data using experimental methods, the programs’ influence on family income, stress, and stability; self-efficacy and mental health; and risky behaviors like substance use or unprotected sex.

B.2.2. School District Data

The School District of Philadelphia (SDP) received identifiers from the City and used their own probabilistic matching algorithm to link the records to school data. SDP has enrollment and graduation records for every public school student, but charter schools do not typically report grades, absences, or suspensions to the District office. Additionally, their enrollment file only records enrollment status in the current academic school year (unlike Chicago where enrollment records are cumulative). Because the data agreement for this project only covered data starting in the 2016–17 academic year, the earliest I can observe whether a study youth is enrolled, dropped out, transferred, or graduated, is that academic year (the year prior to randomization for the 2017 cohort and 2 years prior for the 2018 cohort).

I use the pre-randomization records to identify youth who graduated prior to the program (n = 286). Since they can not have changes in education outcomes by construction, I exclude them from all education analysis. Another 353 youth did not match to SDP data at all and so are missing all school data. These youth may not have matched because they were enrolled in private or parochial schools or schools outside of Philadelphia. It is also possible that some of the youth were enrolled in SDP, but with identifiers different enough from what they used on the WorkReady application that the matching process missed them. I exclude these youth from the main school persistence outcome, since I do not have information on whether they are enrolled or graduated at any given time, but they are included in the imputation robustness checks in Appendix Section F.4.

Another set of youth do appear in the SDP data, but are missing individual variables or years of data for multiple reasons. Charter school students do not typically have grades, absences, or suspensions, although some charters report those variables in an irregular and unreliable way. I therefore set those variables to missing for charter students.37 Students who drop out, transfer, or graduate will be missing data for some years, as will those who become incarcerated, deceased, or otherwise incapacitated. The fact that there is so much missing data, and the fact that the reason for missingness varies – sometimes indicating school success, as in the case of graduation, sometimes indicating school failure, as in the case of dropout, and sometimes indicating a neutral reason like transfer – makes it difficult to draw strong conclusions from educational outcomes.

Because of the missing data, I focus on the outcome I can measure for everyone other than the 353 non-matches and 286 prior graduates: whether a student is still in school or has graduated by a given year. Section F.4 below presents those results, as well as results for other education outcomes under different imputation schemes for the missing data, including the extreme assumptions that all missing data is an indication of either school success or school failure.

C. Main Effects for School Persistence, Family, and Health Outcomes

C.1. Education

As discussed above, the main measure of education is school persistence, defined as either remaining enrolled in school or having graduated. Table A1 shows that there is no significant change in this measure, consistent with previous OSC+ studies. Point estimates are small, ruling out more than a 2 percent increase for those offered the program in either city. Because of the amount of missing data on other education outcomes, discussed above, it is more difficult to estimate effects on other education outcomes with confidence (see Appendix F.4 below for results and discussion).

C.2. Other Family and Health Outcomes

SYEPs provide income, which could reduce family or individual stress. The extra $1–2,000 could improve living conditions at home; in surveys of Chicago participants, almost 80 percent of net wages went to either local businesses or participants’ families (MHA Labs, 2015). SYEPs also aim to develop personal skills and self-efficacy, which could directly shape mental and behavioral health, as well as risky behaviors. Yet prior work on the behavioral effects of these programs has largely been limited to crime and education outcomes.

Table A1 shows WorkReady’s impact on other family and health outcomes in the first year after randomization. There are some potentially promising results. Among those offered the program, there is a marginally significant (p = 0.09) 0.5 percentage point, or 33 percent, decline in the receipt of child protective services, which include in-home safety checks, other family services, and removal of children to residential placements. The LATE is a 1.6 percentage point decline among compliers, which is basically a 100 percent decline relative to the CCM. There is reason for caution about the strength of these results; they are proportionally huge changes, but from a very low baseline, and they do not survive adjustments for multiple testing within this family of outcomes (see Appendix F.3). Nonetheless, there is more precision in some subgroups (including for African-American youth, see Appendix H) and when using a probit, which may handle the low base rate somewhat more effectively than the linear probability model, though also adjusts for fewer covariates to ensure convergence (average marginal ITT effect of 1 percentage point, a 45 percent decline, p = 0.028, see Appendix F.2). Because intervention due to a substantiated call about concerns over child safety is an extreme and costly outcome, even suggestive evidence that WorkReady reduces these services merits attention.

Similarly, although the point estimate on the receipt of behavioral health services is not statistically significant (p = 0.207), it is proportionally quite large. Youth who participated in WorkReady were 3.3 percentage points, or 31 percent less likely to receive these services. As shown in the bottom two rows, most of these services are mental health related, with a little under 1 percent of them addressing substance abuse issues. This is certainly not strong enough evidence to conclude that SYEPs reduce behavioral health problems with confidence. But as discussed above, program effects may be understated; to the extent that WorkReady increases the probability of detecting and reporting family or health issues, service receipt might increase conditional on the underlying issue for treatment youth, attenuating estimated program effects. Some of the subgroup results in the appendix also point toward the possibility that SYEPs may matter for family and health outcomes, especially for subgroups like boys and African-American youth who are at elevated risk of the outcomes. The direction and magnitude of these effects for outcomes that are central to youth well-being should be a priority for future study.

D. The Role of Incapacitation

In multiple prior studies of SYEPs, researchers have shown that the decline in crime is not merely a mechanical result of incapacitation, or keeping youth busy over the summer (Sara B. Heller, 2014; Modestino, 2019; Judd B. Kessler et al., 2021). They show this by documenting the time path of the cumulative treatment effect on crime, month by month. If crime declines were purely a result of keeping youth occupied during the summer, we would expect to see the treatment effect accrue entirely during the summer months, then level off and stop growing as time passed. In fact, prior studies have shown that the treatment effect continues to grow after the summer is over, for at least the year following randomization.

The potential role of incapacitation is also relevant for the heterogeneity discussed in this paper. In theory, one explanation for the bigger program effects among those at higher risk of negative outcomes is that there is more crime to prevent during the summer among those groups. In a world where effects only worked through incapacitation, we should expect to see bigger effects among those who would have committed more crime over the summer months.

Figure A.1 suggests that incapacitation is unlikely to explain the heterogeneity in the paper. It plots the time path of the ITT on total arrests for both OSC+ and WorkReady, with 36 months of data available in Chicago and 16 months available for both cohorts in Philadelphia.38 As in prior SYEP studies, the treatment effect continues to accrue after the summer of the program. Panel A shows that in Chicago, a very small portion of the treatment effect accrues over the summer (to the left of the dotted line). The cumulative effect continues to grow at roughly the same rate over the next 36 months, with little sign of a slow down over that time. Like the main effects reported in the paper, the decline in crime is a bit noisy. But the pattern of a continued decline is not consistent with an incapacitation story, since the program only occurred during the first 3 months.

The Philadelphia figure (Panel B) aggregates over young people who were randomized in waves at different times. This means that the end of the program occurs at different points relative to random assignment; the two vertical lines show the range of program end dates relative to randomization.39 Unlike in Chicago, the crime drop during the program months is distinguishable from zero, which could be consistent with an incapacitation effect. But the treatment effect continues to grow from the end of the program through about month 10, more than doubling in size over that time. It then seems to shrink a bit during the time period that corresponds to the following summer, with the downward slope returning after the 1-year mark.

In both cases, it seems clear that youth participants are taking something from the program that changes their future behavior, even after the program ends. This suggests that the pattern of bigger effects among those at the highest risk of negative outcomes is not simply a mechanical function of keeping different types of youth busy over the summer.

E. Variation across program structure and delivery

Because variation in what the treatment actually is can contribute to difficulty with scale and replication, the research design included experimental variation in both program model and local provider. In practice, in part because of non-compliance, both tests are under-powered. So the main text reports these results, but also highlights what we can learn from across-study variation in program structure and delivery. This section reports more details of the two within-study tests.

In Chicago, I randomly varied the structure of the program across two treatment arms — a job only versus a job, mentor, and civics curriculum. Table A2 shows no significant differences by arm, although the pattern is suggestive: Only the job + mentor group has crime declines that can be differentiated from 0.40 The point estimates on arrests for the mentored group are generally about twice as large as in the job-only group, providing some hint that the additional program elements do increase the program’s impact. But there is not enough statistical power to differentiate the two groups.

To give a concrete sense of this test’s power, I impose the null of equal treatment effects across arms by re-randomizing which arm a treatment youth is assigned and calculate the standard deviation of group differences under that null. I find that the standard deviation for the total arrest ITT difference is 0.021, meaning that the difference between arms would have to be at least 0.041 for the test to reject the null. Substantively that is about 23 percent of the control mean, which is perhaps a plausible difference. But since the overall ITT effect is only −0.023, it would be equivalent to one treatment arm having twice as big an effect as was observed on average, and the other doing nothing. Since substantively important differences could be much smaller than 0.041, we likely need further research to establish whether the changes in mentorship (dedicated mentor versus adult supervisor) and changes in enrichment in curriculum matter.

Random assignment to provider facilitates a different kind of test of how much program structure and delivery matter. As described in the main text, in both Chicago and part of Philadelphia’s 2017 lottery, the experimental design means provider is uncorrelated with youth characteristics and treatment probability, conditional on randomization strata. I test for cross-provider heterogeneity by including random treatment coefficients (i.e., random slopes) for each provider, then testing whether the random effects explain significantly more variation than a single treatment indicator. There is no detectable variation in treatment effects across providers. Across both studies, for only 1 of the 17 main year 1 outcomes can I reject the null that the random effects have no additional explanatory power relative to the single fixed treatment effect.41 It is possible that the strata fixed effects, which capture geographic variation in both cities and age variation in Philadelphia, absorb some of the differences across providers, or that the 19–20 providers in each city’s test are not enough to generate detectable variation.

To assess the power of this test, I generate the randomization distribution of the standard deviation of (random) treatment effects across providers. This is a different test statistic than my main likelihood ratio test, but it is one with a somewhat more interpretable meaning: how much treatment effects vary across providers. For total year 1 arrests, I find that I could statistically rule out that the standard deviation of treatment effects across sites is more than 0.0361 in Chicago and 0.013 in Philadelphia. In Chicago, this would generate a 95 percent confidence interval ruling out a SD of cross-provider treatment effects about 3 times as large as the mean ITT effect, and in Philadelphia it would rule out a SD about 2.5 times as large as the mean effect. In other words, there could be substantively important variation across providers that these tests could not reject.

As such, the more informative argument about how the details of job placement, program staff, organizational structure, or program elements matter likely comes from the cross-study comparisons discussed in the main text.

F. Robustness Checks

F.1. Covariates

In the main analysis, baseline covariates are included in all the regressions. For non-HIPPA data in Philadelphia, covariates include dummy variables indicating age bins, gender, race/ethnicity, any receipt of baseline services by type, number of baseline arrests by type, and prior grade level. For HIPAA data, the baseline covariates are gender, an indicator for being Black, age bins, and indicators for prior receipt of HIPAA-covered services. In Chicago, I include indicator variables for male, race, age bins, number of baseline arrests by type, GPA categories, number of absence bins, and duplicate randomization. To reduce any potential finite-sample misspecification, I have included baseline covariates as dummy variables. Missing baseline covariates are imputed as 0, with missing data indicators included.

Tables A3 and A4 show all the main results without baseline covariates, other than the strata and duplicate indicators needed to ensure treatment is conditionally random. None of the substantive conclusions change relative to the regressions including baseline covariates in the main text. Because there is a slight imbalance favoring the treatment group on a couple baseline variables in Philadelphia, results excluding covariates tend to be larger and more statistically significant there. In Chicago, some results that are right on the border of standard significance cutoffs cross the cutoff without covariates, most becoming more significant. But overall, the findings of both studies are robust to excluding baseline covariates.

F.2. Alternative Functional Form

The dependent variables in the paper are generally either indicators or counts. Table A5 reports the results of either probit or Poisson regression (with robust standard errors to relax the assumption that mean equals variance) to ensure findings are not sensitive to the functional form of the regression. I note that for WorkReady, no observations had more than 1 drug or other arrest during year 1, so those results are from a probit rather than Poisson regression. All coefficients are reported as average marginal effects.

To ensure convergence, the non-linear regressions are run with a limited set of baseline covariates. For Philadelphia non-HIPAA outcomes and Chicago, this includes indicators for Black, male, and whether someone had a baseline arrest. For Philadelphia HIPAA outcomes, it includes indicators for Black, male, and whether someone had any behavioral health service prior to baseline.

Given the rarity of most outcomes in Philadelphia, many randomization strata are dropped due to lack of variation in the dependent variable within stratum. As such, the tables report the N used in the regression, as well as the control mean specific to the sample contributing to identification. The overall pattern of results shows extremely similar changes to the main results. In Philadelphia, the proportional changes are slightly bigger than the linear probability models in the main text. This may be because probit handles the low base rates more effectively, or because the sample only includes strata with crime variation in them. It may also be because the regressions include fewer baseline covariates to ensure convergence, and covariates help to control for the small amount of baseline imbalance. In Chicago, the main difference is that results are slightly less precise due to the inclusion of fewer baseline covariates.

F.3. Randomization Inference and Multiple Testing Adjustments

At the time of random assignment for WorkReady, I set a balance check rule of re-randomizing until each individual covariate was balanced at the p>0.1 level and the joint test of balance had p>0.05. None of the randomizations required more than one draw of random assignment vector to meet this rule. But since this is the rule I would have followed, my intention was for the randomization inference to throw out any potential random assignments that would not have met this rule (Morgan and Rubin, 2012). However, I was legally obligated to delete all our application data prior to receiving de-identified outcome data, and not all of the previously available data was transferred back to us for the outcome analysis. As a result, I no longer have all the information on which I tested balance initially. Since I can not implement the exact randomization rule, I use straightforward randomization inference that does not exclude any randomization vectors. This should be conservative, making the adjusted p-values slightly larger than they would otherwise be.

I implement standard randomization inference separately for each main outcome, which tests the sharp null of no treatment effects for anyone. This test is more robust to outliers and clustering than standard tests of the null of no average treatment effects, and it more directly ties the statistical test to the randomness induced by random assignment. I re-randomize 5,000 times, reporting the probability we would find the treatment coefficient as large (in absolute value) as the one in the actual data given the distribution of treatment coefficients under the null.

Since the probability of a Type I error increases with the number of hypothesis tests performed, I also adjusts inference to control for the number of tests run within families (the WorkReady PAP pre-specifies that I would adjust within families). I control both for the family-wise error rate (FWER), which ensures that the probability I reject any null within the family of tests remains less than α, as well as the false discovery rate (FDR), which relaxes the conditions for rejection in exchange for additional power by allowing for q percent of null rejections to be false (Westfall and Young, 1993; Benjamini and Hochberg, 1995). For the FWER adjustment, I use the Stata command wyoung, written by Julian Reif (which uses bootstrap-based resampling rather than permutation of the treatment indicator; I use 5,000 bootstraps). For the FDR adjustment, I use Michael Anderson’s code based on the original Benjamini and Hochberg approach. I report the adjusted p-values under FWER control and the q-value – the smallest proportion of false rejections I could allow and still reject each null. See Davis and Sara B. Heller (2020) and Sara B. Heller et al. (2017) for additional discussion of the pros and cons of each adjustment.

To perform these adjustments, I group our hypotheses into families to answer the following questions: 1) do any criminal justice outcomes move? (any juvenile incarceration, any juvenile justice services, total number of arrests), 2) what type of crime changes? (violent, property, drug, and other arrests), 3) do family outcomes move? (fertility, receipt of child protective services, homeless shelter use), and 4) do behavioral health outcomes move? (receipt of substance abuse services, receipt of mental health services). Conceptually, we might prefer the fertility outcome to be grouped with behavioral health to capture risky behavior. But because my legal agreements required all HIPAA-covered services to remain in a separate file with limited covariates to avoid re-identification (despite the fact that the data do not include personal identifiers), I can not combine those outcomes into a single family. I make no adjustment to school persistence, since it is the single outcome in the education family that can be defined for everyone, regardless of missing data.42 Note that family 2 is the only relevant family for the OSC+ study given the data available for Chicago.

Tables A6 and A7 show the p-values from the original robust standard error calculation, p-values from randomization inference, and adjusted p- and q-values under FWER and FDR control for WorkReady and OSC+ respectively. Randomization inference changes little; p-values move around a small amount, but substantive conclusions are basically unchanged. Adjusting for the number of tests, on the other hand, often pushes the probability of false rejections above traditional cut-offs. The decline in other arrests in Philadelphia is the only result that remains at or below p = 0.05 across all the adjustments.

If this were the first and only set of tests for SYEP effects, these results might merit considerable caution about drawing conclusions. But there are several arguments for resisting too much caution here. First, the fact that I have two entirely independent samples across the two studies provides a built-in replication. The fact that both drug and other arrests decrease in both settings is informative; the probability we would be falsely rejecting both nulls across two independent experiments is much lower than the probability that any single result is a false rejection. Second, the decline in arrests and incarceration has also occurred in every other study of SYEPs where they have been measured. This suggests that a null hypotheses of no effects might actually be conservative; all of the existing evidence suggests that these effects are unlikely to be zero. I therefore argue that, at least for criminal justice outcomes, the adjustments for multiple testing often move p-values above the typical 0.1 threshold largely because the non-compliance in the study reduces power, not because the probability of false rejections is actually high.

We should be more cautious about the result that child protective services decline. That outcome is only available in Philadelphia, so I lack the built-in replication across cities. And it is not an outcome that has been tested elsewhere, so zero-impact is a more reasonable null hypothesis. Because it is measuring an extremely costly outcome — a substantiated risk of child abuse and neglect — the possibility that SYEPs affect the outcome is worth some attention even if there is some risk of a false rejection. And the fact that there is a more precisely-estimated decline among Black youth, who have elevated rates of these services in the control group, is worth noting, although certainly not dispositive given how many tests are run in the subgroup analysis. As I argue in the main text, the result should be a priority for future research before we draw strong conclusions about SYEPs effects.

F.4. Imputed School Data

Missing education data is prevalent and substantively important. The absence of information is not random; missingness could indicate that students graduated, dropped out, transferred, or are attending a non-public school. Charter school enrollment generates considerable missing data in our context. In Chicago, charters report enrollment and attendance but not grades or disciplinary issues in the administrative records; in Philadelphia, charters only report enrollment consistently. This means grades, suspensions, and in Philadelphia, attendance data are missing for charter school students. This group makes up about 4 percent of the Chicago sample and a little over a third of the Philadelphia sample in the baseline years.

Because of the prevalence of missing data and the possibility that treatment could affect missingness by changing the probability of dropout, graduation, or transfer, the best measured education measure, which is non-missing for everyone who has any kind of education data available, is school persistence, or an indicator for whether a youth has either graduated or remains enrolled in a public school as of a given school year, reported above. This section reports results for other education measures with more missing data—absences, grades, and misconduct—using both non-missing data only as well as a variety of imputation methods. But the amount of missing data here, especially in Philadelphia where charters are so widespread, makes the interpretation of these results more cautious.

Tables A8 through A11 show education outcomes with different assumptions about missing data, separately for each study’s ITT and LATE. All regressions drop pre-program graduates, since they can not have school data by construction. I make one departure from the pre-analysis plan in these outcomes by excluding Keystone test scores as an outcome. This is because the tests were not regularly administered to everyone in the sample in every year, which I did not know at the time of pre-specification.

In each table, Panel A shows the treatment-control difference for absences, GPA, and suspensions for non-missing data only, which implicitly assumes data are missing completely at random. For students with at least 1 day attended in a school year, I assign 0s for suspensions and absences if those variables are missing. Panel B imputes the mean outcome variable by group (treatment and control) and study year for all missing data (excluding pre-program graduates but including those with missing variables or missing all school data). This assumes that the data are missing completely at random conditional on study year and group. Panels C and D make the extreme assumption that all missing data is missing for either the most low-performing or most high-performing students, respectively, in the spirit of Lee bounds. For the low-performing imputation, I assign a 0 GPA, 1 for ever being suspended, and the 95th percentile of non-missing days absent for all missing observations. For the high-performing imputation, I assign the 95th percentile of GPA, 0 for ever being suspended, and 0 days absent for all missing data. The final two panels take a hybrid approach. They assume that data are missing completely at random (conditional on year and group) for those who are marked as transfering out of the district in the school records, and assign transfers the group-year mean. They then assign all-low or all-high values for the remaining missingness.

For WorkReady, the fact that the control means move around so much across imputations emphasizes how much missing data there is. This is largely from the fact that about a third of youth attend charter schools, which do not report grades, misconduct, or days absent to SDP. Because the treatment-control difference on missing data is relatively small, the various imputations do not change the substantive conclusion that there is little significant change in school outcomes. If anything, there is perhaps an indication of small decreases GPA and days absent, and small increases in suspensions, though no results are statistically significant. But the amount of missing data makes it hard to draw clear conclusions, since the point estimates do move around under the different assumptions about missing data.

OSC+ has less missing data, in part because charter schools in Chicago do report attendance data, though not grades. The coefficient on GPA moves around somewhat less as a result, always indicating negative but quite small and statistically insignificant effects on GPA. The coefficient on days absent moves around more across the various imputations, flipping sign but never differentiable from zero. Ever being suspended, on the other hand, consistently increases by between 8 and 11 percentage points for compliers, which is proportionally quite a large increase relative to complier means (between 51 and 73 percent). It is not entirely clear how much stock to put in this result; prior work on OSC+ has not reported misconduct results, because the school district has warned how unreliable the data are (with reporting quite inconsistent, even within schools). And no other SYEP study has found an increase in misconduct.

It is certainly possible that this is a real treatment effect; as shown in Table A13 below, there is also an indication that treatment youth are less likely to still be enrolled or graduate in the second post-randomization year. Table A12 shows that these negative impacts may be concentrated in the jobs-only treatment arm, although the arms are not statistically different from each other. And in year 2, school persistence declines among the job-only group (see Appendix F.4). One possibility is that, by reducing arrests, the program keeps some youth who would otherwise be incarcerated in traditional schools, rather than the schools in juvenile detention or prison. If so, they might be less likely to continue attending (whereas attendance in criminal justice facilities is mandatory), and more likely to commit disciplinary infractions. Or, given the slight concentration of these effects in the job-only arm, it is possible that the stronger connection to the labor force introduces additional problems at school, and the extra supports may help prevent youth being pulled into the labor market before finishing school. This is certainly an issue to which future work should attend.

G. Longer-Run Results

Table A13 shows year 2 results for the 2017 cohort of WorkReady (the only one observed long enough in the data to have year 2 results), and results for years 2 and 3 for OSC+. Since the year 2 results are only a single cohort, they should be compared to just the 2017 cohort’s results in year 1, shown separately in the next section.

In both cities, there is some suggestion that crime reductions last beyond the first post-randomization year. In Philadelphia, property crime declines by 2.2 arrests per 100 compliers, effectively eliminating property crime relative to the control complier mean. Point estimates on total and other arrests remain proportionally large and negative, but not statistically significant. Similarly, point estimates on mental health and child protective services are negative and proportionally large, but not significantly different from 0. In Chicago, year 2 arrest results are typically negative but not significant. But in year 3, the drops in total, property, and other arrests are all marginally significant again.

The education results are less encouraging. There is a marginally significant decline in school persistence in OSC+ of 6.2 percentage points. As discussed in the previous section, the results by treatment arm suggest this is driven by the jobs-only group. So it is possible that a summer job unaccompanied by additional program supports gives youth more of a chance to engage with their employer, pulling them into the labor market and out of school. Explanations not specific to treatment arm are also possible. It could be that summer jobs are pulling youth out of summer school, making it harder for them to complete the credits they need to motivate staying in school. Or it could be that the decline in crime keeps youth out of detention facilities, which force them to attend school (but rarely give them enough credits to graduate). If youth who would otherwise drop out do not get incarcerated, they are more free to unenroll in school. This could decrease enrollment at the margin, even if it does not change eventual educational attainment.

If this decline in school engagement persists, it will be an important caveat to the overall effects of SYEP. But it is also in conflict with findings in other contexts – Philadelphia here, where there is no change in school persistence, as well as zero school effects in prior Chicago and NYC work, and positive impacts in Boston. On the other hand, it is more consistent with the results in Heller and Kessler (2021), in which letters of recommendation increased the employment effects of SYEPs and slowed down graduation for those still in school. Given that many of the youth are not yet old enough to complete their school career, it will be important to follow up after more time has passed to see if these results persist.

H. Subgroup Interactions

For WorkReady, the subgroup divisions included here (gender, race, and age) were pre-specified in the pre-analysis plan. I added one additional split – cohort – ex post, because it adds useful, if unanticipated, variation in baseline risk (the 2018 cohort turned out to be at far lower risk of negative outcomes, helping to identify variation in risk that is useful for the heterogeneity analysis). See Table A14 for baseline descriptive statistics by cohort. I also exclude one split, which the pre-analysis plan specified was more of interest for program implementers than for substantive reasons (whether the youth had previously participated in certain types of City programming). For OSC+, I did not post a pre-analysis plan, since the analysis was intended to follow the previous analyses of OSC+ quite closely. As a result, I focus on the subgroup splits that were previously reported in Davis and Heller (2020): gender and prior arrests. The previous paper also reported splits on whether youth were in school and whether they had worked before. But this 2015 study population was all in school prior to the program, so that split does not apply here. And I do not have access to data on employment records, so we can not test the prior employment split.

I hesitate to over-interpret subgroup differences, because they involve a huge number of hypothesis tests, and because I often lack the statistical power to differentiate between groups. These splits serve mostly to generate variation in baseline rates of the outcomes, used in the main risk-responsiveness analysis. Nonetheless, this section highlights some of the results, which should be considered tentative until validated in future studies that are better powered for this kind of analysis. All estimates come from a model with a one-way treatment interaction with the relevant covariate, but the linear combination of coefficients and standard errors, representing the net effect for each group, are shown to ease interpretation.

Table A15 shows that the 2018 cohort of WorkReady consistently faces a floor effect across many outcomes. The 2018 control group hits the floor on several outcomes, with 0 other types of arrest, no parenthood, and no homeless shelter use. Because the 2018 control means are so low, their program impacts are often significantly more positive than the 2017 cohort; there was no room for the decline that occurred in 2017. This echoes the larger targeting point of the paper: that there are bigger declines for youth at higher risk of the relevant outcomes.

Tables A16 and A17 show a similar pattern by gender. The control means for arrests in both cities reflect boys’ disproportionate involvement in the criminal justice system. Boys in the sample are also considerably more likely to have received substance abuse treatment (1.8 versus 0.2 percent) and less likely to persist in school (89 versus 93 percent in Philadelphia, and 94 versus 96 percent in Chicago). Point estimates on arrests are generally bigger for boys, though in Chicago the declines for women are still statistically significant on their own.

In Philadelphia, males have a proportionally huge and statistically significant decline in substance abuse treatment (a 1.1 percentage point ITT decline relative to a control mean of 1.8 percent). This contributes to an overall drop in the receipt of behavioral health services for boys, 2.5 percentage points, or a 22 percent decline (p = 0.057). Males also show the only significant adverse effect, a 0.9 percentage point increase in parenthood, which is almost a tripling relative to the control mean. While it is certainly possible that the program increases confidence and income in a way that increases risky sexual activity, it is also true that there is more of a margin for increases in reporting of childbirth for fathers than for mothers (who have a negative but not significant point estimate), and that the outcome is quite rare..

Table A18 shows Chicago results separated by whether someone had a prior arrest. As expected, the control means on arrests are much higher for those with a prior arrest record. The point estimates also tend to be much larger, though they are only statistically differentiable from each other for other arrests. Both groups have significant declines in arrests, so it’s not that lower-risk groups do not respond at all. But the groups move proportional to their baseline, so that declines tend to be bigger for the more criminally-involved group.

Table A19 shows Philadelphia results by race. Since some of the racial and ethnic groups in the data are too small to estimate separate effects for, the analysis just divides the sample into Black and non-Black, which includes Hispanic, White, Asian, and other. This is not because these groups necessarily all have the same treatment effect, but rather because there is not enough data to estimate separate effects. Incarceration declines are significantly bigger among Black youth, though the point estimates for arrests are suggestively larger for non-Blacks. The other significant difference is in the receipt of child protective services; Black youth experience a significant 0.9 percentage point (43 percent) decline in the prevalence of these services, with no change for non-Black youth.

Table A20, shows treatment effects for those under 16 versus 16 and over. Significant effects are typically concentrated among the younger group, other than a decline in substance abuse for the older group. But differences between age groups are not significant.

The final interaction table, Table A21, shifts attention to subgroup heterogeneity for the same index I use in the endogeneous stratification analysis (thus results and control means are in standard deviation units, standardized separately for each city). Results are shown for each gender-race/ethnicity-age group. As highlighted in the main text, a major limitation of looking at treatment heterogeneity within a single study is limited power and multiple testing concerns. Many of the standard errors are quite large due to small cell sizes, and as indicated by the dashes, some cells contain no variation in the outcome; there are also a lot of tests in the table. Despite the noise, and although somewhat less true for Black and Hispanic women, one can see the pattern from the main text in many of the cells with large enough sample sizes: higher control means frequently correspond to bigger declines in the LATE.

I. Risk-Responsiveness

I.1. Alternative Index for Endogeneous Stratification

Appendix Tables A22 and A23 show alternative versions of the endogenous stratification exercise using more expansive indices. Rather than limit each index to the outcomes where the SYEP has significant main effects, these tables use all the available socially costly outcomes. This avoids cherry-picking outcomes based on ex-post significance, but it limits power by combining outcomes that the intervention does not move with outcomes that are affected.

In Philadelphia, the index includes juvenile incarceration, juvenile justice services, the 4 types of arrests separately (excluding total since that is the sum of the rest), parenthood, an indicator for child protective services, and an indicator for using a housing shelter. In Chicago, it includes each of the 4 types of arrests separately. In Philadelphia, the pattern of results is quite similar, with the highest-risk groups having the largest point estimates. But as expected, the estimates are noisier. In Chicago, the overall impact on the index is no longer significant, and the point estimates flatten out across risk groups. This suggests that the outcomes not included in the original index (violent and property crimes) may have a somewhat different pattern of responses there. But the standard errors are big enough that this may also just be additional noise from including more skewed outcomes without detectable treatment effects; unlike in Philadelphia and the more limited Chicago index, the standard errors on the highest-risk group are over 5 times as large as the point estimate.

I.2. Selection by the Propensity to Participate and Marginal Treatment Effects

The main text focuses on treatment heterogeneity across the distribution of Y0 . But as shown in the endogeneous stratification results, the risk categories defined by Y^0 appear correlated with probability of take-up (with higher risk groups less likely to take up, at least in Philadelphia). This raises the question of whether it is possible to identify a second margin of treatment heterogeneity: over the probability of take-up, as in the marginal treatment effects (MTE) literature (Brinch et al., 2017; Heckman and Vytlacil, 2005; Kowalski, 2018a; Mogstad et al., 2018; Walters, 2018). In theory, this kind of exercise can illuminate the differences in baseline characteristics, the Y0s , and the treatment effects for people who are more or less willing to participate, potentially helping to explain variation in LATEs across settings.

There is, however, a key feature of this institutional setting that prevents a full MTE analysis using a binary instrument, in the vein of Brinch et al. (2017). The key conceptual underpinning of using a binary instrument to identify the MTE is defining an unobserved propensity to participate, UD , where the instrument shifts the value it takes on, p . With a binary instrument, a MTE analysis requires an ancillary assumption: either that the MTE is linear in p to extrapolate the unobserved untreated outcome for always-takers and the unobserved treated outcome for never-takers, or that there is a different shape restriction such being monotone in p (Brinch et al., 2017; Mogstad et al., 2018). In many cases, this is a reasonable assumption. But in my setting, it creates a problem.

In both the Chicago and Philadelphia studies, the decision to take up the program is not simply the youth’s decision, which might plausibly be linear in p if we think of youth decision-making as based on a Roy model. Rather, it is a combination of two entirely separate processes. First, the provider agency has to decide it wants to serve the young person, successfully locate and contact her, and make the offer. Only then does the young person make her own decision about take-up given the costs and benefits she faces. The role of provider in determining take up is not trivial in this setting; both cities’ providers received more names than they had program slots to ensure they could fill their slots, so they always had to make decisions about whom to approach and in what order. And in Philadelphia especially, providers were completely free to not serve treatment youth and to serve control youth, if they preferred.

This combination of two separate take-up processes undermines the necessary assumption that the MTE is linear or monotone in p . Consider two people who both have some p close to 0, but for different reasons: in one case, the youth is very disconnected and hard to locate, so the provider decides to fill that slot with someone who is easier to serve despite the large gain the youth could get from the program. In the other, the provider makes the offer, but the young person is already employed at a higher paying job and so faces a small gain (or even a loss) from the program. These two young people are unlikely to share similar potential outcomes or treatment effects, but they would both contribute to the never-takers’ observed average p and average untreated outcome. The same issue could arise for compliers and always-takers, with marginal treatment effects within each group depending on how much providers versus individuals determined the composition of each compliance group.

In theory, we might accept that the MTE involves averaging over different types of people in each compliance group. But doing so would lose the clear interpretation of p coming from a Roy model’s selection based on one individual’s benefits and costs of participating, which is the core motivation for assuming linearity. In other words, since each p does not represent one single type of person, but rather a combination of people based on both who they are and the decisions of their assigned provider, the MTE seems likely to vary over the distribution of p in a variety of ways. As such, it seems incredibly difficult to argue that that the MTE would be linear or monotone in p .

Beyond the theoretical argument, Table A24 also provides some empirical evidence against the linearity (or monotonicity) of the MTE in p . It shows the means for always-takers, compliers, and never-takers for the different baseline characteristics, with significance stars indicating differences between each category and the compliers. Although baseline characteristics are an imperfect proxy for potential outcomes, we might think that if the increasing resistance to participating across these three compliance categories were linearly related to an individual’s treatment effect at each p , we should also see roughly linear patterns in the baseline covariates across these categories, especially those that are baseline measures of the outcome variables.43

In fact, moving across the three levels of take-up resistance, not only is there is little indication of linear patterns, but the pattern is not even monotone. Always-takers and never-takers are often more similar to each other than they are to compliers, generating some U-shaped patterns. In Chicago especially, the baseline covariates most related to the main outcomes—the arrest counts—show fewer arrests for compliers but more for both always- and never-takers (significantly different only for the former). This pattern is not consistent with the idea that the propensity to take up is linearly or monotonically related to the propensity for crime Y0 . Of course, baseline characteristics are not perfect predictors of Y0 , and Y0 is not equivalent to the marginal treatment effect. Nonetheless, both the U-shaped patterns and the differences in compliance groups across cities leave the impression that, perhaps due to the variation in provider and individual interactions during recruitment, the level of resistance to take-up is not clearly related to potential outcomes.

So overall, this does not appear to be a setting where the necessary assumptions about the relationship between p and the MTE are particularly palatable. This limits my ability to try to estimate treatment effects for always- and never-takers, or to back out the average treatment effect for these populations. Additionally, as Table A24 shows, the level of behavior is so different across the two cities that it may be difficult to extrapolate one LATE to the other setting, even for similar propensities to participate. As Kowalski (2018b) shows, even without ancillary assumptions, we can still compare untreated outcomes among compliers and never-takers to understand selection heterogeneity. In both cities, compliers engage in more socially costly post-program behavior than never-takers in the absence of the program (i.e., if we use the standardized index from the endogenous stratification exercise as the dependent variable, there is negative selection on Y0 into complying).

For transparency, I note that if we ignore the reasons why linearity is not plausible here and use the index as the dependent variable in a MTE analysis,44 I calculate MTEs that are largest (most negative) for always-takers and smaller but still negative for never-takers in Philadelphia (though the MTE slope is not statistically different from zero). In Chicago, however, I find the opposite: the biggest decreases in costly outcomes occur for the never-takers, with negative but closer-to-zero effects among always takers. I hesitate to over-interpret the MTEs given the apparent failure of the method’s key assumption. But I take the opposite signs on the MTE slopes across cities—despite the fact that the proportion of compliers and never-takers is fairly similar in both cities—as further evidence that the probability of take-up as measured by the different compliance categories is not the key element to understanding treatment heterogeneity in this context.

I.3. Details on Cross-Study Figures

The main text’s discussion of Figure 2 refers to this section for a list of which estimates are included in each panel. The details are as follows: Panel A includes Philadelphia coefficients for main effects on total and other arrests, juvenile incarceration, child protective services, and year 2 property crime arrests. For OSC+, it includes main effects on drug and other arrests, and year 3 arrests for all, property, and other crimes. Panel B adds the decline in violent-crime arrest and increase in years 2 and 3 for property crime arrests in OSC+ 2012; the decline in violent-crime arrests for OSC+ 2013; the decline in incarceration and mortality from NYC’s SYEP; and the decline in violent and property crime arraignments in Boston (Davis and Sara B. Heller, 2020; Modestino, 2019; Gelber et al., 2016). I note that Kessler et al. (2021) was released shortly before this paper. Because that study is still a working paper, and point estimates sometimes move around somewhat over the review process, I do not include those estimates in the graph; all estimates in these graphs have been through peer review and are final. The pattern in their findings is, however, quite consistent with the one shown here.

The subgroups added in Panel C include behavioral health services, incarceration, total and drug arrests, parenthood, and substance abuse for boys in WorkReady; other arrests for girls in WorkReady; incarceration, total and other arrests, and child protection services for the 2017 cohort in WorkReady; incarceration, other arrests, child protection services, and behavioral health services for Black youth in WorkReady; total and property arrests for non-Black youth in WorkReady; total, drug, and other arrests for those under age 16 in WorkReady; substance abuse services for those 16 and older in WorkReady; drug arrests for women in OSC+ 2015; other arrests for those with a prior arrest and drug arrests for those with no prior arrest in OSC+ 2015; incarceration impacts for both age groups, for men, whites, and Blacks, for those not working, and those not living in an empowerment zone in NYC; mortality effects for men, Latinos, younger people, those not working, those living in an empowerment zone, and those observed through the 9-year follow-up in NYC; violent-crime arrests for in-school youth, males, females, those with prior arrests, and those without prior arrests in the pooled OSC+ 2012 and 2013 sample; and drug arrests for in-school youth and property-crime arrests for females in the pooled OSC+ 2012 and 2013 sample. The Boston study does not report control means by subgroup, so no subgroup effects are included from that study. Additionally, the Boston study does not report LATE estimates; the main graph backs out the implied LATE from the reported ITT and the reported take-up rate.

Figure A.2, showing the ITT instead of the LATE, reflects a similar pattern but is slightly more noisy. This suggests that a small part of the differences across groups is in the take-up rate. Although those differences may be informative for policy (e.g., if some types of youth are easier to enroll), some of the take-up differences are really due to differences in experimental designs across studies, i.e., encouragement designs versus studies that strongly enforced compliance. In the ITT version of these graphs, there is also some additional error in our calculation of the NYC subgroup estimates. This is because the original paper only reports point estimates from two-stage least squares. Here I back out intent-to-treat estimates by multiplying IV estimates by the first stage. But the paper only reports the overall first stage. If take-up varies by subgroup, these estimates will mis-estimate the subgroup ITTs.

I.4. Cross-Study Figure Robustness

As in Section I.1 above, one might worry that showing only statistically significant estimates in Figure 2 could falsely generate the appearance of a downward slope in the relationship between control means and LATEs. To see why, note that the standard errors on each point estimate have a mechanical force pushing them to grow as we move from the left to the right side of the graph. Because all the outcomes are either indicators or counts, moving from a low control mean to a higher one (always less than 0.3 in the Figure) will correspond to an increase in the variance of the outcome. A higher variance will generate a bigger standard error, all else equal (though differences in sample sizes and take-up rates will also influence the precision of each estimate in practice). In theory, it is possible that effects would only be statistically differentiable from zero on the right side of the figure as they stretched farther away from the x-axis driven by noise or chance, even if true treatment effects did not have a downward slope.

One way to assess the role of statistical noise in Figure 2’s risk-responsiveness relationship is to run the regression the graph implies (regress LATEs on control means) including inverse variance weights to adjust for the uncertainty inherent in each estimate. This regression is, in fact, a bit awkward to interpret, since point estimates represent changes across different outcomes and different units (either indicators or counts, so a 1 unit increase is always an additional socially-costly incident within the group). But to ensure the pattern in Figure 2 is not generated solely by the larger variances on the right side of the graph, I use the reported standard errors on each point estimate to calculate each point’s variance, then regress the LATE on the control mean with inverse variance weights. Even after weighting, the slope of the line is still negative and statistically significant ( β^=0.158 , p=0.008 ).

In fact, the relationship still holds even if I include insignificant main effects in the regression. Conceptually, it is unclear whether one should include insignificant effects; the figure includes a range of outcomes, and if the program truly has a null effect on set of outcomes, it would not make sense to ask whether those effects are bigger for higher control means. That is, the argument in the paper is not that SYEPs have bigger effects for all outcomes with higher control means, but rather that for those outcomes SYEPs affect, they generate a bigger drop in groups with higher rates of the behavior. On the other hand, some null findings may be due to imprecision rather than true null effects, and some significant effects could be from chance. So there is a risk that focusing on significant effects could be misleading. By including all estimated effects regardless of statistical significance, we can assure ourselves that even in the extreme case that all null effects are worth considering, the relationship between treatment effect magnitude and control means holds. In practice, even including an additional 41 statistically insignificant, inverse variance-weighted main effects does not change the relationship much ( β^=0.145 , p=0.002 ).45

J. Standard Deviation of Outcome Variables

To aid in any potential use of these studies in future meta-analyses, Table A25 reports the standard deviations of the control group for all the outcome variables used in the analysis.

Figure A.1:

Figure A.1:

Time Path of ITT by Month

Note: Each figure shows the cumulative ITT up to the month shown on the x-axis, with the 95 percent confidence interval in gray. The vertical dotted line shows the end of the program in Chicago. In Philadelphia, randomization occurred at multiple dates, such that the program ended at different times for different youth. The two dotted lines contain the range of program end dates.

Figure A.2:

Figure A.2:

Size of ITTs Relative to Control Means

Note: Point estimates and control means taken from this paper, Davis & Heller (2020), Gelber et al. (2016), and Modestino (2019). See text and Appendix I for details.

Table A1:

Family, Health, and School Persistence Impacts in the First Year After Randomization

ITT CM LATE CCM
WorkReady

Graduated or Still Enrolled in School −0.005 (0.007) 0.945 −0.013 (0.019) 0.971
Becomes A Parent 0.001 (0.003) 0.011 0.002 (0.009) 0.000
Any Receipt of Child Protective Services −0.005* (0.003) 0.018 −0.016* (0.009) 0.013
Any Stay in a Shelter 0.002 (0.002) 0.001 0.006 (0.005) 0.000
Any Receipt of Behavioral Health Services −0.011 (0.009) 0.113 −0.033 (0.026) 0.108
 Any Receipt of Substance Abuse Services −0.003 (0.003) 0.008 −0.008 (0.008) 0.009
 Any Receipt of Mental Health Services −0.01 (0.009) 0.110 −0.029 (0.026) 0.104

OSC+

Graduated or Still Enrolled in School −0.001 (0.006) 0.951 −0.003 (0.023) 0.978

Note: WorkReady N=4497, OSC+ N=5405. Table shows estimated intent-to-treat (ITT) and local average treatment effects (LATE), controlling for baseline covariates and randomization block. CM is control mean; CCM is control complier mean, rounded to 0 when estimate is negative. Graduated/still enrolled excludes pre-program graduates and those unmatched to education records; WorkReady N= 3858, OSC+ N= 4077. Becoming a parent is measured across all available outcome years; other outcomes in first post-randomization year only. Robust standard errors in parentheses, clustered by person as the same person can appear in both cohorts.

*

p<0.1

**

p<0.05

***

p<0.01

Table A2:

OSC+ Outcomes in the First Year After Randomization by Treatment Arm

ITT CM ITT CM Test of Difference
Job Only Job and Mentor

Total Number of Arrests −0.015 (0.018) 0.174 −0.031* (0.018) 0.177 0.424
 Number of Violent Arrests 0.004 (0.007) 0.044 0.007 (0.008) 0.031 0.692
 Number of Property Arrests 0.002 (0.006) 0.014 −0.001 (0.006) 0.026 0.605
 Number of Drug Arrests −0.007 (0.007) 0.037 −0.017** (0.007) 0.030 0.146
 Number of Other Arrests −0.014 (0.011) 0.079 −0.020* (0.011) 0.090 0.609

Note: N=5405. Graduated or still enrolled in school excludes pre-program graduates and those unmatched to education records; N= 4077. Table shows estimated intent-to-treat (ITT), controlling for baseline covariates and randomization block. CM is control mean; CCM is control complier mean, rounded to 0 when estimate is negative. The Test of Difference column shows the p-value from the test that the treatment coefficients for each treatment arm are equal. Robust standard errors in parentheses.

*

p<0.1

**

p<0.05

***

p<0.01

Table A3:

Program Impacts in First Year After Random Assignment, No Baseline Covariates

ITT CM LATE CCM
WorkReady

Any Juvenile Incarceration −0.006** (0.003) 0.014 −0.019** (0.009) 0.023
Any Receipt of Juvenile Justice Services −0.004 (0.003) 0.013 −0.012 (0.010) 0.023
Total Number of Arrests −0.012** (0.005) 0.028 −0.034** (0.015) 0.050
 Number of Violent Arrests −0.003 (0.004) 0.011 −0.009 (0.010) 0.019
 Number of Property Arrests −0.002 (0.003) 0.008 −0.007 (0.008) 0.013
 Number of Drug Arrests −0.003* (0.002) 0.004 −0.008* (0.005) 0.008
 Number of Other Arrests −0.003** (0.001) 0.004 −0.010** (0.004) 0.009

OSC+

Total Number of Arrests −0.030* (0.017)  0.176 −0.114* (0.065) 0.193
 Number of Violent Arrests 0.004 (0.006) 0.037 0.013 (0.023) 0.019
 Number of Property Arrests 0.000 (0.005) 0.020 0.001 (0.018) 0.020
 Number of Drug Arrests −0.014** (0.006) 0.034 −0.052** (0.024) 0.058
 Number of Other Arrests −0.020** (0.010) 0.084 −0.076** (0.038) 0.095

Note: WorkReady N=4497, OSC+ N=5405. Table shows estimated intent-to-treat (ITT) and local average treatment effects (LATE), controlling for randomization block. CM is control mean; CCM is control complier mean, rounded to 0 when estimate is negative. Robust standard errors in parentheses, clustered by person for WorkReady, where the same person can appear in both cohorts.

*

p<0.1

**

p<0.05

***

p<0.01

Table A4:

Family, Health, and School Persistence Impacts in First Year After Randomization, No Baseline Covariates

ITT CM LATE CCM
WorkReady

Graduated or Still Enrolled in School 0.000 (0.008) 0.945 0.001 (0.021) 0.957
Becomes A Parent −0.001 (0.003) 0.011 −0.001 (0.009) 0.004
Any Receipt of Child Protective Services −0.008** (0.003) 0.018 −0.023** (0.010) 0.020
Any Stay in a Shelter 0.001 (0.001) 0.001 0.004 (0.004) 0.000
Any Receipt of Behavioral Health Services −0.015 (0.010) 0.113 −0.045 (0.029) 0.119
 Any Receipt of Substance Abuse Services −0.003 (0.003) 0.008 −0.009 (0.008) 0.010
 Any Receipt of Mental Health Services −0.013 (0.009) 0.110 −0.041 (0.029) 0.115

OSC+

Graduated or Still Enrolled in School 0.004 (0.007) 0.951 0.015 (0.025) 0.961

Note: WorkReady N=4497, OSC+ N=5405. Graduated/still enrolled excludes pre-program graduates and those unmatched to education records; WorkReady N=3858, OSC+ N=4077. Table shows estimated intent-to-treat (ITT) and local average treatment effects (LATE), controlling for randomization block. CM is control mean; CCM is control complier mean, rounded to 0 when estimate is negative. Becoming a parent is measured across all available outcome years; other outcomes in first post-randomization year only. Robust standard errors in parentheses, clustered by person as the same person can appear in both cohorts.

*

p<0.1

**

p<0.05

***

p<0.01

Table A5:

Intent-to-Treat Results with Alternative Functional Forms, Probit or Poisson

WorkReady OSC+
AME CM AME CM

Any Juvenile Incarceration −0.008* (0.005) 0.019
Any Receipt of Juvenile Justice Services −0.003 (0.004) 0.017
Total Number of Arrests −0.011* (0.006) 0.028 −0.028* (0.016) 0.176
 Number of Violent Arrests −0.003 (0.004) 0.011 0.004 (0.006) 0.037
 Number of Property Arrests −0.003 (0.003) 0.011 0.000 (0.005) 0.020
 Number of Drug Arrests −0.013 (0.009) 0.020 −0.013** (0.006) 0.034
 Number of Other Arrests −0.006** (0.003) 0.007 −0.019* (0.010) 0.084
Graduated or Still Enrolled in School 0.000 (0.008) 0.940 0.000 (0.006) 0.951
Parenthood 0.000 (0.004) 0.013
Receipt of Child Protective Services −0.010** (0.005) 0.022
Any Stay in a Shelter 0.003 (0.003) 0.003
Receipt of Behavioral Health Services −0.012 (0.008) 0.114
 Substance Abuse Services −0.002 (0.003) 0.011
 Mental Health Services −0.011 (0.008) 0.111

Note: WorkReady N=4497, OSC+ N=5405. Graduated/still enrolled excludes pre-program graduates and those unmatched to education records; WorkReady N=3858, OSC+ N=4077. AME indicates average marginal effects. Total, violent, and property crimes use Poisson regression; all others only take on values 0 and 1, so use probit. CM is control mean. Robust standard errors in parentheses, clustered by person for WorkReady where the same person can appear in both cohorts.

*

p<0.1

**

p<0.05

***

p<0.01

Table A6:

Multiple Testing by Family, WorkReady

Standard P-Value Randomization Inference P-Value FWER Adjustment FDR Adjustment
Any Crime
 Any Juvenile Incarceration 0.08 0.10 0.15 0.12
 Any Receipt of Juvenile Justice Services 0.47 0.50 0.48 0.47
 Total Number of Arrests 0.04 0.05 0.11 0.12
Type of Crime
 Number of Violent Arrests 0.53 0.55 0.73 0.53
 Number of Property Arrests 0.48 0.46 0.73 0.53
 Number of Drug Arrests 0.10 0.11 0.29 0.21
 Number of Other Arrests 0.01 0.02 0.05 0.04
Graduated or Still Enrolled in School 0.50 0.50
Family
 Becomes a Parent 0.81 0.81 0.81 0.81
 Any Receipt of Child Protective Services 0.09 0.13 0.28 0.28
 Any Stay in a Shelter 0.19 0.17 0.36 0.29
Health
 Any Receipt of Substance Abuse Services 0.30 0.26 0.43 0.31
 Any Receipt of Mental Health Services 0.26 0.28 0.43 0.31

Note: N=4497. Table shows the adjusted p-values controlling for the family-wise error rate (FWER) and q-values showing the smallest false discovery rate (FDR) under which each null can be rejected. Graduated/still enrolled excludes pre-program graduates and those unmatched to education records; N=3858.

Table A7:

Multiple Testing by Family, OSC+

Standard P-Value  Randomization Inference P-Value FWER Adjustment FDR Adjustment
Total Number of Arrests 0.13 0.12
Type of Crime
 Number of Violent Arrests 0.35 0.34 0.58 0.47
 Number of Property Arrests 0.93 0.92 0.93 0.93
 Number of Drug Arrests 0.04 0.04 0.15 0.13
 Number of Other Arrests 0.07 0.07 0.19 0.13
Graduated or Still Enrolled in School 0.91 0.93

Note: N=5405. Table shows the adjusted p-values controlling for the family-wise error rate (FWER) and q-values showing the smallest false discovery rate (FDR) under which each null can be rejected. Graduated/still enrolled excludes pre-program graduates and those unmatched to education records; N=4077.

Table A8:

Intent to Treat on Education Outcomes in the First Year After Randomization with Imputations, WorkReady

Grade Point Average Any Suspensions Days Absent
Panel A. Non-Missing Only
Intent to Treat −0.033 (0.036) 0.011 (0.013) −0.035 (0.733)
CM 2.242 0.108 15.3
N 2021 2166 2166
Panel B. Imputed Mean (N=4211)
Intent to Treat −0.009 (0.018) 0.004 (0.007) −0.320 (0.380)
CM 2.244 0.108 15.2
Panel C. Low-Performing Imputation (N=4211)
Intent to Treat −0.023 (0.027) 0.005 (0.011) 0.020 (0.532)
CM 1.081 0.539 34.5
Panel D. High-Performing Imputation (N=4211)
Intent to Treat −0.007 (0.026) 0.003 (0.007) −0.084 (0.415)
CM 3.049 0.056 7.9
Panel E. Low-Performing Imputation for Non-Transfers (N=4211)
Intent to Treat −0.014 (0.028) 0.004 (0.011) −0.227 (0.535)
CM 1.130 0.528 34.0
Panel F. High-Performing Imputation for Non-Transfers (N=4211)
Intent to Treat −0.014 (0.026) 0.004 (0.007) −0.196 (0.409)
CM 3.023 0.057 8.1

Note: Exludes pre-program graduates. Table shows estimated intent-to-treat effects, controlling for baseline covariates and randomization block, for various imputations. Non-Missing shows the treatment-control difference for non-missing data only. Imputed Mean imputes the mean outcome variable by group (treatment and control) and study year for all missing data. Low- and High-Performing Imputation assume all missing data is missing for either the most low- or high-performing students respectively. Low- and High-Performing for Non-Transfers assume data are missing completely at random (conditional on year and group) for those who are marked as transferring out of the district in the school records, then they assign all-low or all-high values for the remaining missingness. CM is control mean. Robust standard errors in parentheses, clustered by person as the same person can appear in both cohorts

*

p<0.1

**

p<0.05

***

p<0.01

Table A9:

Local Average Treatment Effect on Education Outcomes in the First Year After Randomization with Imputations, WorkReady

Grade Point Average Any Suspensions Days Absent
Panel A. Non-Missing Only
Local Average Treatment Effect −0.100 (0.110) 0.034 (0.040) −0.112 (2.265)
CCM 2.422 0.066 12.6
N 2021 2166 2166
Panel B. Imputed Mean (N=4211)
Local Average Treatment Effect −0.026 (0.053) 0.012 (0.019) −0.926 (1.088)
CCM 2.332 0.091 13.9
Panel C. Low-Performing Imputation (N=4211)
Local Average Treatment Effect −0.066 (0.078) 0.015 (0.032) 0.059 (1.523)
CCM 1.134 0.552 34.4
Panel D. High-Performing Imputation (N=4211)
Local Average Treatment Effect −0.019 (0.074) 0.010 (0.019) −0.242 (1.188)
CCM 3.135 0.039 6.2
Panel E. Low-Performing Imputation for Non-Transfers (N=4211)
Local Average Treatment Effect −0.040 (0.080) 0.011 (0.032) −0.656 (1.532)
CCM 1.163 0.542 34.3
Panel F. High-Performing Imputation for Non-Transfers (N=4211)
Local Average Treatment Effect −0.040 (0.074) 0.013 (0.019) −0.568 (1.172)
CCM 3.122 0.038 6.6

Note: Exludes pre-program graduates. Table shows estimated local average treatment effects (LATE), controlling for baseline covariates and randomization block, for various imputations. Non-Missing shows the treatment-control difference for non-missing data only. Imputed Mean imputes the mean outcome variable by group (treatment and control) and study year for all missing data. Low- and High-Performing Imputation assume all missing data is missing for either the most low- or high-performing students respectively. Low- and High-Performing for Non-Transfers assume data are missing completely at random (conditional on year and group) for those who are marked as transferring out of the district in the school records, then they assign all-low or all-high values for the remaining missingness. CCM is control complier mean, rounded to 0 when estimate is negative. Robust standard errors in parentheses, clustered by person as the same person can appear in both cohorts.

*

p<0.1

**

p<0.05

***

p<0.01

Table A10:

Intent to Treat on Education Outcomes in the First Year After Randomization with Imputations, OSC+

Grade Point Average Any Suspensions Days Absent
Panel A. Non-Missing Only
Intent to Treat −0.025 (0.024) 0.024** (0.012) 0.159 (0.559)
CM 2.394 0.153 25.1
N 3461 3851 3851
Panel B. Imputed Mean (N=4102)
Intent to Treat −0.012 (0.021) 0.024** (0.011) 0.083 (0.536)
CM 2.394 0.153 25.1
Panel C. Low-Performing Imputation (N=4102)
Intent to Treat −0.016 (0.029) 0.030** (0.012) 0.436 (0.597)
CM 2.014 0.204 27.8
Panel D. High-Performing Imputation (N=4102)
Intent to Treat −0.005 (0.027) 0.021* (0.011) −0.158 (0.572)
CM 2.604 0.144 23.6
Panel E. Low-Performing Imputation for Non-Transfers (N=4102)
Intent to Treat −0.017 (0.027) 0.029** (0.012) 0.354 (0.566)
CM 2.134 0.178 26.4
Panel F. High-Performing Imputation for Non-Transfers (N=4102)
Intent to Treat −0.006 (0.025) 0.023** (0.011) −0.091 (0.554)
CM 2.538 0.149 24.4

Note: Exludes pre-program graduates. Table shows estimated intent-to-treat effects, controlling for baseline covariates and randomization block, for various imputations. Non-Missing shows the treatment-control difference for non-missing data only. Imputed Mean imputes the mean outcome variable by group (treatment and control) and study year for all missing data. Low- and High-Performing Imputation assume all missing data is missing for either the most low- or high-performing students respectively. Low- and High-Performing for Non-Transfers assume data are missing completely at random (conditional on year and group) for those who are marked as transferring out of the district in the school records, then they assign all-low or all-high values for the remaining missingness. CM is control mean. Robust standard errors in parentheses.

*

p<0.1

**

p<0.05

***

p<0.01

Table A11:

Local Average Treatment Effect on Education Outcomes in the First Year After Randomization with Imputations, OSC+

Grade Point Average Any Suspensions Days Absent
Panel A. Non-Missing Only
Local Average Treatment Effect −0.095 (0.089) 0.087** (0.043) 0.581 (2.035)
CCM 2.422 0.159 25.6
N 3461 3851 3851
Panel B. Imputed Mean (N=4102)
Local Average Treatment Effect −0.043 (0.079) 0.092** (0.042) 0.312 (2.002)
CCM 2.376 0.152 25.8
Panel C. Low-Performing Imputation (N=4102)
Local Average Treatment Effect −0.060 (0.107) 0.112** (0.047) 1.635 (2.234)
CCM 2.064 0.160 26.1
Panel D. High-Performing Imputation (N=4102)
Local Average Treatment Effect −0.020 (0.099) 0.080* (0.042) −0.591 (2.136)
CCM 2.540 0.156 25.8
Panel E. Low-Performing Imputation for Non-Transfers (N=4102)
Local Average Treatment Effect −0.063 (0.101) 0.108** (0.044) 1.327 (2.117)
CCM 2.143 0.147 25.5
Panel F. High-Performing Imputation for Non-Transfers (N=4102)
Local Average Treatment Effect −0.023 (0.094) 0.084** (0.042) −0.342 (2.070)
CCM 2.500 0.156 26.1

Note: Exludes pre-program graduates. Table shows estimated local average treatment effects (LATE), controlling for baseline covariates and randomization block, for various imputations. Non-Missing shows the treatment-control difference for non-missing data only. Imputed Mean imputes the mean outcome variable by group (treatment and control) and study year for all missing data. Low- and High-Performing Imputation assume all missing data is missing for either the most low- or high-performing students respectively. Low- and High-Performing for Non-Transfers assume data are missing completely at random (conditional on year and group) for those who are marked as transferring out of the district in the school records, then they assign all-low or all-high values for the remaining missingness. CCM is control complier mean, rounded to 0 when estimate is negative. Robust standard errors in parentheses.

*

p<0.1

**

p<0.05

***

p<0.01

Table A12:

Education Outcomes by Treatment Arm, OSC+

ITT CM ITT CM Test of Difference
Job Only Job and Mentor

Graduated or Still Enrolled in School (Year One) −0.005 (0.008) 0.947 0.004 (0.008) 0.954 0.296
Days Absent 0.48 (0.690) 24.7 −0.162 (0.708) 25.5 0.439
Grade Point Average −0.039 (0.030) 2.404 −0.012 (0.029) 2.385 0.430
Any Suspensions 0.027* (0.015) 0.143 0.020 (0.014) 0.163 0.697
Graduated or Still Enrolled in School (Year Two) −0.026* (0.012) 0.895 −0.00 (0.011) 0.900 0.138

Note: Excludes pre-program graduates and those unmatched to education records; N=4407. Regressions use non-missing data only, N=3851 for days absent and suspensions and N=3461 for GPA. Table shows estimated intent-to-treat (ITT), controlling for baseline covariates and randomization block. CM is control mean. The Test of Difference column shows the p-value from the test that the treatment coefficients for each treatment arm are equal. Robust standard errors in parentheses.

*

p<0.1,

**

p<0.05,

***

p<0.01

Table A13:

Outcomes in the Second and Third Year After Randomization, by Year

WorkReady Year Two OSC+ Year Two OSC+ Year Three
ITT CM LATE CCM ITT CM LATE CCM ITT CM LATE CCM

Any Juvenile Incarceration −0.006 (0.004) 0.017 −0.023 (0.014) 0.022
Receipt of Juvenile Justice Services 0.001 (0.004) 0.015 0.003 (0.015) 0.004
Total Number of Arrests −0.004 (0.006) 0.027 −0.016 (0.024) 0.021 −0.011 (0.013) 0.123 −0.043 (0.049) 0.133 −0.022* (0.011) 0.110 −0.082* (0.043) 0.165
 Number of Violent Arrests 0.000 (0.004) 0.012 −0.001 (0.015) 0.000 −0.005 (0.005) 0.025 −0.017 (0.017) 0.029 0.004 (0.004) 0.016 0.015 (0.014) 0.019
 Number of Property Arrests −0.006** (0.003) 0.009 −0.022** (0.011) 0.021 −0.003 (0.004) 0.018 −0.011 (0.014) 0.028 −0.006* (0.004) 0.014 −0.023* (0.013) 0.031
 Number of Drug Arrests 0.003 (0.003) 0.003 0.013 (0.012) 0.000 0.002 (0.004) 0.013 0.006 (0.016) 0.016 −0.007 (0.004) 0.018 −0.025 (0.016) 0.025
 Number of Other Arrests −0.001 (0.002) 0.003 −0.006 (0.007) 0.007 −0.005 (0.009) 0.068 −0.020 (0.034) 0.060 −0.013* (0.008) 0.062 −0.050* (0.029) 0.089
Graduated or Still Enrolled in School 0.006 (0.011) 0.902 0.023 (0.039) 0.936 −0.016* (0.009) 0.897 −0.060* (0.034) 0.971
Receipt of Child Protective Services  −0.003 (0.003)  0.012  −0.011 (0.013) 0.010
Any Stay in a Shelter  0.000 (0.002)  0.002  −0.001 (0.008) 0.002
Receipt of Behavioral Health Services  0.001 (0.010)  0.097  0.003 (0.039) 0.061
 Substance Abuse Services  0.004 (0.004)  0.009  0.015 (0.015) 0.000
 Mental Health Services  −0.002 (0.010) 0.093  −0.008 (0.038) 0.067

Note: Only 2017 WorkReady cohort observed for second year of data, N=3,392. OSC+ N=5,405. Enrolled/graduated excludes pre-program graduates and those unmatched to education records; WorkReady N=2869; OSC+ N=4077. Table shows estimated intent-to-treat (ITT) and local average treatment effects (LATE), controlling for baseline covariates and randomization block. CM is control mean; CCM is control complier mean, rounded to 0 when estimate is negative. Robust standard errors in parentheses, clustered by person for WorkReady where the same person can appear in both cohorts.

*

p<0.1

**

p<0.05

***

p<0.01

Table A14:

WorkReady Descriptive Statistics and Baseline Balance by Cohort

2017
2018
Control Mean Treatment Mean Test of Difference Control Mean Treatment Mean Test of Difference
N 2,056 1,336 3,392 655 450 1,105
Demographics
 Age 15.7 15.8 0.98 15.4 15.4 0.80
 Male 0.39 0.41 0.47 0.39 0.38 0.75
 Black 0.77 0.76 0.73 0.83 0.80 0.24
 Hispanic 0.13 0.13 0.80 0.07 0.10 0.13
 White 0.05 0.06 0.25 0.02 0.00 0.04
 Other Race 0.05 0.05 0.88 0.08 0.10 0.35
 Is a Parent 0.02 0.02 0.52 0.01 0.00 0.49
Contact with Justice System
 Ever Incarcerated 0.027 0.024 0.56 0.008 0.000 0.025
 Ever Received Juvenile Justice Services 0.028 0.022 0.37 0.012 0.011 0.866
 Ever Arrested 0.053 0.041 0.10 0.024 0.033 0.392
 Total Number of Prior Arrests 0.068 0.055 0.23 0.037 0.033 0.813
  Violent 0.036 0.020 0.01 0.021 0.016 0.591
  Property 0.021 0.021 0.68 0.005 0.011 0.244
  Drug 0.004 0.003 0.98 0.000 0.000 .
  Other 0.006 0.011 0.19 0.011 0.007 0.470
Receipt of Social Services
 Ever Received Child Protective Services 0.16 0.14 0.15 0.13 0.10 0.29
 Ever Stayed in a Shelter 0.05 0.04 0.08 0.05 0.02 0.02
 Any Behavioral Health Service 0.28 0.27 0.69 0.25 0.21 0.16
  Any Substance Abuse Services 0.02 0.02 0.76 0.01 0.00 0.11
  Any Mental Health Services 0.28 0.26 0.70 0.24 0.21 0.20
Educational Characteristics
 Enrolled in School 0.88 0.90 0.20 0.88 0.90 0.20
 Graduated 0.05 0.05 0.66 0.05 0.05 0.66
 Grade (self-reported) 9.0 9.0 0.85 9.0 9.0 0.85
 Days Absent (if enrolled) 10.8 9.5 0.17 10.8 9.5 0.17
 Grade Point Average (if non-missing) 2.50 2.50 0.96 2.50 2.50 0.96
 Ever Suspended (if enrolled) 0.11 0.14 0.37 0.11 0.14 0.37
P-value on F-test of treatment-control comparison for all non-behavioral health baseline characteristics 0.52 0.28
P-value on F-test of treatment-control comparison for all behavioral health baseline characteristics 0.96 0.28

Note: Grade information is self-reported, N=4482; school absence (N=2336), GPA (N=2228), and suspension (N=2337) measures reflect only youth enrolled in public, non-charter schools in the School District of Philadelphia for whom data was non-missing.One WorkReady participant was missing race information. Enrolled and graduated are mutually exclusive. The difference column shows the p-value from the test that treatment and control means are equal, adjusting for randomization block. Behavioral health services, including substance abuse and mental health services, are held in a separate data set to maintain confidentiality of HIPAA-covered data and are thus a separate F-test from other baseline characteristics.

Table A15:

WorkReady Outcomes in the First Year After Randomization, by Cohort

Intent-to-Treat Local Average Treatment Effect
2018 2017 2018 CM 2017 CM Test of Difference 2018 2017 2018 CCM 2017 CCM Test of Difference
Juvenile Incarceration 0.001 (0.005) −0.007** (0.004) 0.008 0.017 0.155 0.002 (0.008) −0.029** (0.014) 0.009 0.028 0.054
Receipt of Juvenile Justice Services 0.005 (0.005) −0.005 (0.004) 0.006 0.015 0.108 0.009 (0.005) −0.019 (0.014) 0.006 0.028 0.091
Total Number of Arrests 0.001 (0.008) −0.014** (0.006) 0.012 0.033 0.106 0.003 (0.013) −0.056** (0.024) 0.012 0.072 0.030
 Number of Violent Arrests 0.001 (0.007) −0.003 (0.004) 0.008 0.013 0.614 0.001 (0.011) −0.012 (0.016) 0.014 0.018 0.475
 Number of Property Arrests 0.002 (0.004) −0.003 (0.003) 0.003 0.010 0.282 0.004 (0.007) −0.013 (0.013) 0.000 0.024 0.242
 Number of Drug Arrests −0.001 (0.001) −0.003 (0.002) 0.002 0.004 0.440 −0.002 (0.002) −0.012 (0.008) 0.002 0.012 0.236
 Number of Other Arrests 0.000 (0.000) −0.005*** (0.002) 0.000 0.006 0.014 −0.001 (0.001) −0.019** (0.007) 0.001 0.018 0.011
Graduated or Still in School 0.005 (0.012) −0.008 (0.008) 0.945 0.944 0.374 0.008 (0.020) −0.031 (0.031) 0.956 0.983 0.292
Becomes a Parent 0.001 (0.001) 0.001 (0.004) 0.000 0.015 0.947 0.002 (0.002) 0.003 (0.017) 0.000 0.001 0.952
Receipt of Child Protective Services 0.001 (0.004) −0.008* (0.004) 0.006 0.022 0.163 0.001 (0.007) −0.030* (0.016) 0.003 0.023 0.084
Any Stay in a Shelter 0.003 (0.002) 0.002 (0.002) 0.000 0.002 0.669 0.005 (0.004) 0.007 (0.007) 0.000 0.000 0.804
Receipt of Behavioral Health Services 0.000 (0.017) −0.015 (0.010) 0.110 0.113 0.453 0.000 (0.031) −0.057 (0.039) 0.104 0.110 0.252
 Substance Abuse Services 0.001 (0.003) −0.004 (0.003) 0.003 0.010 0.261 0.002 (0.006) −0.016 (0.013) 0.002 0.014 0.211
 Mental Health Services −0.001 (0.017) −0.013 (0.010) 0.110 0.109 0.534 −0.001 (0.031) −0.050 (0.039) 0.105 0.103 0.327

Note: 2018 N=1780, 2017 N=2717. Enrolled/graduated excludes pre-program graduates and those unmatched to education records; N=3858. Table shows estimated intent-to-treat (ITT) and local average treatment effects (LATE), controlling for baseline covariates and randomization block. Coefficients estimated in single regression with interactions; linear combination of estimates and standard errors displayed to show net effect for each group. CM is control mean; CCM is control complier mean, rounded to 0 when estimate is negative. Robust standard errors in parentheses, clustered by person as the same person can appear in both cohorts.

*

p<0.1

**

p<0.05

***

p<0.01

Table A16:

WorkReady Outcomes in the First Year After Randomization, by Gender

Intent-to-Treat Local Average Treatment Effect
Male Female Male CM Female CM Test of Difference Male Female Male CM Female CM Test of Difference
Juvenile Incarceration −0.011* (0.006) −0.001 (0.003) 0.029 0.005 0.123 −0.038* (0.021) −0.003 (0.007) 0.041 0.008 0.105
Receipt of Juvenile Justice Services −0.006 (0.006) 0.000 (0.003) 0.025 0.005 0.320 −0.020 (0.006) 0.001 (0.008) 0.037 0.008 0.307
Total Number of Arrests −0.019* (0.010) −0.005 (0.005) 0.049 0.015 0.201 −0.063* (0.035) −0.013 (0.012) 0.078 0.028 0.162
 Number of Violent Arrests −0.003 (0.007) −0.002 (0.003) 0.019 0.007 0.838 −0.010 (0.023) −0.004 (0.008) 0.025 0.011 0.801
 Number of Property Arrests −0.005 (0.005) 0.000 (0.003) 0.013 0.005 0.422 −0.016 (0.017) 0.000 (0.008) 0.017 0.009 0.396
 Number of Drug Arrests −0.006* (0.004) 0.000 (0.001) 0.009 0.000 0.075 −0.021* (0.012) 0.000 (0.002) 0.021 0.000 0.074
 Number of Other Arrests −0.005 (0.003) −0.003** (0.001) 0.008 0.002 0.575 −0.016 (0.010) −0.008** (0.004) 0.014 0.008 0.461
Graduated or Still in School −0.005 (0.012) −0.004 (0.008) 0.936 0.950 0.932 −0.017 (0.036) −0.011 (0.021) 0.968 0.973 0.884
Becomes a Parent 0.009** (0.004) −0.005 (0.005) 0.003 0.017 0.028 0.029** (0.013) −0.012 (0.012) 0.000 0.014 0.021
Receipt of Child Protective Services −0.007 (0.005) −0.005 (0.005) 0.017 0.019 0.741 −0.023 (0.015) −0.013 (0.012) 0.017 0.012 0.614
Any Stay in a Shelter 0.000 (0.002) 0.003 (0.002) 0.002 0.001 0.188 −0.065 (0.006) 0.009* (0.006) 0.000 0.000 0.226
Receipt of Behavioral Health Services −0.025* (0.013) −0.002 (0.011) 0.116 0.110 0.171 −0.089* (0.047) −0.004 (0.031) 0.126 0.098 0.125
 Substance Abuse Services −0.011** (0.004) 0.003 (0.003) 0.018 0.002 0.003 −0.041** (0.016) 0.009 (0.008) 0.046 0.000 0.003
 Mental Health Services −0.018 (0.013) −0.004 (0.011) 0.108 0.110 0.412 −0.065 (0.046) −0.011 (0.031) 0.102 0.105 0.331

Note: Male N=1780, Female N=2717. Enrolled/graduated excludes pre-program graduates and those unmatched to education records; N=3858. Table shows estimated intent-to-treat (ITT) and local average treatment effects (LATE), controlling for baseline covariates and randomization block. Coefficients estimated in single regression with interactions; linear combination of estimates and standard errors displayed to show net effect for each group. CM is control mean; CCM is control complier mean, rounded to 0 when estimate is negative. Robust standard errors in parentheses, clustered by person as the same person can appear in both cohorts.

*

p<0.1,

**

p<0.05,

***

p<0.01

Table A17:

OSC+ Outcomes in the First Year After Randomization, by Gender

Intent-to-Treat Local Average Treatment Effect
Male Female Male CM Female CM Test of Difference Male Female Male CM Female CM Test of Difference
Total Number of Arrests −0.037 (0.033) −0.014 (0.012) 0.330 0.069 0.49 −0.137 (0.120) −0.053 (0.046) 0.299 0.081 0.51
 Number of Violent Arrests 0.010 (0.012) 0.003 (0.006) 0.058 0.023 0.60 0.035 (0.043) 0.010 (0.022) 0.018 0.009 0.61
 Number of Property Arrests 0.006 (0.009) −0.003 (0.005) 0.027 0.016 0.34 0.023 (0.033) −0.013 (0.019) 0.023 0.017 0.33
 Number of Drug Arrests 0.023 (0.014) −0.005** (0.002) 0.076 0.005 0.21 −0.083 (0.051) −0.020** (0.009) 0.107 0.016 0.22
 Number of Other Arrests 0.030 (0.020) −0.008 (0.007) 0.169 0.026 0.29 −0.112 (0.075) −0.030 (0.026) 0.151 0.039 0.30
Graduated or Still in School 0.006 (0.010) −0.006 (0.008) 0.944 0.956 0.35 0.024 (0.039) −0.021 (0.029) 0.975 0.980 0.35

Note: Male N=2168, Female N=3237. Enrolled/graduated excludes pre-program graduates and those unmatched to education records; N=4077. Table shows estimated intent-to-treat (ITT) and local average treatment effects (LATE), controlling for baseline covariates and randomization block. Coefficients estimated in single regression with interactions; linear combination of estimates and standard errors displayed to show net effect for each group. CM is control mean; CCM is control complier mean, rounded to 0 when estimate is negative. Robust standard errors in parentheses.

*

p<0.1

**

p<0.05

***

p<0.01

Table A18:

OSC+ Outcomes in the First Year After Randomization, by Prior Arrest Status

Intent-to-Treat Local Average Treatment Effect
Arrested Not Arrested Arrested CM Not Arrested Test of Difference Arrested Not Arrested Arrested CM Not Arrested Test of Difference
Total Number of Arrests −0.091 (0.063) −0.009 (0.010) 0.592 0.056 0.20 −0.350 (0.245) −0.033 (0.036) 0.503 0.089 0.20
 Number of Violent Arrests −0.001 (0.020) 0.005 (0.005) 0.112 0.016 0.75 −0.004 (0.077) 0.020 (0.018) 0.012 0.019 0.76
 Number of Property Arrests 0.013 (0.018) −0.003 (0.003) 0.057 0.010 0.37 0.052 (0.071) −0.012 (0.012) 0.036 0.014 0.37
 Number of Drug Arrests −0.036 (0.026) −0.007** (0.003) 0.123 0.008 0.25 −0.139 (0.100) −0.025** (0.010) 0.151 0.029 0.25
 Number of Other Arrests −0.067* (0.038) −0.004 (0.005) 0.299 0.023 0.10 −0.259* (0.149) −0.017 (0.020) 0.305 0.027 0.11
Graduated or Still in School 0.010 (0.017) −0.004 (0.006) 0.906 0.964 0.45 0.040 (0.068) −0.014 (0.023) 0.975 0.979 0.45

Note: N=5405. "Arrested" refers to youth who had any record of arrest prior to randomization, regardless of number of arrests; Arrested N=1210, Not Arrested N=4195. Enrolled/graduated excludes pre-program graduates and those unmatched to education records; N=4077. Table shows estimated intent-to-treat (ITT) and local average treatment effects (LATE), controlling for baseline covariates and randomization block. Coefficients estimated in single regression with interactions; linear combination of estimates and standard errors displayed to show net effect for each group. CM is control mean; CCM is control complier mean, rounded to 0 when estimate is negative. Robust standard errors in parentheses.

*

p<0.1

**

p<0.05

***

p<0.01

Table A19:

WorkReady Outcomes in the First Year After Randomization, by Race

Intent-to-Treat Local Average Treatment Effect
Black Non-Black Black CM Non-Black Test of Difference Black Non-Black Black CCM Non-Black CCM Test of Difference
Juvenile Incarceration 0.007** (0.003) 0.002 (0.004) 0.017 0.003 0.089 −0.02** (0.010) 0.007 (0.015) 0.024 0.001 0.112
Receipt of Juvenile Justice Services −0.001 (0.004) −0.007 (0.005) 0.013 0.010 0.295 −0.002 (0.004) −0.024 (0.017) 0.016 0.026 0.258
Total Number of Arrests −0.009 (0.006) −0.016** (0.007) 0.030 0.021 0.415 −0.025 (0.018) −0.055** (0.023) 0.045 0.052 0.293
 Number of Violent Arrests −0.001 (0.004) −0.007 (0.005) 0.012 0.010 0.383 −0.003 (0.012) −0.022 (0.017) 0.016 0.019 0.335
 Number of Property Arrests −0.001 (0.003) −0.006* (0.003) 0.009 0.007 0.225 −0.002 (0.009) −0.022* (0.012) 0.010 0.022 0.179
 Number of Drug Arrests −0.003 (0.002) −0.002 (0.002) 0.004 0.002 0.780 −0.008 (0.005) −0.007 (0.006) 0.008 0.007 0.928
 Number of Other Arrests −0.005** (0.002) −0.001 (0.002) 0.005 0.002 0.170 −0.013** (0.005) −0.004 (0.005) 0.012 0.004 0.238
Graduated or Still in School −0.006 (0.008) 0.001 (0.010) 0.937 0.974 0.569 −0.017 (0.022) 0.004 (0.034) 0.963 1.001 0.606
Becomes a Parent 0.003 (0.004) −0.006 (0.006) 0.011 0.014 0.163 0.008 (0.010) −0.021 (0.019) 0.000 0.018 0.164
Receipt of Child Protective Services −0.009** (0.004) 0.009 (0.007) 0.021 0.009 0.025 −0.027** (0.010) 0.028 (0.024) 0.025 0.000 0.036
Any Stay in a Shelter 0.002 (0.002) 0.003 (0.003) 0.002 0.000 0.765 −0.038 (0.005) 0.009 (0.009) 0.000 0.000 0.672
Receipt of Behavioral Health Services −0.016 (0.010) 0.008 (0.017) 0.118 0.094 0.223 −0.047* (0.028) 0.028 (0.062) 0.116 0.070 0.268
 Substance Abuse Services −0.004 (0.003) 0.003 (0.006) 0.008 0.007 0.212 −0.013 (0.008) 0.012 (0.021) 0.014 0.000 0.252
 Mental Health Services −0.013 (0.010) 0.003 (0.017) 0.114 0.094 0.419 −0.038 (0.028) 0.009 (0.062) 0.107 0.089 0.478

Note: Black N=3513, Non-Black N=984. Enrolled/graduated excludes pre-program graduates and those unmatched to education records; N=3858. Table shows estimated intent-to-treat (ITT) and local average treatment effects (LATE), controlling for baseline covariates and randomization block. Coefficients estimated in single regression with interactions; linear combination of estimates and standard errors displayed to show net effect for each group. CM is control mean; CCM is control complier mean, rounded to 0 when estimate is negative. Robust standard errors in parentheses, clustered by person as the same person can appear in both

cohorts.

*

p<0.1

**

p<0.05

***

p<0.01

Table A20:

WorkReady Outcomes in the First Year After Randomization, by Age

Intent-to-Treat Local Average Treatment Effect
Under 16 Over 16 Under 16 CM Over 16 CM Test of Difference Under 16 Over 16 Under 16 CCM Over 16 CCM Test of Difference
Juvenile Incarceration −0.004 (0.004) −0.006 (0.004) 0.013 0.016 0.648 −0.010 (0.010) −0.021 (0.014) 0.015 0.025 0.505
Receipt of Juvenile Justice Services 0.001 (0.004) −0.006 (0.004) 0.010 0.015 0.217 0.004 (0.004) −0.019 (0.014) 0.015 0.021 0.186
Total Number of Arrests −0.011 (0.007) −0.009 (0.007) 0.028 0.028 0.817 −0.030* (0.018) −0.031 (0.023) 0.054 0.036 0.964
 Number of Violent Arrests 0.002 (0.005) −0.006 (0.005) 0.009 0.014 0.234 0.005 (0.012) −0.02 (0.016) 0.016 0.016 0.207
 Number of Property Arrests −0.004 (0.004) 0.000 (0.004) 0.010 0.007 0.368 −0.011 (0.010) 0.001 (0.012) 0.016 0.006 0.421
 Number of Drug Arrests −0.003** (0.002) −0.002 (0.002) 0.004 0.004 0.490 −0.009** (0.004) −0.005 (0.008) 0.009 0.005 0.657
 Number of Other Arrests −0.006*** (0.002) −0.002 (0.002) 0.005 0.004 0.222 −0.014*** (0.005) −0.007 (0.007) 0.012 0.008 0.368
Graduated or Still in School −0.001 (0.010) −0.008 (0.010) 0.943 0.946 0.646 −0.003 (0.024) −0.024 (0.030) 0.957 0.987 0.591
Becomes a Parent −0.001 (0.002) 0.002 (0.006) 0.004 0.018 0.578 −0.003 (0.006) 0.008 (0.019) 0.001 0.000 0.591
Receipt of Child Protective Services −0.006 (0.005) −0.004 (0.004) 0.024 0.012 0.760 −0.017 (0.014) −0.015 (0.013) 0.015 0.011 0.910
Any Stay in a Shelter 0.001 (0.001) 0.003 (0.003) 0.001 0.002 0.522 −0.033 (0.004) 0.01 (0.009) 0.000 0.000 0.451
Receipt of Behavioral Health Services −0.013 (0.012) −0.009 (0.012) 0.121 0.104 0.825 −0.034 (0.032) −0.033 (0.042) 0.085 0.139 0.990
 Substance Abuse Services 0.001 (0.003) −0.006* (0.004) 0.004 0.012 0.103 0.003 (0.009) −0.023* (0.013) 0.000 0.024 0.081
 Mental Health Services −0.013 (0.012) −0.007 (0.012) 0.120 0.099 0.730 −0.033 (0.032) −0.025 (0.042) 0.084 0.131 0.873

Note: Under-16 N=2223, Over-16 N=2274. Enrolled/graduated excludes pre-program graduates and those unmatched to education records; N=3858. Table shows estimated intent-to-treat (ITT) and local average treatment effects (LATE), controlling for baseline covariates and randomization block. Coefficients estimated in single regression with interactions; linear combination of estimates and standard errors displayed to show net effect for each group. CM is control mean; CCM is control complier mean, rounded to 0 when estimate is negative. Robust standard errors in parentheses, clustered by person as the same person can appear in both cohorts.

*

p<0.1

**

p<0.05

***

p<0.01

Table A21:

Effects on Standardized Index by Race/Ethnicity, Age, and Gender Subgroups

Intent-to-Treat Local Average Treatment Effect
Male Female Male CM Female CM Male Female Male CCM Female CCM
WorkReady (Philadelphia)

Younger
 Black −0.129* (0.068) 0.011 (0.043) 0.145 N=724 −0.046 N=1059 −0.394* (0.202) 0.024 (0.090) 0.274 −0.142
 Hispanic −0.141 (0.136) −0.022 (0.104) 0.147 N=87 −0.048 N=167 −0.311 (0.225) −0.135 (0.555) 0.12 0.215
 Other Races −0.035 (0.053) - −0.015 N=80 - −0.134 (0.156) - −0.058 -
Older
 Black −0.083 (0.109) −0.080** (0.037) 0.156 N=664 −0.061 N=1066 −0.304 (0.373) −0.252** (0.113) 0.117 0.116
 Hispanic 0.025 (0.073) −0.112 (0.090) −0.116 N=120 −0.123 N=162 0.105 (0.208) −0.331 (0.218) −0.296 0.049
 Other Races −0.007 (0.045) 0.078 (0.123) −0.144 N=105 −0.191 N=157 −0.015 (0.072) 0.383 (0.278) −0.32 −0.574

OSC+ (Chicago)

Younger
 Black −0.179* (0.099) −0.008 (0.020) 0.431 N=661 −0.134 N=973 −0.590* (0.326) −0.025 (0.064) 0.094 −0.132
 Hispanic −0.112 (0.088) −0.031 (0.023) 0.033 N=210 −0.160 N=300 −0.718 (0.560) −0.199 (0.162) 0.162 −0.057
 Other Races - - - - - - - -
Older
 Black 0.005 (0.069) −0.025 (0.015) 0.191 N=960 −0.145 N=1438 0.018 (0.238) −0.088 (0.055) 0.253 −0.119
 Hispanic −0.027 (0.056) −0.023 (0.017) −0.084 N=279 −0.182 N=422 −0.142 (0.287) −0.120 (0.089) 0.376 −0.111
 Other Races 0.032 (0.053) - 0.477 N=43 0.173 (0.218) - −1.197 -

Note: Table shows the effect of each program on a standardized index of the outcomes that have significant changes in the main estimates: incarceration, total arrests, and receipt of child protective services (WorkReady) and drug and other arrests (OSC+). "Younger" indicates youth under the age of 16 for WorkReady and under the age of 17 for OSC+; "Older" indicates youth 16+ and 17+, respectively. A horizontal dash indicates there was no variation in the outcome within that cell. CM is control mean for each group, and CCM is control compiler mean. Robust standard errors in parentheses, clustered on individual for WorkReady.

*

p<0.1

**

p<0.05

***

p<0.01

Table A22:

WorkReady Treatment Effect on Combined Index by Predicted Risk Level

Panel A: Intent to Treat Effect, Index
Full Sample −0.045* (0.026)
Panel B: Intent to Treat Effect by Predicted Risk Level
Predicted Risk Level Repeated Split Sample Leave One Out CM First Stage
Low −0.011 (0.020) 0.038 (0.026) −0.184 0.444
Medium −0.019 (0.024) −0.069* (0.040) −0.075 0.280
High −0.099 (0.063) −0.115 (0.072) 0.243 0.288

Note: N=4497. Estimation from the Abadie, Chingos, and West (2018) procedure. CM is the control mean for each group, assigned in the leave one out estimation. Table shows the effect of WorkReady on a standardized index of all the socially costly outcomes available in the WorkReady data: juvenile incarceration, juvenile justice services, the 4 types of arrests separately, excluding total since that is the sum of the rest, parenthood, an indicator for child protective services, and an indicator for using a housing shelter. Each variable is standardized on the control group, averaged together, then re-standardized so that the standard deviation of the index is 1.

*

p<0.1

**

p<0.05

***

p<0.01

Table A23:

OSC+ Treatment Effect on Combined Index by Predicted Risk Level

Panel A: Intent to Treat Effect, Index
Full Sample −0.023 (0.022)
Panel B: Intent to Treat Effect by Predicted Risk Level
Predicted Risk Level Repeated Split Sample Leave One Out CM First Stage
Low −0.014 (0.010) −0.022* (0.010) −0.208 0.263
Medium −0.017 (0.019) −0.009 (0.019) −0.163 0.265
High −0.012 (0.063) −0.013 (0.063) 0.374 0.273

Note: N=5405. Estimation from the Abadie, Chingos, and West (2018) procedure. CM is the control mean for each group, assigned in the leave one out estimation. Table shows the effect of OSC+ on a standardized index of the 4 types of arrests separately, excluding total since that is the sum of the rest. Each variable is standardized on the control group, averaged together, then re-standardized so that the standard deviation of the index is 1.

*

p<0.1

**

p<0.05

***

p<0.01

Table A24:

Baseline Statistics for Always Takers, Compliers, and Never Takers

WorkReady
OSC+
Always Takers Compliers Never Takers Always Takers Compliers Never Takers
Estimated Proportion 0.161 0.340 0.498 0.204 0.265 0.530
Demographics
 Age 15.6** 15.4 15.8*** 17.4 17.3 17.4
 Male 0.42** 0.34 0.43** 0.38 0.42 0.40
 Black 0.78 0.82 0.76** 0.83 0.82 0.68***
 Hispanic 0.11 0.10 0.14* 0.14 0.16 0.29***
 White 0.06** 0.02 0.05* 0.01 0.01 0.01
 Other Race 0.05 0.07 0.05 0.02 0.01 0.02
 Is a Parent 0.011 0.007 0.024
Contact with Justice System
 Ever Incarcerated as a Juvenile 0.016 0.014 0.027
 Ever Received Juvenile Justice Services 0.014 0.024 0.024
 Ever Arrested 0.041 0.040 0.046 0.262 0.219 0.212
 Total Number of Prior Arrests 0.046 0.052 0.063 0.795** 0.478 0.628
  Violent 0.023 0.037 0.023 0.261*** 0.136 0.182
  Property 0.016 0.007 0.026 0.118 0.091 0.091
  Drug 0.000 0.001 0.005 0.103 0.061 0.074
  Other 0.007 0.007 0.01 0.314* 0.190 0.281
Education
 Enrolled in School 0.90 0.92 0.83*** 0.7* 0.75 0.76
 Graduated 0.05 0.03 0.11*** 0.27 0.24 0.23
 Grade 9.6*** 9.3 9.7*** 10.8 10.7 10.8
 Days Absent 14.4 14.2 21.3*** 24.1 23.9 24.1
 Grade Point Average 2.55 2.50 2.34 2.35 2.35 2.44
 Ever Suspended 0.07*** 0.17 0.16 0.24 0.22 0.17*
Receipt of Social Services
 Ever Received Child Protective Services 0.12 0.15 0.15
 Ever Stayed in a Shelter 0.04 0.06 0.03*
 Any Behavioral Health Service 0.21* 0.27 0.27
  Any Substance Abuse Services 0.00 0.01 0.02
  Any Mental Health Services 0.21 0.27 0.27

Note: WorkReady N=4497; OSC+ N=5405. Always-taker means estimated from participants in the control group; never-taker means estimated from non-participants in the treatment group; and complier means backed out from the other estimates.

Compiler means do not adjust for randomization strata (see text for details). Stars indicate whether each group is significantly different from the compilers using bootstrapped standard errors. Other variable details and list of Ns by variable the same as in Table 2.

Table A25:

Control Group Standard Deviations by Year After Randomization

WorkReady (Philadelphia) OSC+ (Chicago)
Year One
 Juvenile Incarceration 0.12
 Receipt of Juvenile Justice Services 0.11
 Total Number of Arrests 0.19 0.71
  Violent Arrests 0.12 0.22
  Property Arrests 0.09 0.18
  Drug Arrests 0.06 0.28
  Other Arrests 0.07 0.41
 Graduated or Still in School 0.23 0.22
 Becomes a Parent 0.11
 Receipt of Child Protective Services 0.13
 Any Stay in a Shelter 0.04
 Receipt of Behavioral Health Services 0.32
  Substance Abuse Services 0.09
  Mental Health Services 0.31
Year Two
 Juvenile Incarceration 0.13
 Receipt of Juvenile Justice Services 0.12
 Total Number of Arrests 0.20 0.56
  Violent Arrests 0.13 0.19
  Property Arrests 0.10 0.16
  Drug Arrests 0.05 0.16
  Other Arrests 0.07 0.36
 Graduated or Still in School 0.30 0.30
 Becomes a Parent 0.11
 Receipt of Child Protective Services 0.05
 Any Stay in a Shelter 0.32
 Receipt of Behavioral Health Services 0.30
  Substance Abuse Services 0.10
  Mental Health Services 0.29
Year Three
 Total Number of Arrests 0.51
  Violent Arrests 0.13
  Property Arrests 0.14
  Drug Arrests 0.20
  Other Arrests 0.33

Note: WorkReady year one N=2711 (N=2346 for Graduated or Still in School); year two N=2056 (N=1764 for Graduated or Still in School). OSC+ N=2911 (N=2192 for Graduated or Still in School).

Footnotes

1

E.g., Perry Preschool, LIFE and CEO jobs programs, and Tennessee STAR among others.

2

The Philadelphia study registration includes a pre-analysis plan detailing primary and secondary hypotheses, as well as methods to address multiple testing concerns. The OSC+ study was pre-registered but without a pre-analysis plan, largely because the outcome definitions follow the prior studies of OSC+ exactly, limiting the scope for any potential data mining. I nonetheless perform similar multiple testing adjustments, described in the methods section below.

3

The main text focuses on criminal justice impacts, since those are the outcomes most consistently measured across settings, providing the most opportunity to analyze impact heterogeneity across studies. The appendix also reports education impacts, which are generally null but suffer from considerable missing data due to high charter school enrollment in Philadelphia. Family and health measures from social service records available only in Philadelphia, which have not been measured elsewhere, are also in the appendix. There are promising indications that participation may decrease the need for child protective services, and perhaps substance abuse and mental health services, especially among boys and Black youth. But estimates are too imprecise to survive adjustments for multiple hypothesis testing, and some outcomes are quite rare in the sample. Because these were pre-specified outcomes, I include these results in the individual treatment heterogeneity analysis that follows. But the imprecision means that confirming SYEPs’ impacts on family and individual health and well-being should be a priority for future work.

4

This pattern is not driven by differences in take-up rates; the relationship between intent-to-treat effects and control means looks very similar. Nor is it a by-product of selecting statistically significant outcomes; the pattern remains when accounting for each estimate’s variance and when including all main effects.

5

Changes to program details are common in SYEPs; OSC+ has continued to change since 2015. Although there was also a 2013 OSC+ study, I focus here on comparing 2015 to the original 2012 cohort. The 2013 program involved explicitly different eligibility criteria, recruiting only males, some from the criminal justice system, to test for effect heterogeneity. So the 2012 cohort, where recruitment worked like the current study, is the most useful comparison for isolating what happens when the same approach scales.

6

In practice, program providers were quite resistant to removing the additional mentorship, so they replaced the mentors with “adult supervisors” who provided less personalized and intensive support, but who were nonetheless available as needed. While this makes the test of the difference slightly less informative, site observations suggested that there was still a difference in the amount of adult support offered across treatment arms, just less of a difference than was originally intended. Because the program only officially provided mentors to half the sample, the increased scale does less to identify the challenge of recruiting additional mentors as the program grew than it does to identify the challenge of finding new job placements (from 2012 to 2015, mentors were provided to an additional 300 youth). One other program change is worth highlighting, since it complicated recruitment: The program obtained a waiver against the city’s minimum wage increase, so youth were still paid $8.25 per hour, compared to a concurrent shift to a $10 minimum wage in the regular labor market.

7

See https://www.civicleadershipfoundation.org/curriculum for details. This curriculum was quite different from prior OSC+ studies, where the programming focused on developing socio-emotional skills like self regulation, goal setting, and perspective-taking.

8

All three models focused on developing “21st-Century Workforce Skills” and offered an hourly wage, but they varied in how like a private-sector summer job they were. Both the number of providers and the program models have continued to evolve since the study took place.

9

As part of the study, the research team conducted open-ended interviews with 18 study participants (13 in the treatment group) as well as field observations. One of the core conclusions of that work was how much the experience of the program varied across individuals based on job assignment; timing, content, and instructional details of professional development activities; and relationships with supervisors. One might hypothesize that this kind of treatment variation would be likely to generate effect heterogeneity, especially given other work finding that two-thirds of examined multi-site workforce development and education studies show significant variation in treatment effects across sites, and variation is more likely when program details are not highly codified (Weiss et al., 2017).

10

As is true in all studies using administrative police records, arrests are an imperfect measure of true offending behavior. They capture both police and individual choices. This generates both a downward bias in the measurement of crime, since many offenses do not result in arrest, as well as an upward bias, since not all arrests are for something an individual actually did. There is plenty of evidence that these biases do not affect all types of youth the same way, and likely vary systematically by race, neighborhood, and other characteristics of both the individuals and the arresting officers (Goncalves and Mello, forthcoming; Hinton et al., 2018; Ridgeway and MacDonald, 2009). One key benefit of the randomized design is that the study does not need to assume that arrests are a perfect measure of underlying criminal behavior; it is clear they are not. But the biases in the data-generating process affect both the treatment and control groups equally, and treatment effects measure the difference between the groups. Mismeasurement in the dependent variable might attenuate estimated treatment effects, but it does not bias them. Rather, the key assumption is that the treatment does not affect the probability of being arrested conditional on committing (or not committing) a crime. This is not entirely trival, since treatment could teach youth to interact with police more constructively, thus avoiding arrests that would otherwise have taken place. But even if that is driving some of the estimated treatment effects, the fact that criminal justice system involvement is so damaging to future individual and family outcomes (e.g., Aizer and Doyle, 2015; Charles and Luoh, 2010; Dobbie et al., 2018; Holzer et al., 2006; Mueller-Smith, 2015) means there still a large social benefit to reducing arrests, even if some of the change is not driven by changes in underlying crime.

11

Given the low baseline means of some outcomes, estimates of CCMs for indicator or count variables are sometimes negative due to the sampling error in the LATE. I round these cases to 0.

12

In theory, including provider-by-treatment interactions and testing whether they are all equal is another option for this test. But estimating the separate provider fixed effects adds uncertainty from the estimation of the fixed effects, reducing the power to distinguish effects across providers. Since there is random variation in provider assignment, the assumptions for random effects are met by construction; provider assignment is not correlated with any other covariate, conditional on strata. So I focus on the random effect approach to aid power.

13

For WorkReady, 2 of 26 tests have p<0.05 , about what would be expected if all covariates were independent (which they are not, since some are sums of other covariates shown). The joint tests at the bottom of the table exclude variables that are linear combinations of the others. It is worth noting that one of the chance imbalances in the WorkReady study is on the primary pre-specified outcome, with the treatment group having fewer pre-program violent-crime arrests p=0.01 . Although imbalance on this outcome is unfortunate, the difference is controlled for by including baseline covariates in all outcome regressions. Additionally, as discussed below, the overall level of violent-crime arrests (and in fact, all arrests) ended up being lower than expected at the outset. So the results separated by crime type are less informative than expected at the time of pre-specification.

14

As a rough benchmark for how big the target population of those who might benefit from an SYEP that reduces arrests is relative to the 5,405 applicants, around 13,000 people under age 17 were arrested in Cook County, where Chicago is located, in 2015 (Gleicher, 2017). The ACS reports a little over 150,000 people between 15 and 17 living in Chicago. So while there is likely continued scope for program expansion without major changes in the participant population, a universal program for everyone under 17 would likely dramatically change the criminal justice involvement of the participants (as we see in Philadelphia).

15

Although the study sample is not perfectly representative of the WorkReady program as a whole (applicants without pre-existing relationships with providers are over-represented, see Appendix A), it does reflect a broad subset of the program, with youth at almost all providers participating.

16

Complier means are calculated with kappa weights that account for the varying treatment probabilities across strata, per Abadie (2003).

17

As shown in the appendix, there is no change in school persistence, the best-measured education outcome. There are some suggestions of improvements in family and health outcomes that could be related to income, time use, or personal skills like self-efficacy. Child protective service receipt shows a marginally significant decline, especially among Black youth, and behavioral health services decline among boys. But these results are sensitive to adjustments for multiple testing, so they serve to generate hypotheses for exploration in future work rather than establish clear impacts. Appendix Table A8 through A11 show no clear improvements in other education outcomes, regardless of how missing data are imputed. If anything, the point estimates tend to be negative; see Appendix F.4 for discussion.

18

The drug arrest result in Philadelphia is exactly at the standard significance threshold, p = 0.100, which I treat as marginally significant.

19

Given the similarity in number of arrests prior to random assignment, the drop in violent crime involvement among controls does not seem to be due to a fundamental change in the population served. It likely is due partly to the large secular drops in violent-crime rates citywide over time, as well as changes in policing that decreased arrest rates among youth (see, e.g., https://home.chicagopolice.org/statistics-data/statistical-reports/annual-reports/). So it is possible that the lack of a significant treatment effect on violence is because the lower occurrence and recording of violent events makes behavioral changes harder to detect in the data.

20

In practice, all else might not be equal. For example, the cost of serving different types of individuals may vary, such that policymakers would need to balance bigger benefits with higher costs. Or the social costs of a behavior could also vary by θ , as would be the case if policymakers cared about the distributional impacts of a program. I return to this point in the discussion below.

21

One interpretation of the effects of active labor market programs more broadly is that they follow this kind of pattern. Many short-term programs that target those with the highest barriers to employment, like the long-term unemployed or those returning from prison, have historically had mixed to no effects (Berk et al., 1980; D. Bloom, 2010; Card et al., 2017; Cave et al., 1993; Doleac et al., 2020; Heinrich et al., 2013; MDRC, 1980). Some, like the JTPA, even have adverse crime effects on those already at elevated risk of crime (H. S. Bloom et al., 1997). Many of the more recent programs that do have large, positive employment effects perform purposeful screening upfront, potentially finding those close to the margin of success but still at risk of a bad outcome as a way to ensure programs effectively help participants cross the relevant margin (Fein and Hamadyk, 2018; Roder and Elliott, 2020; Schaberg and Greenberg, 2020). Within-study heterogeneity is also consistent with this idea, as in the classic Friedlander (1988) finding that the biggest responses occurred among those who were in a middle-risk tier, neither too well off nor too disadvantaged at baseline.

22

Note that this analysis was not pre-specified, so should be considered exploratory.

23

In this context, a joint analysis across studies is logistically impossible. The data are not all held in the same place; Chicago and Philadelphia data are on separate institution’s servers due to data agreement limitations. Even within Philadelphia, the behavioral health data are completely separate, with very limited Xs available, to comply with HIPAA regulations. So I can not include behavioral health data in this exercise. I limit the index to statistically significant main effects to avoid diluting the index with additional noise, but Appendix I.1 shows that I get a generally similar pattern of results when using an index of all the available measures of socially costly outcomes. The pattern is clearer in Philadelphia than in Chicago, but with less precision than the results in the main text.

24

I exclude drug and other arrests since they are already part of the total arrest outcome.

25

I rely on the user-written Stata command estrat for this analysis, with a small adjustment to the code to ensure that results are exactly replicable within the same seed.

26

This pattern raises the question of whether doing the reverse exercise—estimating treatment effects across the distribution of how likely someone is to take up the program, as in the marginal treatment effects (MTE) literature (Brinch et al., 2017; James J Heckman and Vytlacil, 2005; Kowalski, 2018a; Mogstad et al., 2018; Walters, 2018), rather than across the distribution of Y^0 ―could uncover additional insights. Doing this across different studies rather than within a single study may provide particular insight into external validity (Kowalski, 2018b). In practice, however, the key assumption in this literature that MTEs are linear or monotonic in the propensity to take-up is unlikely to hold in this setting; see Appendix Section I.2 for further discussion about how this literature can inform the current analysis.

27

See the discussion in Heller and Bhanot (2021), a companion project which explores different barriers to take-up. That paper uses a separate “nudge” experiment, along with non-experimental variation in the level of administrative enrollment support, to demonstrate which strategies help the more responsive population overcome the higher barriers to participation they face.

28

The power issue may be part of the reason that Davis and Heller (2020) did not find statistically significant treatment heterogeneity on crime outcomes within previous studies of OSC+.

29

Table 1 summarizes the details of the different programs, and Appendix I.3 lists which estimates are included in each panel. To ease interpretation, I focus only on the outcomes where negative effects are desirable. This includes crime, incarceration, and mortality effects in Davis and Heller (2020), Gelber et al. (2016), and Modestino (2019). The Boston paper reports only the ITT effects; I use the reported take-up rate to scale the ITT, backing out the LATE. One could also do this exercise with the positive educational effects in Leos-Urbel (2014), Schwartz et al. (2015), and Modestino and Paulsen (2019), but I avoid that here in part because those results differ from the education effects in this paper, while the other results are more consistent across studies.

30

Standardizing the outcomes could be a partial solution. But since so many outcomes are indicator variables, focusing on mean changes rather than standard deviation changes is more directly interpretable. And importantly, switching to standard deviation units would dramatically limit our ability to include estimates from other studies, since none of them report outcome standard deviations, either overall or by subgroup. Regardless, since all outcomes are either indicators or counts, the units are still roughly comparable: An increase of 0.01 for any of the outcomes reflects 1 extra occurence of a negative incident in a group of 100 youth offered the program.

31

Both here and below, point estimates in the plot are not always independent. For example, some arrest categories are included in the total arrest category, and some subgroups compose part of the overall estimates from the same study. But I avoid plotting estimates that are purely linear combinations of each other (e.g., if I show effects for OSC+ 2012 and 2013 studies separately, I do not also show the pooled estimate).

32

I am cautious not to over-interpret any given subgroup estimate given the number of hypothesis tests in all these interaction effects. A few findings may merit further attention in future work powered to distinguish subgroup effects. In Philadelphia, males have a proportionally huge and statistically significant decline in substance abuse treatment, a 1.1 percentage point ITT decline relative to a control mean of 1.8 percent, and a significant overall drop in combined behavioral health services. Black youth show a large and significant drop in child protective services (ITT = 0.9 percentage points, a 43 percent decline). Males also show the only significant adverse effect, a 0.9 percentage point increase in parenthood, tripling relative to the control mean. While it is certainly possible that the program increases confidence and income in a way that increases risky sexual activity, it is also true that there is more of a margin for increases in reporting of childbirth for fathers than for mothers (who have a negative but not significant point estimate). The 2018 cohort of WorkReady consistently faces a floor effect across many outcomes; their program impacts are often significantly more positive than the 2017 cohort, because their baseline rates are so low that there was no room for a decline. This is consistent with the overall lesson from this analysis: that targeting youth at higher risk of these outcomes will generate larger effects.

33

These more-responsive youth also have higher employment rates in the control group. So this is consistent with the main pattern above of bigger responses for higher control means. But it has different implications for outcomes like employment that have social benefits rather than social costs.

34

This is both because of the compliance issues and because program providers were quite resistant to removing the additional mentorship. They replaced the mentors with “adult supervisors” who purposefully provided less personalized and intensive support, but who were nonetheless available as needed. Site observations suggested that there was still a difference in the amount of adult support offered across treatment arms, but the difference in practice was not as stark as originally intended.

35

Both the number of providers and the program models have continued to evolve since the study took place.

36

Note that there is a research report on a Philadelphia summer jobs program implemented in the late 1990s, but that was a different program was run by a different organization (McClanahan et al., 2004). Although that study was intended to be a randomized controlled trial, the appendix reports dropping all control youth who actually participated, then estimating the difference between participants and non-participants (not treatment and controls). So to my knowledge, there are no experimental estimates of any Philadelphia summer jobs programs.

37

There is one case where a student is enrolled in a District school prior to randomization but has absence data from a charter school. This results in the student’s absence record being missing despite having non-missing grade and suspension information. The student graduated prior to the program, so this does not change any results. But it does explain why the baseline N for suspensions and attendance do not match exactly.

38

Arrests are the most useful outcome for this exercise, since they are associated with an exact date. Many of the other outcomes in the paper are captured only annually.

39

About 47 percent of the sample was randomized 5/1/17, with 27 percent randomized on 5/12/17 and 25 percent on 6/3/18. The program ended on 8/11/17 and 8/12/18, meaning that about a quarter of the sample finished the program just over 2 months after randomization, about a quarter finished at about 3 months, and about half the sample finished at about 3.5 months. There were also 40 individuals randomized early at one provider on 4/7/17, but since there are so few of them, the figure does not include a separate line just past 4 months.

40

Because I do not have separate measures for who actually worked without a mentor and who received real mentorship, I can not instrument separately for the two types of activities. As a result, the table focuses just on the ITT effects by arm. Take-up was slightly higher in the mentor arm, with a first stage of 0.29 in the mentor group and 0.24 in the job-only group. But the difference is not statistically significant (p = 0.25).

41

The significant variation is for drug-crime arrests in Chicago. One out of 17 is about what we would expect by chance.

42

This is a departure from the pre-analysis plan for Philadelphia, driven by the unexpected amount of missing data from charter schools. It is, however, consistent with our prior analysis of OSC+ in Davis and Sara B. Heller (2020). I show other outcomes without multiple testing adjustments in Appendix Section F.4.

43

This analysis is in the vein of Kowalski (2018b), which performs a similar exercise with baseline ER visits to argue that linearity is likely to hold in her empirical setting. I use her Stata code, mtebinary, for this analysis, with some trivial edits to keep additional digits and to write out results to files outside of Stata. It is worth noting that the complier means reported here differ very slightly from those in the main text. This is because the main text uses Abadie’s kappa weighting to deal with the stratified random assignment (Abadie, 2003), while the MTE literature has not, to my knowledge, worked through the proofs of how to adjust for randomization strata with varying treatment probabilities. As such, the complier means and difference tests in Table A24 are unadjusted for strata, making them only approximate. Given how substantively trivial the differences in complier means are with and without adjustment, however, I stick with reporting the unstratified version here, so that it reflects the same inputs as those in the MTE calculation I mention below.

44

To deal with stratification here, I use the residuals from a regression of the index on strata fixed effects for the MTE analysis. Given the bootstrapping, this should not affect inference. But further theoretical work on the appropriate way to adjust for stratification and what, if any, additional assumptions are needed would be useful. Substantively, conclusions do not change if I use the unadjusted index and ignore stratification entirely.

45

In particular, I include all the crime and health outcomes (excluding totals across categories that are reported separately) for WorkReady and OSC+ 2012, 23013, and 2015 (from the current paper and Davis and Sara B. Heller (2020)). I can not include the insignificant crime effects from Boston (Modestino, 2019), because the paper does not report exact control means for insignificant results. In NYC, the two main effects on incarceration and mortality are already included, since they are statistically significant. I do not include insignificant subgroup estimates, since there would be hundreds of them.

References

  1. Abadie Alberto (2003). “Semiparametric instrumental variable estimation of treatment response models”. Journal of econometrics 113.2, pp. 231–263. [Google Scholar]
  2. Abadie Alberto, Chingos Matthew M., and West Martin R. (2018). “Endogenous Stratification in Randomized Experiments”. Review of Economics and Statistics 100(4), pp. 567–580. [Google Scholar]
  3. Aizer Anna and Doyle Joseph J. (2015). “Juvenile Incarceration, Human Capital and Future Crime: Evidence from Randomly-Assigned Judges”. The Quarterly Journal of Economics, pp. 759–804.
  4. Allcott Hunt (2015). “Site selection bias in program evaluation”. The Quarterly Journal of Economics 130.3, pp. 1117–1165. [Google Scholar]
  5. Anderson Michael L. (2008). “Multiple Inference and Gender Differences in the Effects of Early Intervention: A Reevaluation of the Abecedarian, Perry Preschool, and Early Training Projects”. Journal of the American Statistical Association 103(484), pp. 1481–1495. [Google Scholar]
  6. Athey Susan and Imbens Guido (2017). “The Econometrics of Randomized Experiments”. In: Handbook of Field Experiments Ed. by Banerjee Abhijit Binayak and Duflo Esther. North-Holland. [Google Scholar]
  7. Benjamini Yoav and Hochberg Yosef (1995). “Controlling the False Discovery Rate: A Practical and Powerful Approach to Multiple Testing”. Journal of the Royal Statistical Society.Series B (Methodological) 57(1), pp. 289–300. [Google Scholar]
  8. Berk RA, Lenihan KJ, and Rossi PH (1980). “Crime and Poverty - Some Experimental Evidence from Ex-Offenders”. American Sociological Review 45(5), pp. 766–786. [Google Scholar]
  9. Bhatt Monica P., Guryan Jonathan, Ludwig Jens, and Shah Anuj K. (2021). “Scope Challenges to Social Impact”. NBER Working Paper No. 28406
  10. Bloom Dan (2010). “Transitional Jobs: Background, Program Models, and Evaluation Evidence”.
  11. Bloom Howard S., Orr Larry L., Bell Stephen H., Cave George, Doolittle Fred, Lin Winston, and Bos Johannes M. (1997). “The Benefits and Costs of JTPA Title II-A Programs: Key Findings from the National Job Training Partnership Act Study”. The Journal of Human Resources 32(3), pp. 549–576. [Google Scholar]
  12. Brinch Christian N, Mogstad Magne, and Wiswall Matthew (2017). “Beyond LATE with a discrete instrument”. Journal of Political Economy 125.4, pp. 985–1039. [Google Scholar]
  13. Card David, Kluve Jochen, and Weber Andrea (2017). “What Works? A Meta Analysis of Recent Active Labor Market Program Evaluations”. Journal of the European Economic Association 16(3), pp. 894–931. [Google Scholar]
  14. Cave George, Bos Hans, Doolittle Fred, and Toussaint Cyril (1993). “JOBSTART: Final Report on a Program for School Dropouts”. MDRC
  15. Charles Kerwin Kofi and Luoh Ming Ching (2010). “Male Incarceration, the Marriage Market, and Female Outcomes”. The Review of Economics and Statistics 92(3), pp. 614–627. [Google Scholar]
  16. Davis Jonathan M.V., Guryan Jonathan, Hallberg Kelly, and Ludwig Jens (2017). “The Economics of Scale-Up”. NBER Working Paper No. 23925
  17. Davis Jonathan M.V. and Heller Sara B. (Oct. 2020). “Rethinking the Benefits of Youth Employment Programs: The Heterogeneous Effects of Summer Jobs”. The Review of Economics and Statistics 102(4), pp. 664–677. [Google Scholar]
  18. Dobbie Will, Goldin Jacob, and Yang Crystal S. (2018). “The Effects of Pretrial Detention on Conviction, Future Crime, and Employment: Evidence from Randomly Assigned Judges”. American Economic Review 108(2), pp. 201–240. [Google Scholar]
  19. Doleac Jennifer L., Temple Chelsea, Pritchard David, and Roberts Adam (2020). “Which prisoner reentry programs work? Replicating and extending analyses of three RCTs”. International Review of Law and Economics 62. [Google Scholar]
  20. Fein David and Hamadyk Jill (2018). “Bridging the Opportunity Divide for Low-Income Youth:Implementation and Early Impacts of the Year Up Program”. Pathways for Advancing Careers and Education (PACE)
  21. Fisher Ronald Aylmer (1935). The Design of Experiments Oliver & Boyd. [Google Scholar]
  22. Friedlander Daniel (1988). “Subgroup Impacts and Performance Indicators for Selected Welfare Employment Programs”. MDRC
  23. Gelber Alexander M., Isen Adam, and Kessler Judd (2016). “The Effects of Youth Employment: Evidence from New York City Lotteries”. Quarterly Journal of Economics 131, pp. 423–460. [Google Scholar]
  24. Gleicher Lily (2017). “Juvenile justice in Illinois, 2015”. Illinois Criminal Justice Information Authority
  25. Goncalves Felipe and Mello Steven (forthcoming). “A Few Bad Apples? Racial Bias in Policing”. Conditionally Accepted at American Economic Review
  26. Heckman James J., Smith Jeffrey, and Clements Nancy (1997). “Making the Most Out of Programme Evaluations and Social Experiments: Accounting for Heterogeneity in Programme Impacts”. The Review of Economic Studies 64(4), pp. 487–535. [Google Scholar]
  27. Heckman James J and Vytlacil Edward (2005). “Structural equations, treatment effects, and econometric policy evaluation 1”. Econometrica 73.3, pp. 669–738. [Google Scholar]
  28. Heinrich Carolyn J., Mueser Peter R., Troske Kenneth R., Jeon Kyung-Seong, and Kahvecioglu Daver C. (2013). “Do Public Employment and Training Programs Work?” IZA Journal of Labor Economics 2. [Google Scholar]
  29. Heller Sara B. (2014). “Summer jobs reduce violence among disadvantaged youth”. Science 346, pp. 1219–1223. [DOI] [PubMed] [Google Scholar]
  30. Heller Sara B. and Bhanot Syon (2021). “Overcoming Application and Take-Up Barriers for Summer Youth Employment Programs”.
  31. Hinton Elizabeth, Henderson LaShae, and Reed Cindy (2018). “An Unjust Burden: The Disparate Treatment of Black Americans in the Criminal Justice System”. Vera Institute of Justice
  32. Holzer Harry J., Raphael Steven, and Stoll Michael A. (2006). “Perceived Criminality, Criminal Background Checks, and the Racial Hiring Practices of Employers”. The Journal of Law and Economics 49(2), pp. 541–80. [Google Scholar]
  33. Jepsen Christopher and Rivkin Steven (2009). “Class Size Reduction and Student Achievement: The Potential Tradeoff between Teacher Quality and Class Size”. The Journal of Human Resources 44(1).1, pp. 223–250. [Google Scholar]
  34. Kessler Judd B., Tahamont Sarah, Gelber Alexander M., and Isen Adam (2021). “The Effects of Youth Employment on Crime: Evidence from New York City Lotteries”. NBER Working Paper No. 28373
  35. Kling Jeffrey R., Liebman Jeffrey B., and Katz Lawrence F. (2007). “Experimental Analysis of Neighborhood Effects”. Econometrica 75, pp. 83–119. [Google Scholar]
  36. Kowalski Amanda E (2018a). “Behavior within a clinical trial and implications for mammography guidelines”. National Bureau of Economic Research Working Paper 25049 [DOI] [PMC free article] [PubMed]
  37. — (2018b). “Reconciling seemingly contradictory results from the Oregon health insurance experiment and the Massachusetts health reform”. National Bureau of Economic Research Working Paper 24647 [DOI] [PMC free article] [PubMed]
  38. Leos-Urbel Jacob (2014). “What is a Summer Job Worth? The Impact of Summer Youth Employment on Academic Outcomes”. Journal of Policy Analysis and Management 33, pp. 891–911. [Google Scholar]
  39. MDRC (1980). “Summary and Findings of the National Supported Work Demonstration”.
  40. Modestino Alicia Sasser (2019). “How Do Summer Youth Employment Programs Improve Criminal Justice Outcomes, and for Whom?” Journal of Public Policy Analysis and Management
  41. Modestino Alicia Sasser and Paulsen Richard (2019). “School’s Out: How Summer Youth Employment Programs Impact Academic Outcomes”. Northeastern University Working Paper
  42. Mogstad Magne, Santos Andres, and Torgovitsky Alexander (2018). “Using instrumental variables for inference about policy relevant treatment parameters”. Econometrica 86.5, pp. 1589–1619. [Google Scholar]
  43. Mueller-Smith Michael (2015). “The Criminal and Labor Market Impacts of Incarceration”. University of Michigan Working Paper
  44. Ridgeway Greg and MacDonald John (2009). “Doubly Robust Internal Benchmarking and False Discovery Rates for Detecting Racial Bias in Police Stops”. Journal of the American Statistical Association 104(486), pp. 661–668. [Google Scholar]
  45. Roder Anne and Elliott Mark (2020). “Stepping Up: Interim Findings on JVS Boston’s English for Advancement Show Large Earnings Gains”. Economic Mobility Corporation
  46. Ross Martha and Kazis Richard (2016). Youth Summer Jobs Programs: Aligning Ends and Means The Brookings Institution. [Google Scholar]
  47. Schaberg Kelsey and Greenberg David H. (2020). “Long-Term Effects of a Sectoral Advancement Strategy: Costs, Benefits, and Impacts from the WorkAdvance Demonstration”. MDRC
  48. Schwartz Amy Ellen, Leos-Urbel Jacob, and Wiswall Matthew (2015). “Making Summer Matter: The Impact of Youth Employment on Academic Performance”. NBER Working Paper No. 21470
  49. Al-Ubaydli Omar, List John A., and Suskind Dana (2019). “The Science of Using Science: Towards an Understanding of the Threats to Scaling Experiments”. NBER Working Paper No. 25848
  50. Walters Christopher R (2018). “The demand for effective charter schools”. Journal of Political Economy 126.6, pp. 2179–2223. [Google Scholar]
  51. Weiss Michael J, Bloom Howard S, Verbitsky-Savitz Natalya, Gupta Himani, Vigil Alma E, and Cullinan Daniel N (2017). “How much do the effects of education and training programs vary across sites? Evidence from past multisite randomized trials”. Journal of Research on Educational Effectiveness 10.4, pp. 843–876. [Google Scholar]
  52. Westfall Peter H. and Young S. Stanley (1993). Resampling-Based Multiple Testing: Examples and Methods for p-Value Adjustment Wiley-Interscience. [Google Scholar]

RESOURCES