April 28, 2026
Scaling, part the second
Friends, it’s a beautiful spring day here in Connecticut. All the leaves on the trees came out last week like somebody flipped a switch, and the flowering trees are particularly lovely. But don’t worry, the weather’s supposed to get worse again tomorrow. Thank you for subscribing to the Coaching Letter—you’re terrific.
In the last CL, I talked about scaling and its complications. I talked about Logic Alpha, which assumes that scaling is always linear, hierarchical, and managerial. And sometimes scaling should be that—no question. But we argue that scaling should sometimes be much more organic, but not a free-for-all: tight on practice, loose on everything else. Call it Logic Phi, standing for the Greek phroenesis, which means wise action. Andrew and I have been working on a way to represent the different ways to represent scaling—not quite ready to share that yet, stay tuned.
But—at least, I think this is true—however you think about scaling, you should also be aware of the ways in which scaling fails. And the best source for this is a book called The Voltage Effect by John List. You can listen to him talking about the book on several podcasts, including this one, or read List himself here.
List is not interested in redefining scale or in describing how it unfolds in complex systems. He is interested in why things that appear to work fail when expanded. His central concept—“voltage drop”—captures the loss of effectiveness as an intervention moves from small, controlled conditions to larger, more variable contexts. To make the connection to the last Coaching Letter more explicit, List is diagnosing Logic Alpha when it is not executed well or appropriately. What List is really showing is not just that scaling is hard, but that the most common way we think about scaling—take something that worked and replicate it—is systematically vulnerable to failure. His cautionary tales are not tied to any one model of scaling; they are structural risks that show up whenever we try to move from small-scale success to broader adoption
Whatever your approach to scale—whether you lean more toward Logic Alpha or Logic Phi—List’s failure modes are worth keeping in mind. . Here are the five failure modes:
- False positives – The initial success was not real or generalizable; it was an artifact of small samples, bias, or unusually favorable conditions, or there is a confounding variable that is the true source of success.
- Population mismatch – The intervention worked for one group of people but does not transfer to others.
- Context dependence – The success depended on specific contextual conditions (leadership, culture, resources) that do not hold elsewhere.
- Spillovers and system effects – Scaling changes the system itself in ways that undermine the original effect.
- Cost and diseconomies of scale – The resources required to maintain quality become prohibitive as expansion occurs.
So I’m going to take up most of the rest of this CL giving examples from education, because I think they’re useful. I am confident that many of these examples will resonate with you…
False positives.
A district hears about a neighboring district’s focus on PD on higher order thinking, and assumes that their promising AP scores are a result of the PD. But actually, there have also been big changes in the high school department meetings protocols, so that high school teachers are spending a lot more time working on enacting practices that keep students thinking than they spend in whole-school PD. The district adopts the PD, but fails to understand the importance of the focus on enactment. AP scores do not go up.
A middle school math teacher uses vertical whiteboards and random groups in one class and sees dramatic gains in participation and thinking. Her principal sees this happen, and starts telling other teachers that they should use whiteboards. But those other teachers don’t realize the extent to which the first teacher is engineering the tasks to be low-floor, high-ceiling thinking tasks, and how she ramped up the challenge of the tasks over time, so when the other teachers start implementing whiteboards without the same attention to task, the kids are not more engaged because they are doing the same work as before but this time standing up, and they complain.
Population mismatch
A coaching model that works well with early-career teachers is extended to veterans. New teachers appreciate detailed feedback and concrete suggestions; experienced teachers experience the same approach as intrusive and controlling and complain about the coaching. Conversely, a coaching approach built around reflective questioning works beautifully with experienced teachers but leaves novice teachers frustrated because they need more direct guidance.
One school uses a literacy intervention model that is very successful and leads to big gains in standardized testing. Another school implements the same program without realizing that the first school has screened carefully for need (decoding, comprehension, vocabulary) and has only used the intervention with kids who fit the profile. The second school implements with all kids below a certain percentile score in reading, and is disappointed when they don’t see the same results.
Context Dependence
A school successfully uses weekly teacher team meetings because the principal protects the time, attends regularly, and empowers the instructional coach to keep the focus on instruction. When the assistant superintendent for instruction mandates weekly meetings in every school, many become perfunctory because other principals do not understand the importance of creating the same conditions, and approach them as exercises in compliance—e.g. they check attendance and monitor the use of protocols, instead of paying attention to the learning cycles that are the real engine of improvement. Teachers feel like they are being treated like kids in detention and silently rebel.
A district launches a new literacy program that has a really solid research base and was very successful in a single-school pilot. But in scaling up from the pilot to the full implementation, the district does not provide the same training in teaching literacy or the intricacies of the program, and without those, implementation more broadly suffers and the program looks like a failure.
Spillovers and System Effects
A district highlights one school as a model site. Over time, the principal and teachers at that school become overwhelmed with visitors and presentations, leaving them less time to sustain the actual work.
A district adopts a strong emphasis on foundational skills in reading. As the work spreads, teachers narrow the curriculum to isolated drills because they believe that is what the district values. They associate the new program with more direct instruction and assume that this represents a new direction for the district rather than simply an approach to teaching early reading.
Cost and Diseconomies of Scale
A nonprofit works with one state to support districts in improving their lowest achieving schools. Another state asks for the same thing, but balks at the cost. So the districts in that state get the same program in name only; in practice, the amount of time between support meetings with the nonprofit mean that the work never really takes hold.
A district depends on one extraordinary central office leader who personally facilitates every meeting and supports every school. The model works only as long as that one person can sustain it. As List says, you can’t scale people.
Honestly, I could keep going. Many of the examples I came up with cut across multiple categories—I had to get ChatGPT to help me simplify them back down to just one category.
So I just want to add a word about pilots in general. Pilots are intended as proof of concept, but they are often treated as proof of scalability—I don’t remember if List says that explicitly, but it’s definitely true. In other words, a pilot answers the question “Can this work?” but is often misread as answering “Will this work everywhere?” And to know the answer to that, you have to keep trying things out. That’s one of the reasons why I push very hard against the language of launch, roll-out, turnkey, and other similar ideas—because I think they are indicators of an underlying mental model that you can take a program/initiative/whatever and lay it over another context like putting a tablecloth on a picnic table. You have to learn your way to scaling in education, so that’s a future Coaching Letter.
I’m facilitating a book study on Friday for the book What’s Your Problem, by Thomas Wedell-Wedellsborg, which is about how we frame problems—more on that in a future Coaching Letter. But how we frame solutions also matters—replicating success is not easy, straightforward or, in many cases, possible. There’s a great line in the book Developmental Evaluation by Michael Q. Patton about how sometimes you just have to have the whole program in order to have the best chance of success—I feel that way about HQI Live! I’m sure that many people to whom I describe that 5-day event think, “Really? You need 5 days for that?” and think they could do it in one or two. But there’s a reason that it’s five days, and a mistake to think you could accomplish the same goals in a shortened period of time, yet we do that kind of thing all the time in attempting scaling in education.
OK, what started off as a nice clean set of examples to elucidate List’s failure modes is turning into a soapbox, so I’m going to stop. All that’s left is to tell you that registration is still open for the Building Thinking Classrooms conference the last two days in June—you should absolutely come to that if you can. The conference is filling up, but there is plenty of room in the pre-conferences, which cover BTC 101, leading for BTC, and coaching for BTC. Hope to see you there! And Maximizing Opportunity to Learn is out! And can I just reiterate that, while I hope everyone reads it, it’s not designed as a book for teachers to read and become better teachers. It’s a book about the kind of instructional program that a district should build, which includes features that teachers do not themselves control, such as the availability of high-quality tasks. You can order your copy here, and we also created a website with supporting materials, and if you’re looking for something there that you can’t find, please let me know! Thanks again for reading. Best, Isobel
Isobel Stevenson, PhD PCC
Author of The Coaching Letter



