Welcome to the new subscribers. Inevitable & Obvious is a publication about Earth systems stabilization, the cooling interventions and other tools we’re going to need that buy time while decarbonization and carbon removal catch up. I spend most of my days talking with the people actually working on this, and what you’ll read here is my view on where the field is and where it’s headed.
If you’re new, How to Learn Everything You Need to Know About Climate Cooling is the place to start.
When it comes to climate interventions, we spend a lot of time asking whether we can do this as well as whether we should. Asking “can we?” leads to lab research and modeling work. Asking “should we?” leads to conversations about social legitimacy and governance. Both are important, for concrete reasons. Social legitimacy is a critical component of any large-scale intervention project, because to deploy these interventions durably you need the support of policymakers, and policymakers need the support of the people affected. Governance matters because these will be multi-decade megaprojects with hard decisions and tradeoffs to work through, now and for decades after deployment starts, and that requires institutions capable of working through them well. Any large intervention project will depend on those things as much as it depends on the finance and the engineering and the ongoing operations and maintenance.
What’s interesting to me is how little time we spend, by comparison, on a third question, which is how we would actually do any of this at scale. By this I mean, it’s one thing to ask will marine cloud brightening work to shade coral reefs or even provide larger-scale cooling benefits. It’s an entirely different question to ask, ok then how many MCB sprayers does that require, on the back of how many ships, where should they be located, how much will this cost, who will support this, and more.
I think this current state of unknown scalability is a natural result of how the field has evolved. For a long time, can-we was a science question, and scientists worked on it. Over the last decade plus, as interventions started to become more real, the should-we people arrived, from policy shops and think tanks and the social sciences. The people who work on operational scaling mostly haven’t arrived yet, and why would they? There’s been no money in it, and we’re only just now seeing the first funded companies. It does mean that when these conversations happen, hardly anyone in the room is thinking about what it would take to run one of these projects at scale.
If this question is interesting to you, there will be an entire session to discuss how to operationally scale interventions at the Global Risk and Stabilization Summit in Reykjavik, October 12-13.
Apply for registration at https://globalstabilization.org/.
My background is in engineering, and I spent years building a company in the neighboring field of carbon removal. A lot of what I thought about then, and still think about now, is how you scale something like this in the way most likely to succeed, and what you would have to start doing today if you worked backwards from that.
So this piece is about the how. I’m going to make the case that stabilizing climate interventions comprise a reference class. However different these projects look from each other, they’re going to run into similar challenges around social legitimacy and establishing governance. Who decides, and where does their authority to decide come from? How do the projects get funded, and who pays? What technology goes into them, who manages and runs them, and how does that work get contracted out? How does the engineering and operational scaling actually happen, and how are these things maintained over years and decades? How do they evolve as our understanding of the Earth systems they touch evolves, and how do we monitor their impacts and make decisions about what we learn? Every one of those is incredibly complex, but the interventions will be thought about similarly and in relationship to each other. And the best place to start on the how is the track record of extraordinarily large projects, which is the subject of How Big Things Get Done by Bent Flyvbjerg and Dan Gardner.
Flyvbjerg is an Oxford megaproject researcher with a database of more than 16,000 large projects across 136 countries, from home renovations to nuclear power plants to the Olympics. And what his database shows is that only 8.5 percent of the projects in it came in on time and on budget. Only 0.5 percent(!), one project in two hundred, came in on time, on budget, and delivered the benefits they promised. That is a horrifically poor record of performance. The pattern is consistent enough that Flyvbjerg calls it “the iron law of megaprojects.”
So why does this keep happening? A big part of Flyvbjerg’s answer is that projects get committed to before they get planned well, and the people who want them have incentives to commit early. Politicians want to be first, or to build the biggest, because that is what gets remembered, and they know how hard funding is to win. He summarizes this as a tactic where politicians want to just start “digging the hole,” because once the hole exists, people susceptible to the sunk cost fallacy will keep committing money rather than admit the project was a mistake.
Where I live in Seattle, we lived through this in the 2010s. When we replaced the Alaskan Way Viaduct, the design called for a double-deck tunnel with two lanes in each direction, running underneath downtown to connect the south end of the city with the north end. To bore it, the state commissioned Bertha, the largest tunnel boring machine ever built at the time. Five months into the dig, Bertha hit an abandoned steel pipe and broke, and it sat under the city for about two years while the contractor dug a rescue shaft, lifted the four-million-pound front end to the surface, and rebuilt it. Once that machine was in the ground there was never any question of stopping. The tunnel is completed and done now and it’s actually fantastic to have, but it was much more expensive than planned and took years longer than planned, and the stall alone accounted for a couple hundred million dollars in overruns.

Bertha’s problem wasn’t tunneling, because that’s a mature discipline with a century of accumulated practice behind it. Bertha’s problem was that no machine like it had ever been built, so all the myriad ways in which it could break or run into problems hadn’t been put to real-world testing. The same dynamic shows up at small scale too. I learned it personally by buying a Fisker Ocean, a first-of-a-kind car from a brand-new EV company. Fisker went bankrupt shortly after delivering the first ~5,000 vehicles (of which mine was one), and so I’ve spent a lot of money on repairs that should have been covered under warranty had the company survived. One of the problems with the company was an inexperienced COO (who was married to the CEO) who took charge for selecting parts and components, invariably choosing the cheapest option. Look in the photo below—do you see the white inserts in the door handles? They’re purely decorative, but it turns out they were made of a plastic material that happens to crumble into dust when exposed to UV light. I don’t know about you, but I drive my car outside, and so after about 18 months, I had to replace the door handles entirely. I’ve totally learned my lesson. The next EV I buy will be from a major manufacturer who knows how to build cars and, more importantly, run a business that builds cars. This is a common pattern in climate tech startups too. When I’ve talked with climate VCs over the last year or so, they’ll typically say: we don’t want to fund first-of-a-kind tech. First-of-a-kind applications are great, but use off-the-shelf components. Components that have been through years of use have had their failure modes found and fixed. New components haven’t.

Big projects also go better when they’re delivered by teams whose experience compounds, because so much of what an experienced team knows is tacit, the kind of knowledge you get from doing and can’t fully write down. Sort of how if you try to explain to a child learning to ride a bicycle how to balance, you can’t really. They just have to figure it out for themselves based on feel. That feeling is tacit knowledge.
Flyvbjerg showcases the Olympics as the archetypal example where this problem arises. The games are chronically over budget and behind schedule, partly because they move from city to city, so each host starts over completely. Maybe you’d think cities would hire vendors with experience from previous Olympics, and that does happen to a degree, but Flyvbjerg found that in most cases politicians promise contracts to local vendor companies (read: inexperienced vendors) in order to win public support for funding the games. This is another reason to think in reference classes. A field that treats its projects as one class can carry teams and their accumulated experience from project to project, while a field that treats every project as unique rebuilds that experience from zero each time, at each project’s expense.
Flyvbjerg calls this underlying error the uniqueness bias. Everyone believes their project is special, and in one sense every project is, the same way every human being is an individual. But humans share a set of common characteristics, and projects do too. The Sydney Opera House is unique in its location and design, and it is also an opera house that needs good acoustics, a certain number of seats, entryways, and different classes of tickets to sell. Daniel Kahneman calls the two ways of seeing a project the inside view, where my project is special and I forecast how it will go based on its particulars, and the outside view, where my project is one of a class and I forecast it based on how that class has actually performed. Almost everyone takes the inside view, and Kahneman and Flyvbjerg both consider that a fatal forecasting error.
Imagine renovating your kitchen. You measure everything, price out the materials, and add it up to get a total budget. And then when the work begins, the countertop breaks on delivery and you find mold under the floorboards. Dealing with those setbacks weren’t in your itemized plan, but they are in the average cost of kitchen renovations, because other people hit surprises too. So that’s actually how the book recommends forecasting what a project will take to do. Start from that average for the reference class (i.e. other kitchen renovations in your city) and adjust for what’s genuinely different about your project. That might seem too simple to be correct, but Flyvbjerg has empirical evidence that shows it reliably produces better forecasts. He turned this into a formal method called reference class forecasting.
Now apply this to climate interventions. When I’ve suggested that these projects have a lot to learn from each other, the pushback from inside the field has usually been some version of: “they’re too unique between them,” “SAI has nothing to teach an Antarctic seabed curtain,” and “a curtain has nothing to teach a reef-shading program.” I think that’s the inside view, and it’s precisely the error that the forecasting research warns against. Whatever the intervention, you have to get people and equipment to remote, difficult places, power what they’re doing, and keep them supplied. You have to measure whether the thing worked. You have to earn and keep the support of the people affected. Those are common characteristics of the class. And when you view this from the perspective of a government contending with catastrophic climate risks, these projects are plainly related. Tornadoes are different from hurricanes are different from earthquakes are different from extreme wildfires, but FEMA is the same agency in the US that responds to all of them. If you view these Earth systems from a risk-first perspective, then stabilizing interventions are related and intertwined.
Let’s go back to that 0.5 percent success rate for very large projects. If only one project in two hundred delivers on time, on budget, and with the promised benefits, then we should assume that any stabilizing intervention will run over time, over budget, and/or fail to deliver what it promised. This is what history shows us to be true, and it is a terrible track record to start from!
For this field, that default is untenable. For one, with temperatures accelerating higher, coral reefs reaching a tipping point, and AMOC collapse risk looming, we don’t have the time to waste. But also the money and the mandate will eventually have to come from governments and international agencies, which requires social legitimacy. Public support is what lets a government fund a decades-long project, and it erodes fast when a project is late, over budget, and visibly failing to deliver what it promised. These projects will also outlast the governments that start them. Support has to survive elections, and given the risk of derailment (the risk that society’s capacity to even respond to climate change will degrade because of the damages caused by climate change), the political environment these projects live in will get harder over time.
A stuck megaproject ends in one of two ways: either you keep going and absorb the cost and time or you pull the plug—which might also cost you. Once the Bertha machine was under the city, there was no way they weren’t going to finish the job, whatever it cost. The book’s other example is Japan’s Monju fast-breeder reactor, which took about thirty years and more than $9 billion dollars to build, operated for roughly 250 days in total, and never delivered meaningful power before the government decided in 2016 to decommission it. An intervention project that languishes is not neutral. It consumes money, attention, and public patience that the rest of the field needs.
Policymakers and the public are going to view these projects as a class whether or not the field does, the same way FEMA’s disasters are a class. If one intervention fails badly in public, it damages the ability to do all the others. (Nuclear power already demonstrated this. After its most public failures, the entire technology class froze for a generation.) So I think we have to hold these projects to incredibly high standards of operational excellence—on time, on budget, delivering what was promised—because the whole portfolio’s permission to proceed depends on it.
The reference classes for costing this work already exist. The International Thwaites Glacier Collaboration has been running science expeditions to one of the least accessible places on the planet on a five-year, fifty-million-dollar budget, so we know roughly what it costs to get scientists and equipment there and back. But an intervention would need more than visits. If we’re going to attempt a seabed curtain to slow the collapse of Thwaites, then phase one is a permanent remote sensing hub at the glacier, collecting data continuously instead of expedition by expedition. Marianne Hagen and I talked about this on the most recent episode of my podcast. That hub is a bigger ask than what the science teams do today, and reference class forecasting would handle it the normal way, by starting from what expeditions cost and adjusting upward for continuous presence. The subsea cable industry lays and maintains fiber across ocean basins as a routine commercial service, and that industry is the nearest class for putting a curtain on the seabed. Militaries have spent a century learning how to supply remote bases. When a piece of an intervention really is first of a kind, the move is still to find the nearest class, take its average, and adjust, instead of pricing the project from its particulars and finding the surprises later.
The book’s answer to the iron law problem is modularity. Flyvbjerg uses Lego as the model and asks you to figure out what the small repeatable piece of your project is, your Lego block, because the projects in his database that reliably come in on time and on budget are the ones built by repeating a small unit many times. He compares solar and wind farms, which are assembled from repeated panels and turbines and reliably deliver, against nuclear plants, which are enormous custom builds and sit near the bottom of his rankings. The Empire State Building went up fast for the same reason, one repeated floor at a time, with the crews getting better at each one. Repetition is how experience accumulates. For stabilization, I think the Lego blocks mostly won’t belong to any single intervention. When you decompose these systems into components, the same technologies keep showing up: underwater drones and autonomous platforms, cold-rated sensors and communications, remote power generation, the logistics of moving equipment to hard places, etc. They’re shared across sea ice, ice sheets, coral reefs, the stratosphere, and more. If you build those as repeatable products instead of expedition one-offs, then every intervention that uses them inherits the learning from all the others. That’s one more reason the reference class matters. The Lego blocks only become visible when you look at the projects together.
So here’s the paradigm I think this field should adopt while it’s still early enough to adopt one. Treat stabilizing interventions as a reference class with shared characteristics (African Futures Tech Lab already put out their own climate interventions taxonomy, which sort of gets at this), and take the outside view when forecasting them. Cost interventions from the classes that already exist (polar expeditions, cable ships, remote bases, stratospheric planes), then adjust for what’s genuinely new. Prefer proven components underneath novel applications, and treat any component that has never existed before as a schedule and budget risk. Structure the work as repeatable programs built from shared modules, so experience compounds across interventions instead of being relearned by every project the way the Olympics relearn hosting every four years. None of this slows anything down. The iron law is what punishes moving fast without planning, and given the urgency and the stakes, we cannot afford to languish our way through this. If we’re going to deliver on these projects, and we have to, this is what delivering takes.



