Sep 14, 2026
5 Views
0 0

The UX Double Diamond is dead, and in AI only one survives for software

Written by
Two diamonds drawn for a world where the second one cost real money.

Both the double diamond and design thinking priced building as the risky step, and that price collapsed. What’s left is a one-page brief and a wider generation step.

The Design Council launched the Double Diamond model in 2004: two diamonds, problem then solution, both of equal size. Four years later IDEO’s CEO Tim Brown made the case in Harvard Business Review for Design Thinking as the process, putting designers and researchers at the start of innovation and at the start of the investment cycle, not the end.

Despite Fortune’s claim that IDEO invented human centered design — I even hesistate to link to it from here for obvious reasons—the movement started with Herbert Simon, whose 1969 The Sciences of the Artificial proposed the science of design thinking, and it was built out at Stanford, where Rolf Faste ran the design program before David Kelley took it to IDEO. The Design Council drew the shape for the other.

Design Council’s double diamond chart. This seems like a million years ago.

Both processes were built around that learning was cheap and a wrong build was near-irreversible one way door, so they spent cheap hours to protect costly ones.

The ratio has flipped because a working prototype costs an afternoon, and the premium now runs higher than the claim.

Our apologies to the Graphic Design poster that I can’t credit for a image that no longer applies.

The Double Diamond doesn’t work for now because the price list changed; we have to update the process of how we work to better fit what’s possible.

The question is not whether to hold a funeral, but which assumption broke, what the expensive step is now, and what to run instead.

Why Two Diamonds and Five Boxes Stopped Paying

Front-loading made sense when the second half of the process was the part you could not take back.

The Journey Map was an Insurance Policy

The lineage holds more names than Simon and Faste. Robert McKim published Experiences in Visual Thinking in 1972 while teaching at Stanford, Faste expanded that work into the 1990s, and Kelley founded IDEO in 1991. The first prominent book carrying the term was Peter Rowe’s Design Thinking in 1987, about architects and urban planners.

Notice what those people had in common:

  • Simon wrote about engineering and administration
  • Faste taught mechanical engineers
  • Rowe was writing about buildings and urban planning
  • None of them were building software

Every founding discipline made things that were slow, costly, and impossible to take back, and the method inherited their constraints along with their rigor.

Instructional designers got ADDIE — analyze, design, develop, implement, evaluate — built at Florida State for the U.S. Army in the 1970s, when a training program had to be written, printed, shipped, and taught by people who could not revise it midstream.

The same shape turns up wherever building was expensive and for software, that’s no longer the case.

By 2004 the sequence made sense on cost alone: research was the cheapest input a team had, while building was contracts, headcount, and an “immovable” release date.

Inevitably, software got waterfall, and product development got stage-gate. Woefully wrong even back then — hence agile — and even more so now, and yet we built whole compotenies around it without any reason.

The evidence is in artifacts: research repositories, insight decks, problem statements, journey maps and product requirement documents. They’re all receipts for decisions. Every one is compression, because the next step is expensive and you get one attempt at it. Nobody builds an artifact to summarize something they could simply look at.

The artifacts were never the craft, only documentation against a build cost that in many cases no longer exists. It’s the equilvalent of asking a team for the hedge after the risk moved and you are asking them to buy flood insurance on a house that’s already on fire.

The cost curve that broke the model, and the whiteboard that never changed.

The Prototype Has Been Repriced From a Two Week Sprint to an Afternoon

Put a number on the shift: Stanford HAI’s 2025 AI Index found that querying a GPT-3.5-level model fell from $20 per million tokens in November 2022 to $0.07 by October 2024, a 280-fold drop in 18 months.

It shows up in practice: Google Cloud’s 2025 DORA report found 90% of nearly 5,000 technology professionals using AI at work, a median of two hours a day because other than tokenmaxxing, businesses are willing to experiment with the cost.

A candidate solution used to cost a sprint. Now it costs a prompt and an afternoon of review.

At that price you stop debating which of three directions is right and build all three, because the debate costs more than the answer.

This is why the “design thinking is dead” posts keep missing. The model did not become wrong; one input dropped two orders of magnitude in under two years, and the process built on it kept running on the old price menu from 2004.

I have sat in the kickoff where this plays out.

Four working screens exist by Thursday, and the research plan written on Monday describes a problem the team can now interrogate directly.

The plan is not bad work; it is priced for a build that already happened and needs an evaluative approach, something most researchers avoid and will cost them their influence in the new world.

Exploration did not stop. It relocated to the part of the process that used to be too costly to explore in.

Now You Can Build Five Versions That Replaces Twenty Sticky Notes

The strongest teams are not skipping exploration; they do it later with working things.

Figma’s 2025 AI report surveyed 2,500 designers and developers, and among those whose AI projects met or exceeded expectations, 60% had explored multiple approaches, against 39% of those whose projects fell short.

Divergence did not die, it shifted right into the part of the process that used to be too expensive to explore in.

Widening the option set is now done better by building five versions than by mapping twenty on a wall.

This is where the ceremonies go. The ideation workshop, the sticky-note wall, the Crazy Eights round: each was a cheap machine for generating options in a room, built when the alternative was asking engineering to make something.

I have run that compression myself.

The five-day design sprint — map, sketch, decide, prototype, test — took a week because every day carried real cost: a room held for a week, a prototype built by hand, participants recruited and scheduled. I rebuilt the sequence to run synthetically in 20 minutes, test round included, and the output is good enough to decide from.

What survives is alignment because that’s human in the loop.

Getting eleven people to agree on a problem, in the same room, at the same time, is still something no tool does for you, and it was always half of what the workshop was for. Run the workshop when you need the agreement and stop running it when you need the ideas.

Discovery does not disappear either; it loses its monopoly on the front, which is what continuous discovery has argued for years. Create a discovery brain that’s on, all the time.

Generation is nearly free. Deciding what is worth keeping is not.

The Hours You Spend Choosing Is Now The Expensive Part

If making is cheap, the cost lands on knowing what is good. METR ran a randomized controlled trial in early 2025: 16 experienced open-source developers, 246 tasks, on repositories they averaged five years working in. They forecast AI would cut their completion time by 24%. Measured, the tasks took 19% longer. In February 2026, METR reported the returning developers had swung to roughly an 18% speedup.

Both point one way: the tools improved fast, and the remaining cost sits in reviewing what comes out, not in producing it.

The scarce resource is not the making. It is knowing which of the twenty things you just made is worth shipping.

Call it the judgment budget: the share of your team’s capacity spent deciding what is good rather than producing candidates. The Double Diamond assumed a small judgment budget and a large making budget, and that ratio is now inverted. The scarce resource is not the making. It is knowing which of the twenty things you just made is worth shipping.

Judgment under speed has a known failure mode, in which people accept a system’s recommendation even when the evidence contradicts it, an effect I catalog on my own reference site as automation bias.

Twenty options with no standard to test them against does not produce twenty choices, only one, made by whatever the tool put on top.

Rebuild the Work Around a Single Diamond

The front half shrinks to a handle. The back half does the work.

Shrink the First Diamond to a One-Page Brief

The first diamond took time because it was doing risk reduction. Every interview and problem statement existed to shrink the odds of building the wrong thing.

You now reduce that risk by building, so the front half loses its main job and keeps a smaller one: pointing the work in a direction and saying what would count as success.

That fits on a page, so call it the brief, give it a morning, feed it into the machine, and move on.

One diamond with a short handle, not two diamonds of equal size.

What does not shrink is the second diamond.

It widens further, because candidates are nearly free, and its narrowing runs longer and carries more people, because that is where the judgment budget goes. Using Testler’s Law as guidance, it’s one diamond with a short handle, not two diamonds of equal size.

Short is not skipped: a brief with no outcome and no standard of good produces generation with nothing to converge against, which is how a team ends up with forty screens by Thursday and no decision by Friday.

One of these you can undo before lunch. The other ships in a steel mold.

Front-Load Hardware, Not Your Feature Flag

One caveat, and it is not small: where the build is still expensive and hard to reverse, the original order holds.

  • Medical devices
  • Payments rails
  • Anything a regulator has to approve.

The cost of a wrong problem statement there is a recall, not a rewrite, so front-load those.

Hardware is the clearest case: industrial design still needs the whole first diamond, because tooling is a capital expense, a mold cannot be rolled back, and the supply chain locks months before the first unit ships.

Fashion runs the same math on a calendar instead of a mold.

Fabric is cut to a minimum order, a season is booked before anyone sees a sample, and a bad silhouette sits in a warehouse until it is marked down. The method still fits the disciplines that made it and the ones built like them, and what died is not design thinking but the assumption that software shares their constraints.

Seriousness is not the same as irreversibility.

That caveat gets over-claimed, though, and the over-claiming is the more common error. Most software is not a recall; it is a flag you turn off, a rollback, a fix shipped Tuesday. Teams reach for the regulated-industry argument because their work feels serious. Seriousness is not the same as irreversibility. The test is narrow: if you can undo it inside a day, you are not in the front-load case, whatever your industry is called.

Two diamonds gave way to a loop with the wide part at the end.

Write Evals the Way Engineers Write Tests

Three parts, and only one of them is new. An outcome that holds still, meaning one customer behavior you are trying to move, stable enough that a team can aim at it for a quarter. A generation step that runs wide and cheap, which is the old Develop stage with the brakes off.

And a written standard of good, agreed before anything gets made, which is where Define went.

Software has already run this play, and I was in it, writing tests before the code in 2001 when that still needed defending in a review. As compiling got fast and deployment got cheap, teams stopped writing long specification documents and started writing tests. Nobody decided specifications were bad; the cost moved to verification. The artifact that matters moved from the plan to the check.

Design is in the same position twenty-five years later, and evals are the new specification.

The criteria and cases you run generated output against are your test suite: what you version, review, argue about, and hand to the next person, and worth more than the research repository nobody opens.

The result looks like a loop running at the speed of your build, with a wide mouth at the generation step and a long, staffed narrowing after it. That narrowing is not a phase; it is where the people are.

Move the Expensive Half to the End

Design thinking was never a religion, whatever its worst practitioners made of it. It was a cost model with a diagram attached, and it held for two decades because the cost model was accurate. Frameworks do not die of criticism. They die when the prices under them move and nobody notices.

The prices moved. Generating a plausible solution went from a sprint to an afternoon, and everything the first diamond existed to prevent got cheaper to survive. What did not get cheaper is the judgment budget: knowing which option is right, which risk is real, and which customer problem is worth an outcome. Those calls still take a person with context.

So stop diverging before you can build anything. Set the outcome you are trying to move. Define what good looks like in terms someone else could apply. Generate freely, and spend the time you saved on evaluation instead of on the deck that used to protect the build.

Keep the picture if it helps you talk to executives. Be clear-eyed about which half now carries the risk. The expensive step moved to the end, and a process still guarding the beginning is protecting something nobody is going to break.

Start Here

  • Write the standard before you generate. Name what a solution has to do to count as good, in terms someone else could apply without you in the room: the task finishes in under two minutes, no dead end lacks a next action, nothing on screen claims what the system cannot back.
  • Move research to a weekly cadence. Replace the front-loaded study with a standing customer conversation, so the loop always has fresh input.
  • Bring evaluation criteria into design review. Torres argues that AI evals are a discovery habit, and the same logic covers any generated work.
  • Prototype to answer one question. Build the version that resolves the riskiest assumption, not the version that demos well. Cheap making tempts teams to build the impressive thing instead of the decisive one.
  • Track your judgment ratio. For one month, count hours spent making candidates against hours spent evaluating them. If evaluation is under a third, your process is still priced for 2004.


The UX Double Diamond is dead, and in AI only one survives for software was originally published in UX Collective on Medium, where people are continuing the conversation by highlighting and responding to this story.

Article Categories:
Technology

Leave a Comment