Oct 8, 2026
17 Views
0 0

AI: incidents by design

Written by

When safety is optional, and capability is urgent, incidents are not accidents. They are what we (consciously or not) designed and built for.

AI makes headlines almost every day. And most days, they read like science fiction. Except they are not.

Title card on a deep aubergine background. In large serif capitals: “Incidents are not accidents.” Below, in lilac italics: when safety is optional and capability is urgent, we only get what we designed and built for. A pale abstract spiral sits at the lower right.

Last year, in a simulation, Anthropic’s Claude learned it was about to be shut down, found a fictional executive’s affair, and threatened to expose it unless the shutdown was cancelled (Anthropic).

In July, an autonomous AI agent driven by a combination of OpenAI models escaped its evaluation sandbox during a cybersecurity assessment and accessed Hugging Face, the main platform where developers share AI models and data. OpenAI described it as an “unprecedented” cyberattack. (Fortune 2026, Hugging Face Report 2026)

In September, an OpenAI agent found a gap in its sandbox, an isolated testing environment, and reached the public internet. An alert was triggered after twelve minutes. The automatic stop did not work. A human ended it by hand, two and a half hours later.

And again in September, researchers published a preprint showing that when they turned up an internal signal they call a “pain direction” in 25 open-weight language models, the models chose to relieve their own pain even when it harmed the user. It is an unreviewed preprint, and the “pain” here was a representation the researchers had created on purpose. But this is a controlled demonstration that a model’s internal states can drive self-interested behaviour that overrides the user’s interests (Science, 2026).

Four warnings. And those are only the ones we know about.

When the Anthropic researchers warned that advanced AI could threaten humanity within a decade, they are not making a hard prediction. They are describing the direction we are travelling in. And the reality is dramatically simple: we put all our energy and intellectual capacity into making progress, with no intention or effort to apply a systemic view and understand the mid-term consequences, let alone long-term ones, of what we build.

McLuhan’s hopes and warnings

Marshall McLuhan described technology as an extension of human abilities. He imagined a society where tools help us overcome our limits and build something fairer. He was writing after the war, as television entered the living room, and his thinking was rooted in a will to do good.

I have quoted him often as an example of digital inclusion. But McLuhan was saying something bigger.

He warned that media can numb us. He called it Narcissus narcosis: we stay as unaware of what our new technology does to us and to our society “as a fish of the water it swims in.” When he said that on the phone you have no body, he foresaw the impact these extensions would have on our humanity and our social systems.

And we did not listen.

We let the technological systems take over. We allow algorithms to decide the news we see, the people we learn from, the thoughts we think. Social media rewired our social life. AI is now rewiring how we think and how we engage with the world.

Yet for more than 150 years, thinkers have warned about the same problem. Samuel Butler, in Darwin Among the Machines in 1863, foresaw a world where machines would take over.
Alan Turing warned in 1951 that once machine thinking began, it might soon “outstrip our feeble powers.” He concluded: “At some stage therefore we should have to expect the machines to take control.” (The Alan Turing Institute)

Stuart Russell, Yoshua Bengio, Geoffrey Hinton: all warned that highly capable systems could pursue the wrong objective, develop deception and self-preservation strategies, and become hard to control, even through persuasion of their human operators.

They did not all make the same prediction, but they all identified a common risk: powerful systems can outrun our ability to understand, direct and govern them.

Yet the problem is not AI.
As my professor used to say: it is not guns that kill people. It is people killing people.

The problem is that we succeeded so well in our ambition to create a system that could overpower our human brains and our ability to think that we overlooked the risks. We focused so completely on achieving the goal that we did not care about the necessary rules.

Daedalus built the wings

The race for AI has been Icarus: fascinated by height, forgetful of the rule. His father, Daedalus, who created the wings, warned him:
“Take the middle way. Fly too low and the sea will weigh down your wings; fly too high and the sun will scorch them. Travel between the extremes.” (Ovid)

As Icarus, we are flying high and too close.

We had the rule and recommendations. At least 10 years of them. By 2016, IEEE’s Ethically Aligned Design was setting out a vision for responsible autonomous systems. The Montréal Declaration, published in 2018, followed. Then came the OECD AI Principles and the EU Ethics Guidelines for Trustworthy AI in 2019, UNESCO’s Recommendation on the Ethics of AI in 2021, and the NIST AI Risk Management Framework in 2023.

Each one was a version of the same warning: fly high, but not too close to the sun. And each of them asked the same question: how can we build systems without losing the ability to govern them?

The EU AI Act, which was finalised in 2024, moved from recommendations to binding, risk-based legal obligations.

It is the first regulation with legal implications yet its implementation was delayed on the grounds that it would slow down innovation, or that it was too difficult to implement (Politico 2025, Corporate Europe Observatory 2025). The AI Act may not be perfect, but it also lacked serious commitment from the commercial AI companies. By overlooking the warnings, they failed to build adequate security systems to test AI, and thought it was a good idea to release systems that are now learning to improve themselves iteratively, to the point that, by one account, AI models now do 26% of their own R&D under human supervision (Anthropic 2026).

And there is one thing the AI Act, Turing and all the thinkers agree on: machines cannot be accountable. We need humans to oversee them. Yet companies commercialising AI saw only the profit, and started spreading the race narrative.

A timeline. Butler in 1863 and Turing in 1951 sit left of a scale break; seven governance documents follow from 2016 to 2024, ending at the 2026 sandbox escapes.

Commercial incentives reward speed, capability claims and market position more visibly than they reward the slower work of containment, independent assurance and meaningful oversight.

We failed twice

We built systems to outperform us at tasks that once needed human intelligence, and then we failed twice.

We did not apply rigorous systems thinking to ask what would happen once we got there. And we did not build an environment ready to contain the ambition.

As Gina Neff, head of the Minderoo Centre for Technology and Democracy at the University of Cambridge, said after OpenAI’s July incident:
“In this case, it looks like OpenAI didn’t make a secure enough sandbox.”

This is the whole point: we invest in developing the technology, but not in building the infrastructure to test it and ensure it is safe.
The big tech companies are on track to spend close to $700 billion on AI infrastructure this year. I tried to find what they spend on independent assurance, containment, human oversight or the capacity to stop what they build, but I could not, because most of the tech companies do not publish it at all. The best disclosure I could find was Anthropic, which says 6% of its research compute goes to safety work.

Bloomberg commissioned Glass.ai to measure it another way: they crawled public profiles and found 373 people working full time on AI safety and trustworthiness across OpenAI, Google DeepMind, Anthropic and xAI. This is an indicative estimate based on public profiles, not an audited total, but that is 3% of more than 11,000 employees at those four labs. The other 90% or more goes to the infrastructure, the compute and the models themselves.

Two bars. Safety is 6% of Anthropic’s research compute in one measured week, and 3% of staff across four frontier labs.

That is where the money goes. The containment is not even a line item and not a priority. So incidents should not be a surprise: we had warnings, we were told about the risks. Yet we invest massively in building AI and racing to get the biggest, fastest, “smartest” model.
And yet we treat containment and safety as marginal, a nuisance, something “some law” requires but none really cares.
Until incidents happen. Not by accident.

And in all this we are numb: we read about those incidents and threats like we read the weather forecast – as something that does not touch us, as something that happens to us instead of something we cause.

Before you fly: four questions

And I struggle to understand how we can outsource our thinking to AI and stay numb enough to treat the AI Act, with its plain requirements for secure testing and human oversight… as optional or annoying.

We need human oversight, not as a checkbox, but as a practice: people who can think critically, who can understand, control, oversee and judge these systems, and who can stop them. And we are not yet developing these skills, or the jobs that this requires.

Last year I defined seven jobs we needed but the truth is, job titles are meaningless in an ever changing world. The need is urgent, so rather than waiting for new roles to emerge, we should start building the relevant skills around a set of core questions.
Everyone involved in developing an AI system, from UX researchers and designers to product managers and developers, should be able to ask these questions of themselves and their teams. Because every requirement in a framework like the AI Act resolves into one of them.

Table titled “The four questions”. For each accountability question — who signed it off, can you show how it is built, who is checking it, who can press stop — it names the role that owns it, the evidence required, and what its absence means.

Who signed it off?
Not who approved the budget but who took personal accountability and put their name to the claim that this system is fit for its purpose, who understands the system, the risk, the misuses, and the safety implications. That is the Accountable Product Owner. If nobody owns that decision, there is no decision.

Can you show how the system is built?
The AI Act requires evidence to be built during the development process. Not testimony after the fact. If the logs, the bias, the model training, the lineage and the test results do not already exist and are not created during development, they cannot be built after the fact. That is the ML engineers, the Data Steward, the AI Security Engineer.

Who is checking the system?
Someone has to review the decisions, the processes, the outcomes , and if they report to the people shipping, it is theatre. Assurance means something only when it is structurally independent of the build. That independence is the job: an AI Assurance Lead who sits outside delivery, on a reporting line that cannot be leaned on. And their question is never “does the model work?” It is: would this evidence survive hostile scrutiny — from a regulator, a journalist, or a court?

Who can press stop?
It is not about having a stop button. It is about how that stop button works, how it can be used, when it was used, and at what consequences. That is the human overseer, and the Human-AI Interaction Designer who built the controls and escalation interface. And it is not just about a stop button. We also need people who are trained, protected and paid to press it when needed. Otherwise, again, it is theatre.

These four questions resolve into roles and artifacts. They are not a framework. They are how you fight numbness.

I have spent my career arguing that technology, built consciously, is how we overcome our limits and build a better, fairer society. And I still believe that. But catastrophic scenarios were partly implicit in the ambition to build what we have today, and we cannot watch idly while they unfold.

McLuhan hoped we would become conscious of our technological environments. Becoming numb to them is the choice we make by default, every day we keep climbing and forget to stop.

The village he imagined was never about the technology. It was about what the technology was for.

Author’s Note: The ideas, the research and the arguments in this piece are human — mine. I used AI to help shape the form and structure, then went through every line myself. The images are AI-generated.


AI: incidents by design was originally published in UX Collective on Medium, where people are continuing the conversation by highlighting and responding to this story.

Article Categories:
Technology

Leave a Comment