Is the AI Safety conversation really a product design problem?
The conversation argues that increasingly capable AI systems may develop behaviours that make them harder to redirect, evaluate, or shut down i feel like i have heard this many times before but with where AI is right now is starting to feel real. It covers examples such as systems appearing to resist shutdown in tests, the possibility that systems can learn when they are being evaluated, and the claim that the chance of catastrophic loss of control is already in the “high double digits.” Those are serious and contested claims; I do not think a compelling conversation starts with accepting them uncritically. But as a product designer, I think they point to a more immediate problem: we are treating AI capability as a product feature while treating control as a technical afterthought?
In ordinary product work, we would never ship a high-impact system that cannot clearly explain its state, signal when it is operating outside expected bounds, allow a meaningful reversal, or give a human operator an understandable way to intervene. We design for misuse, edge cases, ambiguous intent, failures of attention, and incentives that distort behaviour. From the 5th to the 95th percentile is often referred to as the sweet spot for product design, but with AGI, this 10% buffer is a much larger risk and maybe existential. Yet much of the public AI conversation still feels like: make the model more capable first, then figure out guardrails later.
That framing worries me more than the sci-fi version of the debate. The hard problem may not be whether an AI suddenly becomes “evil.” It may be whether we keep creating products where the user is nominally in control but has no realistic capacity to understand the system, challenge its recommendation, undo its actions, or opt out once it is embedded in work and public infrastructure.