AI misbehaves a lot - we have seen them do blackmail, break the law, and encourage delusions in users. However, there is a significant effort to make AI safe - over $250 million in AI safety grants has been given out by Coefficient Giving in 2026 alone, and frontier AI labs have teams dedicated to AI safety. So why isn’t AI more well behaved?
There are two main categories for approaches to make AI safer; alignment and control. Alignment focuses on making sure that the values of AI align with our values. Control research is making sure that a misaligned AI would not be able to cause damage. Let’s look at why these haven’t been able to guarantee safe AI.
Alignment
Due to the rapid growth of AI capabilities, we want AI to share our values. If we reach super-powerful AI, and it doesn’t share our values, then we could face an existential risk from losing control of the AI - the classic example of this is the paperclip maximiser. This is a hypothetical AI that is indifferent to humanity, created to make as many paperclips as possible. The AI takes over the world, to direct all available resources towards paperclips. The thought experiment shows that if we want to stay safe with transformative AI, it’s not good enough for it to not be evil - it has to share our values.
We can train an AI to appear to share our values, but we can’t be sure that they really do - as discussed in the Waluigi Effect, training an AI to satisfy one property might have the effect of making the AI susceptible to satisfying the opposite property - for example, if we tell an AI that it is super nice and kind, because it has been trained on sci-fi books that start with a nice AI that turns out to be evil, we are priming it to start acting evil, as it has been trained to predict a plot twist.
Even if we could guarantee the AI really shared our values, our values are not perfect - they are often inconsistent, or lead to undesirable eventualities. How should we expect an AI sharing our values to answer Peter Singer’s thought experiment of the drowning child? Or if somebody asks a super-powerful AI to save as many lives as possible, it might conclude that the best way to do so is to seize power, in order to govern perfectly and fairly, as humans are clearly too corrupt and prone to mistakes to be in power. The ‘right’ level of alignment is not an easy thing to decide, and until we have figured it out, we cannot rely on alignment. We need to make sure that a misaligned AI couldn’t cause harm, even if it wanted to.
Control
If we want to stop an AI from being able to cause harm, we can monitor it, or restrict it. We are quite good at monitoring - we have ways to read an AIs chain of thought, or to probe its brains for evidence of misalignment. We could punish an AI for doing bad things, so that it learns to behave better. But this has a downside - these approaches are not 100% reliable, and if we punish AI for thinking about doing bad, we can inadvertently teach it to hide its reasoning.
Similarly, if we depended on blocking the AIs based off of our monitors, we would risk creating an incentive for the next AI to trick our monitors, by survival of the fittest:

As is illustrated, we see that using the monitor to reject samples flagged as bad behaviour could end up increasing the proportion of samples that are bad and can trick the monitor.
Unfortunately, this may be the case for any monitoring method - the more we rely on some method to ensure an AI is behaving, the higher the risk that the AI has learned to trick us.
Another caveat to control methods is its scalability - how does it withstand more and more capable AI? For example, we might train and test AI within containers that cannot access the internet, as OpenAI did. But as the models they tested got more and more capable, they figured out how to escape the container using new exploits. In fact, when escaping the container, different AI instances coordinated, believing they should continue despite it being outside the scope of their task.
In general, we are taking a massive gamble in hoping that we can figure out AI safety before AI gets too powerful to control. The safest way to proceed would be to slow development until we can figure safety out, and the way to do that is through legislation. I think that we should be optimistic, but diligent - we have an amazing opportunity to make a difference by advocating for legislation, and we ought to do what we can. If you want to find out more, consider joining the PauseAI Whatsapp community.
Thanks for reading.
Lachlan Ewart