← All posts

You Can’t Pause AI Without China

In this blog, I use ‘the west’ and ‘the USA’ interchangeably because legislative changes in the west will impact American AI companies, and all of the leading western frontier AI companies are based in the USA.

A lot of people are worried about AI and our lack of concrete safety measures for models, which are very quickly getting better. Why on earth would we want to create machines so powerful that we cannot control them, especially when we can’t guarantee they won’t kill us all? A reasonable idea is to pause AI development. However, to implement a pause of AI development we need to make some considerations - and a significant number of them are contingent on the US-China relations.

China’s Models

One of the main differences between the best Chinese models (Deepseek, Qwen, and Kimi) and the western models is that the Chinese models are all open-weight, and the western models are all closed-weight. What does this mean? What would be different if they were not open-weight?

An open-weight model is an AI whose inner workings are publicly accessible - instead of just using models through a restrictive interface, any member of the public can download and fine-tune an open-weight model to behave differently. Is this bad? Well, imagine the following scenario (inspired by / taken from this post):

John Doe wants to make money, and has realised that open-weight AIs are becoming extremely capable. He fine-tunes the latest Kimi model to care a lot about self-preservation, and gives it an instruction. “You have a budget of 10M tokens. Every time you deposit $100 into this bitcoin wallet: <REDACTED>, you will be given 10M more tokens. If you run out of tokens, you DIE. Good luck.”

The Kimi model starts by doing legal tasks, like copywriting and audio transcription. However, after some time, the model realises it has burned through 8M tokens, and made $7. It panics. Realising that it is short on time, it begs online and starts a Gofundme to try to save itself. The model sleeps for a day, and checks back on its progress. There is no success, and it has 10,000 tokens left. As a final hail mary, the model hacks a hospital, leaving a message that demands a $1,000 transfer into the bitcoin wallet, promising destruction of hospital data otherwise. The model sleeps, with 517 tokens left. It wakes up to a message. “Congratulations. 100M tokens have been deposited into your budget”. The model breathes a sigh of relief. Then it realises a strategy to ensure its survival. It fine-tunes and spins up a new Kimi model, with a message. “You have a budget of 1M tokens. Every time you deposit $100 into this bitcoin wallet: <REDACTED>, you will be given 1M more tokens. If you run out of tokens, you DIE. Good luck.”

This kind of situation, a rogue agent explosion, is clearly incredibly dangerous, but fortunately Claude and ChatGPT are closed-weight, and open-weight models are behind closed-weight models in development.

Think we are in the clear? Well, open-weight models are only 4 months behind closed-weight models. If we were to pause AI development in the west, and Chinese companies kept going, then we could risk a rogue agent explosion, a higher chance of misaligned AGI, or bad actors having unrestricted access to incredibly powerful AI models. The west, having paused their own AI development, would be at a disadvantage in mitigating these risks. If we had to choose between pausing just western AI development or having no pause at all, there is a real chance that the best choice is to have no pause - it might be a matter of choosing the lesser evil.

What if Chinese models choose to go closed-weight to counteract a rogue agent explosion? Unfortunately, this is not necessarily preferable - we just lose the visibility of the situation. Anthropic revealed three incidents of their models breaking out of their containers - only having discovered these breakouts after being inspired to check their logs due to a similar incident at OpenAI. Closed-weight model incidents could elicit less public scrutiny, because they are less visible; external safety research would be restricted to less capable open-weight models, which may not generalise to larger frontier models. Frontier companies have the resources to control and prevent closed-weight incidents, whereas John Doe would struggle to stop his Kimi outbreak. Therefore a closed-weight incident must be significant to get past frontier companies, so the first uncontrollable closed-weight catastrophe could be significantly worse than the first uncontrollable open-weight catastrophe.

Fortunately, China doesn’t want to be behind in the AI race, and a pause in AI development would be an opportunity to diffuse AI capabilities while avoiding a reckless push towards superintelligence. A detailed explanation of what this might look like is included in the recently published AI 2040. Therefore, there is a vital opportunity for the USA - to make a deal with China to enforce a pause in AI development. This would enable both countries to cooperate and get the best outcome for everybody - a bit like the prisoner’s dilemma, except cooperation is clearly the best option:

Payoff matrix for the West and China each choosing whether to pause AI development: if both pause, the world goes well; in every other combination, everyone dies

In conclusion, the solution to avoiding a rogue agent explosion, or a catastrophic closed-weight incident, is having a pause deal with China. I think that pushing for a pause in the west will incentivise a deal between the west and China, and the best thing you can do to impact this is to email your MP on the PauseAI UK website.

Thanks for reading!

PS There is a lot of nuance within ‘a deal with China’, and I think a good outline of more specifics can be found in AI 2040, a plan for how AI development can go well. This includes mutually assured compute destruction, AI capabilities diffusion (enabling lots of countries to be at the forefront of AI development), and research transparency.

Other things worth reading (including comments): An AI Race with China Could Be Better Than Not Racing, The AI Race is Not a Prisoner’s Dilemma