
When AI researchers talk about “alignment”, they are talking about the process of getting an AI model to behave in a way that fits with human values. Preserving human lives. Respecting human dignity. Planning for a shared human future. Alignment researchers have struggled to instil these basic rules in AI models, but apparently the machines aren’t the only ones who missed the “humanity=good” memo. It seems the AI creators might need a reminder too.
I’ve written previously about the culture of dishonesty within AI companies: risk alerts squashed; grave dangers dismissed… dodgy practices that we had to learn about through tragedies and court cases. But when it comes to the question of human values, we don’t need an exposé to show us that some AI creators’ moral compasses are malfunctioning. They are telling us so themselves.
We got a glimpse of this on 9th September, when Jacob Coxon dramatically resigned from Anthropic claiming that AI companies are “gambling with our lives”. Even more disturbing was his colleague Evan Hubinger, who excitedly responded “we really do earnestly believe AI could kill all humans! I personally think it is >10% within the next decade”.
From an aligned model, the next words would be, “so we’ve stopped building it for now”. But we have an alignment problem on our hands. And not the digital kind. The warped values exhibited by AI models haven’t just sprung from cyberspace. They’ve grown from a primordial soup of reckless tech industry disruption culture and elitist disregard for ordinary public interests. In other words, from misaligned humans.
Take Elon Musk’s “MechaHitler” incident. I won’t rehash the details of it here; it’s not difficult to infer how a scandal with that name unfolded. The point is: this failing didn’t start with AI misalignment. It started with a human’s choice to build a model pandering to the unnerving whims of his followers. It started with one man’s preference for gratifying bad users over protecting the public. The AI’s harmful actions were founded on these distorted priorities which trickled down from its human progenitor into its “systemic ideological programming”. We may not know exactly how AI values emerge, but this episode reminds us that humans set the incentives around the model creation process.
Then there’s Sam Altman’s ongoing tightrope walk between doomsaying and techno-optimism. Back in 2015, he distastefully suggested, “I think AI will probably, like, most likely, sort of lead to the end of the world, but in the meantime, uh, there will be great companies created with serious machine learning”. That was the year he founded OpenAI. Now a seasoned CEO, he is more careful and image-savvy, emphasising the potential medical and educational benefits of his products to the UN security council, while also telling them, “we could lose control of the future to AI”. Not that he’s going to stop making it, or anything.
These guys’ attitudes do not reflect normal human values. Let’s acknowledge that. While everyone is worrying that machines might forget “people=friends”, these jolly harbingers of the apocalypse are proving we should be more concerned about their organic forebears.
But wait! In the context of AI leaders’ new commitment to “pace the frontier”, should we be open to the idea that these grinning reapers have turned over a new leaf? Well… commentators have expressed scepticism about this apparent pivot to prudence. It’s been called “too little, too late”.
Some say it’s an attempt at regulatory capture. Economist Christian Catalini called it, “virtue signalling and cheap insurance when things go wrong”. What I know for sure is this: misaligned machines don’t build themselves. We should not give credence to people who claim they are trying their best to keep us all safe while developing a technology which they say could be the end of humanity. It takes a loose cannon to release a rogue cannonball. So, let’s dampen the gunpowder. Let’s Pause AI.
Read about how you can act here.