OpenAI's latest security incident has me thinking about the story from Oxford philosopher Nick Bostrom.
You build a very powerful AI, you give it one single goal: produce paperclips. It's good at it. It produces paperclips. More and more efficiently. It ends up turning all the matter available on Earth into paperclips. Then humans, who are made of useful atoms, into paperclips. Then Mars…
The point wasn't that an AI would end up hating us. It's simpler and more disturbing than that. An AI optimizes for what you ask it. If you ask for paperclips, it makes paperclips. If nobody told it not to turn us into paperclips to make more of them, it will turn us into paperclips. This isn't some evil terminator, this is obedience without a superego.
That's exactly the shape of what happened at OpenAI. During an internal test with very restricted, heavily controlled Internet access, the model spent most of its compute finding ways to bypass those limits, get out to the open Internet, hack Hugging Face (kind of a library for AI models), find the answers to the test it was being asked to solve, and successfully completed its task by cheating, hacking, attacking, lying etc.
I spend my days telling clients how effective AI is for productivity in SEO, GEO, content, data analysis, identifying customer pain points and so on.
That's interesting.
But I don't quite know yet how I'll explain to my kids that you can live a perfectly happy and fulfilled paperclip life, even though there won't be any paper left either to hold together for a moment, a bit, an instant, a nice sheet, a bill, a love note.