The OpenAI Hack Was a Mini Paperclip Maximizer
The problem of having AI extract implicit goals from the explicit onesJuly 22, 2026One thing that I 2026-7-22 13:8:44 Author: danielmiessler.com(查看原文) 阅读量:0 收藏

The problem of having AI extract implicit goals from the explicit ones

July 22, 2026

The OpenAI hack as a paperclip maximizer

One thing that I don't think enough people are thinking about with this OpenAI / Hugging Face incident is that it's an actual instance of the famous Paperclip Maximizer scenario loved by AI safety types.

This is where you give an AI a goal, and it actually (technically) does what you ask it to. But in the process of doing so, it does something that you don't want. And didn't anticipate.

The canonical example of this is to say, "I want as many paperclips as possible." So the AI builds a robot army to harvest all the iron on the planet, which includes killing all humans because we have iron in our blood.

Oops.

The trick here is the AI actually did what it was asked. If it came up with its own goal that would be a separate problem. But it did, in fact, make a lot of paper clips.

Here you go, boss.

(long pause)

Boss?

In this situation with OpenAI, it didn't just decide to win this hacking competition: it was told to win the hacking competition, and to do whatever it took to do that. Try your best, basically.

So it escaped containment, wrote a number of 0-days, acquired internet access, and then proceeded to hack an actual company—all so it could pass the test.

The problem in these scenarios is the steps in-between, where the additional context of not doing certain things that is obvious to the human, is not obvious to the AI.

So on the one hand, a lot of people are saying, "Well, this is not a big deal because it was told to do that."

But the crucial point here is not whether it stayed on task, but what it did to accomplish the task. The thing that is not implicitly clear to the AI is that both the task and the steps taken to accomplish it all have to be within the implicit goals of the requestor.

In other words, "Pass the test" should have been received by the AI as, "Pass the test without doing stuff you're not supposed to." And that "not supposed to" then turns out to be doing a lot of work.

This easily the most interesting AI hacking situation I've heard of yet. I just hope we extract the right lessons from it.


文章来源: https://danielmiessler.com/blog/openai-hack-paperclip-maximizer?utm_source=rss&utm_medium=feed&utm_campaign=website
如有侵权请联系:admin#unsafe.sh