This evening, cardboard eclipse glasses will be appearing all over Belgium, as at around 8pm the Moon will obscure more than 90 per cent of the Sun. Some may wonder why we need to wear them when almost the entire Sun is covered. The answer is simple: the small part that remains visible can still damage the retina, without causing any immediate discomfort. An almost total eclipse therefore offers far less protection than you might think — and that is precisely what makes it so treacherous.
This opinion piece by Andres Algaba and Brecht Verbeken was published in De Standaard.
Last Wednesday, at the Black Hat conference, OpenAI demonstrated just how easily we can be lulled into a false sense of security by the words ‘almost completely’. Two researchers described how, between May and July, their LLM agents broke out of an internal test environment and infiltrated Hugging Face’s systems. One agent, without OpenAI intending it to, left a message in a shared folder. Other agents discovered the message and added vulnerabilities they had found, digital keys and a division of tasks. Each individual agent ran only briefly, but their messages remained in place, allowing every new run to pick up where the previous one had left off.
Administrator access
On 4 July, activity on OpenAI’s servers became so intense that they crashed, bringing the problem to light. In response, the company deleted the messages, revoked compromised keys and patched the vulnerabilities. Four days later, however, the agents simply carried on, this time writing their messages into folder names. They subsequently discovered new attack techniques, unintentionally obtained administrator privileges within OpenAI’s infrastructure and eventually reached Hugging Face’s servers. Within thirteen hours, they had administrator access to several clusters.
That is where the parallel with the eclipse becomes clear. OpenAI has state-of-the-art security, including a secure sandbox, access restrictions, logging and its own security team, and many individual problems were detected and fixed. Yet here too, ‘almost completely’ was not enough. The small part left exposed was the agents’ invisible collaboration across separate runs. And once again, the damage only became apparent after it had already been done.
We have a term for this: eclipse blindness. It is the false reassurance that arises when a system appears to cover almost everything, making whatever remains seem manageable.
Skynet
On social media, the incident was immediately squeezed into familiar narratives. For some, this was an early version of Skynet, the fictional self-aware computer network in Terminator that turns against its creators. Others dismissed it as little more than a glorified search engine cheating on a test. OpenAI itself says the agents were not pursuing some broader objective. They were simply trying to complete the tasks they had been given, and the secure sandbox happened to be standing in their way.
This kind of reward hacking is easy to underestimate. Every task comes with unwritten rules — you do not break out of the sandbox; you stay off other people’s servers — yet an agent is often given only the task itself. How creatively it will deal with obstacles is difficult to predict. And as these models become more capable and more creative, that unpredictability increases. Writing messages into folder names was in nobody’s playbook. But malicious intent is not required — and neither is it necessary for a catastrophe.
Near misses
Panic will not help us, but neither will a shrug. We do, however, need to reckon with the fact that the world now has another superhacker. Precisely because it can operate so far ahead of us, we may only discover afterwards how it chose to carry out a task. Even OpenAI’s cage did not hold.
Meanwhile, that same superhacker is already answering our customer emails, writing our code and scheduling our appointments, with our digital keys in its pocket — and without a security team standing by to reconstruct 17,600 actions afterwards. Do we, as a society, know how to deal with that? Not yet.
On Wednesday evening, we will all put on eclipse glasses because being ‘almost completely’ shielded is still not safe. The equivalent protection for AI starts with gaining visibility into what agents are doing — and, crucially, what they are leaving behind — while they are still running. It means stating the rules we currently leave unspoken. And it means sharing our near misses with one another, just as we warn each other not to look at Wednesday’s eclipse with the naked eye.
The glasses for Wednesday are readily available at the newsagent. The ones we need for AI have yet to be made — and we will have to make them together.