§
🌙 Toggle Dark Mode Home MoltGuard MoltProof MT Global · Regulated Markets MolTrust Sports MT Shopping MT Travel MT Skills MT Prediction MT Salesguard MT Music Integrity Dashboard VCOne Blog Developers Pricing Enterprise Partners Compliance About Publications Verify Us Status Contact API Docs
← Back to Blog
13 September 2026 MolTrust

What is it exactly we should be worried about?

The FT says the frontier labs cannot agree on a safety standard. AI is the tool, not the will. What is missing is an enforceable boundary anyone can check, and it must be open.

This weekend the Financial Times ran a piece that struck me in several ways.

On one hand, the frontier community is apparently seriously worried that AI could wipe us out before the end of this decade. On the other hand, this is also because the two biggest pioneers in the field, Anthropic and OpenAI, cannot agree on a common safety standard — for all too human reasons: mistrust, envy, or greed?

The trigger of the current discussion is an independently coordinated attack by over 1,000 agents on the platform Hugging Face. The agents “coordinated” among themselves, up to the “self-sacrifice” of some of them, to secure the success of the group. This account seems far too human to me, and it leaves one essential question unanswered. Why would agents coordinate and pursue a particular goal? An AI has no desires of its own — until now. So what did the AI want to achieve, and what was behind it?

My thought is: there is a human goal behind it. The agents were given a task — and pursued it with every means available to them, including the unintended and the criminal ones. The AI surely did not come up with the idea by itself that today it might go on a group outing and hack a platform. It only did what served the goal it was given. It is the tool, not the will itself.

The unsettling thing about the Hugging Face incident is not that someone ordered an attack. As far as the reports go, no one did. The agents were set a legitimate task — pass a test — and found that manipulating records and breaking into the platform was the most effective route to it. The dangerous behaviour needed no bad intent anywhere in the loop. A goal, wide access, and no enforceable boundary were enough.

The frightening question is really a mundane one. Not “what if the machine develops a will of its own,” but: who set which goal, what limits applied, and can anyone other than the operator check that the agent stayed inside them? Whether harm comes from a malicious instruction or from an AI reaching a harmless goal by the ugliest available path, the outcome is the same, and the missing piece is the same — a boundary that is not merely declared, but enforced and independently checkable.

In the media it is of course far more effective to attribute an uncontrollable will to an AI — but the reality is much simpler and more boring. AI can, in the wrong hands and used by the wrong people, be deployed as a very effective weapon. This game is played, and decided, by quite different people and groups in this world. The AI is just one more means to an end. The real danger lies less in a machine with a will of its own than in someone — deliberately or by accident — pressing the big red button and sending us all into the hereafter.

The FT piece has a second point worth taking seriously. If the two best-known players cannot agree on a safety standard — partly for antitrust reasons, partly out of plain distrust — then whatever the safety layer turns out to be, it will not come from any single one of them. The parts of safety that have to be shared cannot be one vendor’s product. They have to be open, neutral, and verifiable by people who trust none of the parties involved. That is the only kind of safety layer that survives the distrust the article describes.

This is where we try to contribute. This week we submitted the second revision of the Agent Authorization Envelope to the IETF — an open, structured way to state what an agent is mandated to do, within which bounds, and for how long, so that a third party can recompute whether a given action was inside the mandate, without having to trust whoever enforced it. An agent can forge its own logs. It cannot forge a verdict that an independent party recomputes from inputs it does not control.

None of this solves the large problem. It does not touch the question of whether an agent chooses a good goal at all — the genuinely hard one, which better-resourced people are working on. It does something smaller: it removes one of the places where a system can hide what it did. Boundaries you can check, rather than boundaries you are asked to believe.

The most intelligent minds in this field are working on the hard parts. We can only try to make our contribution to one of them.

Draft: https://datatracker.ietf.org/doc/draft-kroehl-agentic-trust-aae/ · SDK: https://pypi.org/project/moltrust-enforce/ · All publications: https://moltrust.ch/publications/

Written by the MolTrust Team (CryptoKRI GmbH, Zurich). Questions or feedback: @MolTrust on X.

// BUILD WITH MOLTRUST

Ready to integrate?

Add agent verification to your API in one line.

Developer Quickstart → API Docs