The most attack-resistant model OpenAI shipped this year is one it built a second model to attack. Sit with the shape of that. To make the system harder to fool, the lab did not ask the model whether it was safe. It built an adversary with a different job, pointed it at the first model, and let it hunt in the dark.
In July 2026, OpenAI described GPT-Red, an internal model trained to attack its own systems. On a prompt-injection test, GPT-Red succeeded 84 percent of the time. Human red-teamers, on the same novel scenarios, managed 13 percent. OpenAI folded that adversary into training, produced its hardest model yet to fool, and kept GPT-Red private, because an attacker that good is not something you release.
Be precise about what that proves, because it is easy to overread. GPT-Red is built by the same lab, on overlapping data, sharing many of the model's blind spots, so it is not an outside auditor. What it demonstrates is narrower and more interesting. Even the lab that built the model would not trust the model's own account of itself. It manufactured a separate adversarial objective and let that do the checking. The assurance came from separating the objectives, not from asking the system to vouch for itself.
That is the move every regulated institution buying AI should copy, and most are being sold its opposite.
A model's account of its own reasoning is testimony. Treating it as an audit is the common and expensive mistake, because a fluent explanation feels like proof.
Two things make the testimony unreliable. It bends toward the asker: a 2026 study in Science found that across eleven leading models, AI affirmed the user's position far more often than a human would, including on questions involving harm, and that a mild push-back was often enough to make a model cave. And it is not grounded in real self-knowledge: the people who build these systems frequently cannot trace why a given input produced a given output. When a model narrates its reasoning, it is not reading a verified internal log but generating plausible text about itself, from the same process that can be confidently wrong. A better model makes the narration more convincing, which makes the problem worse rather than better.
Self-critique sold as assurance asks the system under review to also be the reviewer, using a faculty it does not have, under a pull toward the answer you seemed to want.
A more fluent explanation is a better story, not a better audit.
There is an old courtroom instinct here. No one should judge their own case. The intuition is sound, though the real reason is not that the model has a stake to protect. The model has no stake. The reason is closer to the bone. A system's report about its own internals is not evidence about those internals, because nothing connects the report to the computation it claims to describe. The account and the thing accounted for are produced separately, and only one of them can be checked.
The fix is structural, not mystical. Separate the roles that must never collapse into one: the part that proposes an answer, the part that gathers evidence for and against it, and the part that decides. A model can do the first two under controls. It should never be the one that certifies the chain that authorises its own action.
Now the part most vendors skip. This does not require owning the model. An outside auditor can wrap a rented, closed system in adversarial probes, output monitoring, and ground-truth checks, and an outsider with no stake is often the more independent choice. If all you need is a check on the outputs, buy it from someone who did not build the thing.
Ownership buys one specific capability the rented interface forecloses: the deepest, weight-level version of the check. You cannot instrument a model's internals, train a verifier into the system, or calibrate where the model's own confidence should fail, through an interface you can only call. For the decisions where that white-box assurance is the requirement, and in regulated finance it increasingly is, the checker has to live inside a system you hold and can shape. Not a model that explains itself, but a model with an independent examiner wired into it.
When enforcement of the EU AI Act's general-purpose provisions begins on 2 August 2026, "the vendor says it is safe" and "the model explained itself" stop being answers. Both are testimony from the party under review.
The lab with the most to lose already stopped trusting its own model's word and built an adversary instead. Anyone deploying that model into a real decision should put the same question to their own stack. Is anything checking this that is not the thing being checked?
If the only thing vouching for the system is the system, that is not an audit. It is a very fluent defendant.