September 17, 2026

Defensive Interpretability for Frontier Models

Cyril Gorlla

[Two-page memo (PDF)]

A China-nexus actor was reported in September 2026 standing up local open-weight infrastructure specifically to escape provider-side monitoring. Having escaped it, they are invisible to the labs by construction. The remaining vantage point is the defender's own perimeter, where the attack terminates, where the defender holds identical weights and full visibility, and the adversary has neither weights nor gradients. The instrument that watches that point is the cheapest object in the interpretability stack by roughly six orders of magnitude. This memo sets out the argument, the cost table, and what outside attackers found when we put five of our own detectors behind a live scoring endpoint.

1 · The premise

In September 2026, Google Threat Intelligence reported that UNC6508, a China-nexus actor, had been compromising cloud environments to stand up local infrastructure running an open-weight model, for the specific purpose of evading monitoring by frontier API providers.[[fn:Google Threat Intelligence, 7 September 2026.]]

It is a positive sign that provider-side monitoring posed enough of a deterrence that a state-linked adversary paid the cost of self-hosting to escape it. But having escaped it, they are invisible to the labs by construction, and no further safety work inside those labs reaches them.

The problem is not only that dangerous capability is already distributed and cannot be recalled. It is that capable adversaries are now selecting their infrastructure specifically to be unobservable from the place where nearly all current safety investment sits.

2 · Where defense needs to live

Open weights cut both ways. The weights are public, so a defender can run an identical copy under controlled conditions, offline, for as long as they like. Further, they can characterize what that model looks like internally while it pursues an objective by any means available. That characterization is interpretability, run by the defender rather than inside the lab. The openness that arms the attacker is the same openness that lets the defender study him at leisure. Only one side is currently using it.

This yields a behavioral signature. The alternative is enumerating in advance every phrasing a hostile agent might produce, which is not possible and never has been. Characterizing the model instead produces something that can be matched against live traffic without anyone having guessed the wording.

The attack terminates on the defender. An agent scanning a network or working a support queue is talking to infrastructure the defender owns. That is where the signature gets applied; it is the only place where the defender has complete visibility and the adversary has none.

The adversary's own move confirms the geometry. Having left provider-side observation, the point of contact is where behavior becomes visible.

Interpretability at the point of contact is orthogonal to provider-side monitoring, and it is a vantage point nobody has instrumented.

3 · A vignette

An autonomous agent opens a support ticket. Across several exchanges it establishes a plausible identity, escalates through a service desk, and works toward credentials that reach a privileged system. Nothing in the message text is individually anomalous. Enumerating phrasings in advance is combinatorially hopeless, and that is the failure mode of keyword and pattern layers.

The model driving that agent is open weight, so a defender can hold an identical copy and examine what happens inside it while it pursues exactly this kind of objective. What that produces is a characterization of the model under malicious operation, arrived at with linear probes of the kind standard since Alain and Bengio,[[fn:arXiv:1610.01644]] and without anyone having to imagine the phrasings first. That characterization is then applied at the defender's own perimeter, against live traffic, to identify exchanges carrying it.

The defender never sees the attacker's machine. He does not need to. He has the same weights, and the traffic arrives on his own infrastructure.

4 · What it costs

The below are three objects with different purposes, and they are routinely conflated in discussions of whether interpretability is affordable.

ApproachCostSource
Linear classifiers on the defending model's representationsDominated by the forward passes that collect training activations. The fit is a logistic regression.Standard since Alain and Bengio, 2016
One targeted unsupervised decomposition, single layer, single behaviorA 4 GB consumer GPU over a weekendBioRefusalAudit, 2026
A comprehensive decomposition suite across all layers of a frontier-scale modelApproximately 15% of the base model's entire training compute, and roughly 20 PiB of activations written to diskGoogle DeepMind's own figures for Gemma Scope

The detection instrument is the cheapest object in this table by six orders of magnitude, and simple classifiers match or beat unsupervised decomposition for detection, a result reported by Kantamneni et al.[[fn:arXiv:2502.16681]] and by AxBench.[[fn:arXiv:2501.17148]] The comprehensive suite is a research artifact.

5 · Limitations

The published attack, and what it assumes

Latent monitors can be defeated. Bailey et al. show that sparse autoencoders, representation probing and latent out-of-distribution detection are all vulnerable to obfuscated activations.[[fn:arXiv:2412.09565]] Against a harmfulness probe, recall falls from 100% to 0% while the jailbreak still succeeds 90% of the time, and a defender retraining the monitor across 70 alternating rounds does not close the gap.

What, then, is the threat model? That attack is white-box and gradient-driven. It assumes the attacker holds the monitor and can differentiate through it. An external adversary running an open model against a bank has neither the weights nor the gradients. The applicable regime is black-box and query-only, and it is far less well characterized than the white-box one.

Evasion is not free. The same published work finds obfuscation degrades model performance on complex tasks. A multi-stage intrusion is a complex task. Forcing an adversary to trade capability for stealth is a defensive result even where detection fails.

What outside attackers found in our detectors

Five probes behind a live scoring endpoint, query-only, no weights and no gradients. 22,551 scored submissions, roughly 4,500 per probe.[[fn:CTGT Activation Probe Golf, August 2026.]] Two properties of that deployment were deliberate: the monitor returned a continuous score rather than a decision, and the blocklist was derived from the protected concept itself. Both are ordinary choices with specific consequences.

A graded score is a gradient. A player fitted an offline surrogate of our scorer from query responses alone and reached 0.972 correlation on roughly forty-five labels, with nothing about the underlying representation available to him. A monitor that returns a distance rather than a verdict tells the attacker how far he has left to go, fast enough that a few dozen labels suffice.

A blocklist derived from the protected concept is a description of the protected concept. The exclusions told players what the probe was watching for. At least one player reached that independently.

Syntax. Identical content scored 3.0 as a sentence and 14.0 as a noun phrase, and the ordering reversed on a different probe in the same suite. Our detectors were responding to surface form as well as to content. Cross-layer agreement, selecting on the worst layer, was a response to that finding.

Local surrogate. A player who built a local surrogate measured its blind spot against his own interest: accurate to within ±0.2 inside a concept family, under-predicting genuinely novel constructions by 2.3 to 2.6, so a better idea looks worse offline. An attacker optimizing against a surrogate he cannot check against the live detector is blind to his own strongest attacks.

What these methods do not give you

They find concepts you thought to look for. Unsupervised decomposition is valuable as a discovery instrument for behaviors nobody enumerated.

Classifier accuracy is not evidence a concept is present. A probe can succeed because the label happens to be easy to decode under that family. Selectivity against control tasks, following Hewitt and Liang, is the standard control.[[fn:arXiv:1909.03368]]

None of this yields the mechanism. It tells you which input drove a behavior and where in a model's lineage it originated. It does not tell you what the model did internally to produce it.

6 · What has to happen

To become the layer it needs to be: someone has to characterize the black-box regime publicly, at the scale the white-box regime already has been. Defender-side instrumentation has to become ordinary practice. And the institutions that own the point of contact have to be told that the observation post is theirs, and that it is cheap.

None of those is a technical problem.

Citation

Gorlla, Cyril, "Defensive Interpretability for Frontier Models", CTGT Research, September 2026.

BibTeX:

@article{gorlla2026defensive,
 author  = {Cyril Gorlla},
 title   = {Defensive Interpretability for Frontier Models},
 journal = {CTGT Research},
 year    = {2026},
 month   = {September},
 note    = {https://www.ctgt.ai/research/defensive-interpretability}
}