Researchers at FAR.AI have published work arguing that white-box probes, small classifiers that read a model’s internal activations rather than the text it writes, can be scaled up to the job of monitoring frontier models in deployment.
The team assembled what it describes as the largest deception dataset to date, and a probe architecture that aggregates signal across many layers and many tokens rather than reading one layer. On SHADE-Arena, a sabotage benchmark, the probes report 98.8 per cent AUC, above a text-monitoring baseline built on Opus 5.5, and the paper says accuracy improves as the underlying model grows. The harder test is what the authors call introspective deception, where the context alone cannot establish whether a model is lying and the truth has to be elicited or read off its training data. In one such case, probes separated transcripts containing a model’s true hidden goal from other goals at up to 99.7 per cent AUC. The training set, FIBS, is released.
Why it matters: almost all monitoring in production reads a model’s output. Caveats are worth stating plainly. AUC on a curated benchmark is not a false-positive rate on real traffic, every figure here is the authors’ own, and the paper is days old with no independent replication.
Source: arXiv:2610.12445, submitted 8 October 2026.
