Britain’s AI Security Institute has released Transect, an open-source Python tool that turns an agentic evaluation transcript into an interactive report. The problem it addresses is mundane and real: transcripts can run to billions of tokens, and a final score reveals little about how the agent got there.
Transect puts activity labels and token use on one timeline, so a reviewer can follow a run and open the passages behind each label. It is built on Inspect Scout, the institute’s existing transcript-analysis library, with support from Meridian Labs. The code is on GitHub and a paper on arXiv.
The labelling is done by language models acting as judges, and AISI says plainly that their judgements can be wrong or misleading. That is why the tool surfaces both the supporting passages and how the analysis was produced, rather than only the verdict. The reviewer is expected to check the judge, not trust it.
This is infrastructure rather than a finding, which makes it easy to skip past. But the bottleneck in evaluating agents is no longer running the evaluation, it is working out what happened inside a transcript too long for anyone to read. A regulator releasing its own reading tools also makes its evaluations easier to argue with. No licence is named, so check the repository before reuse.
Source: AI Security Institute, Transect: Making large-scale agentic evaluations easier to understand.
