More informative than benchmark percentages, because it captures what practitioners actually care about: not whether an agent can do a task, but how far it can go before it goes off the rails.
Published estimates put current frontier models in the range of a few hours of human work at a fifty per cent success rate — measured on software and technical tasks, which are agents’ best domain.
