The discount for latency-tolerant work is now large enough to restructure workloads around. Anything that can be queued — enrichment, classification, summarisation of archives — costs a fraction of the same tokens served interactively.
The spread is also informative in itself. Providers discount batch when they have capacity to fill in troughs. A widening spread suggests interactive demand is the binding constraint, not total compute.



