xAI claims halved transcription errors at the same price

xAI released Grok Voice Transcribe 2.0 on 18 September, a speech-to-text model it says is twice as accurate as its predecessor at the same price: $0.10 per hour of audio for batch transcription and $0.20 per hour for streaming.

The feature list is the competitive part. Speaker diarisation is included at no extra charge, which several rivals bill separately. The model also does word-level timestamps with confidence scores, up to eight-channel audio, key-term biasing of up to 100 terms per request, filler-word removal, and automatic language detection that follows a switch mid-recording.

On accuracy, separate the two kinds of claim. The headline figure — word error rate on short phrases across 19 languages falling from 20.6% to 6.8% — is xAI measuring its own models against each other, which says something about the upgrade and nothing about rivals. The one external claim is a first place for accuracy among 32 streaming models on the Artificial Analysis leaderboard, a test xAI did not run. It also says it leads on 8 kHz telephony audio, the sort of thing customer support recordings are made of.

The published comparison chart names competing models but not their scores, so the size of any lead is not stated.

Source: xAI, “Grok Voice Transcribe 2.0”.


Related