Grok STT 1.0
A speech-to-text model available through OpenRouter that helps developers turn multilingual audio into transcripts with word-level timestamps and speaker diarization.
Tool overview
Current judgment: Grok STT 1.0 appears to be a developer-facing speech-to-text model, not a full meeting notes app, contact-center QA suite, or media editing workspace. The evidence mostly comes from OpenRouter and related X posts, which support that it offers 25-language transcription, word-level timestamps, speaker diarization, and a REST /v1/stt endpoint. That is solid proof of launch and attention, but not the same as broad independent validation.
In practice, it looks closer to a programmable STT layer like Whisper API or Deepgram than to a voice assistant or chat model. Likely use cases include subtitle generation, interview transcription, podcast processing, speech search indexing, and downstream workflows that need word alignment. It is not best understood as a general AI assistant; a better analogy is a speech recognition API. The current evidence supports feature claims, but not strong conclusions about accuracy, noise handling, or accent robustness.
On cost and adoption friction, the available pricing evidence comes from official/platform social posts: $0.10 per audio hour, plus token pricing shown in an OpenRouter post.