US Edition
Your source for latest news
TechnologyAI Models

Microsoft's MAI-Transcribe-2 Claims Speed and Price Edge Over OpenAI, Google Rivals

Microsoft AI's newest speech-to-text model says it beats competing systems on speed and cost, intensifying competition in AI transcription.

PT
By PressTemps Technology DeskPublished Today, 09:02 ET · 3 min read
Microsoft's MAI-Transcribe-2 Claims Speed and Price Edge Over OpenAI, Google Rivals
Photo: Artotem / Flickr, CC BY 2.0 (illustrative image of a Microsoft office building)
What to know
Microsoft AI released MAI-Transcribe-2, pricing it at $0.10 per hour of audio through the end of 2026
Microsoft says the model ranks first on the 60-language FLEURS benchmark with a 5.2% average word-error rate and is up to 10x faster than OpenAI's GPT-Transcribe
New features include speaker diarization, word-level timestamps, code-switching and phrase-list biasing for domain terms
The model is in public preview via Microsoft Foundry, the MAI Playground and OpenRouter, part of Microsoft's push to reduce reliance on OpenAI's models

Microsoft's in-house AI division has released a speech-to-text model it says outperforms rivals from OpenAI, Google and ElevenLabs while charging a fraction of their prices. Microsoft AI announced MAI-Transcribe-2 this week, describing it as the fastest, most accurate and cheapest transcription system it has built, with a limited-time price of $0.10 per hour of audio through the end of 2026.

The company said the model ranks first on the FLEURS benchmark, a 60-language speech dataset built by Google researchers, posting a 5.2 percent average word-error rate across those languages. On the independently run Artificial Analysis speech-to-text leaderboard, Microsoft said MAI-Transcribe-2 placed second by accuracy while recording the fastest processing speed of any model tracked, turning roughly an hour of audio into text in about ten seconds. Microsoft said the model is up to ten times faster than OpenAI's GPT-Transcribe, seven times faster than ElevenLabs' Scribe v2, and five times faster than Google's Gemini 3.5 Transcribe.

According to a companion post on the Microsoft Foundry engineering blog, the new model adds capabilities its predecessor lacked, including speaker diarization, word-level timestamps, automatic language identification, code-switching for conversations that mix languages, and "phrase-list biasing" that lets developers prime the system with domain-specific vocabulary such as medical or legal terms. Transcripts can be returned in either a verbatim mode or a cleaned-up style that strips filler words, and Microsoft said the model holds up better than earlier versions on noisy, real-world audio rather than studio-quality recordings.

Pricing pressure on rivals

The model card published by Microsoft AI lists the $0.10-per-hour introductory rate as a steep cut from the $0.36 per hour it charged for MAI-Transcribe-1.5, and well below list prices commonly charged by competing transcription APIs. The model is available now in public preview through Microsoft Foundry, the company's Azure-hosted AI developer platform, as well as through the MAI Playground and the third-party router OpenRouter. Microsoft pitched the release for uses including clinical note-taking, legal documentation, closed captioning and voice-agent applications, where transcription cost and latency directly affect how the tools can be deployed at scale.

The release is the latest in a string of in-house "MAI" models Microsoft has shipped over the past year, following MAI-Voice-1, MAI-Image-2 and the original MAI-Transcribe-1, part of a broader effort by the company to reduce its dependence on models built by OpenAI, its longtime partner and now also a rival supplier of frontier AI systems. Independent outlets covering the release, including a report from Unite.AI, noted that Microsoft's benchmark claims rely on the company's own testing and the Artificial Analysis leaderboard rather than a single universally agreed measure, since transcription accuracy is commonly evaluated across several benchmarks that do not always rank the same systems in the same order.

Speech-to-text has become a competitive front in the broader AI race as voice interfaces, meeting-transcription software and voice-controlled agents proliferate, with OpenAI, Google, ElevenLabs and smaller specialists all racing to cut latency and cost per hour of audio processed. Lower per-hour pricing matters disproportionately for products that transcribe continuously, such as call-center analytics platforms or always-on captioning services, where costs scale directly with the volume of audio processed rather than with a fixed number of user queries.

Microsoft did not say how long the promotional pricing would last beyond this year, and the company's public documentation notes the model remains in preview rather than general availability, meaning enterprise customers evaluating it for regulated uses such as clinical or legal transcription will likely wait for a full production release, along with the compliance certifications those industries typically require, before wide deployment.

More on this story

All Technology