Industry
23 Nov 2026
X
Min read

Welcome to medical speech, ElevenLabs

ElevenLabs launched Scribe v2 Medical, a speech-to-text model built specifically for clinical audio. We're pleased, genuinely, because the launch lands on a conclusion we've been talking about for a while. Getting most words right isn't enough when the words you miss are drug names, dosages, measurements, or clinical terminology. A model can post a strong overall Word Error Rate (WER) and still miss the words that matter most in a clinical setting. That was one of the problems we set out to investigate with Symphony.

Medical speech needed better benchmarks

When we published our Symphony research in May 2026, we ran into a second problem underneath the first. We didn't have enough reliable, reproducible datasets to evaluate medical speech recognition. So we built them.

MedDictate and MedTerm are the result. MedDictate tests clinical dictation, including expected output formatting, across specialities and languages. MedTerm puts greater pressure on medical terminology and structured clinical language.

We released both datasets publicly. The goal was to give developers and researchers better tools to test medical speech systems and advance the field, not just to prove Corti performed well. That distinction matters. A benchmark you only use on yourself is a marketing asset. A benchmark other teams can pick up is infrastructure for the field.

Four months later, ElevenLabs used both datasets to evaluate Scribe v2 Medical. That is exactly what we hoped would happen. You can read the full Symphony research on arXiv.

We agree on a lot

ElevenLabs' results reinforce several findings from our research. General-purpose speech recognition struggles with clinical language. Overall WER doesn't tell you enough about clinical performance. Medical terminology needs specific examination. Context matters enormously when recognizing difficult clinical terms.

They highlight the failure mode plainly. A drug such as etodolac, presented without surrounding clinical context, can become something entirely different. Every clinician reading this can picture what happens next when that error makes it into a note.

None of this is a disagreement. It's confirmation from a second team, working independently, reaching the same place. That kind of agreement is worth more than a benchmark win, because it says the problem is real and not a quirk of one lab's setup. Getting the medical words right, though, is only the beginning.

This is where it gets harder

A medical speech system doesn't operate neatly on a finished audio file. Clinicians dictate in real time, have conversations with patients, correct themselves, and use abbreviations, specialty terminology, drug names, dosages, and units. Some workflows need an immediate transcript. Others can run asynchronously. The output must be clinically useful.

Our Symphony research looked beyond raw transcription accuracy. We evaluated a system built around separate stages for speech recognition, clinical formatting, and contextual correction, across both real-time and asynchronous workflows. It also examined medical terminology and structured entities directly. Recognizing a word is one problem. Knowing it's a medication, preserving the dosage and unit, formatting it correctly, and doing all of that while the clinician is still speaking is another.

ElevenLabs' launch focuses on transcription performance: WER, medical-term recall, and term-specific WER. Those are useful measures. They don't yet show how the system handles structured formatting, contextual correction, or live clinical use.

One more thing. Their comparison used MedDictate and MedTerm to put Scribe v2 Medical against several other speech models, but left Symphony out.

Let’s add Symphony back in:

Here's the honest caveat. The metrics in our original research and those in ElevenLabs' evaluation aren't all directly comparable. Whole-transcript WER and WER calculated over annotated medical terms measure different things. Rather than set two mismatched numbers side by side and declare victory, we've recomputed our own model's results using the same metric definitions ElevenLabs reported (term recall and term-WER) on the published MedTerm subset, so the numbers we're showing measure the same thing theirs do. If you advocate for better benchmarks, you should use them properly.

The most interesting thing here isn't Corti versus ElevenLabs. It's that another major speech company has decided healthcare needs dedicated models, dedicated evaluation, and a deeper understanding of clinical terminology. That's good for the field, and it's why we made our benchmarks public. Welcome to medical speech. We're genuinely glad you're here.

Run your own testing with the comparison tool, or read the research.

Build faster. Ship safer. Scale smarter.

Get started with healthcare-native APIs built to power real clinical workflows.

More stories from Corti

View all