Accuracy, measured
How accurate is Dictro, and how would you know?
Every dictation app on the market claims high accuracy, and almost none of them show you how they got the number. Here is how word error rate actually works, exactly how we measure our own, and why there is no single percentage at the top of this page.
Apple Silicon · macOS 15+ · Free plan · 100% on-device
The measure itself
Word error rate, briefly
Record something. Write down exactly what was said — that is the reference. Run the recogniser and compare its output word by word, counting three kinds of mistake: words it swapped for the wrong word, words it dropped, and words it invented. Add those up, divide by the length of the reference, and you have the word error rate.
A WER of 5% means one word in twenty came out wrong. Lower is better, and the difference between 3% and 5% is very noticeable in practice because errors do not spread themselves evenly — they cluster on names, technical terms and anything unusual, which is exactly the vocabulary that carries the meaning of what you said.
Our method, in full
How Dictro measures its own accuracy
The harness ships inside the app and runs as Dictro eval. It takes a folder of test cases — each one an audio file paired with a text file holding exactly what was said — and runs the real pipeline over them, not a simplified stand-in.
Scored per stage
Raw recognition, then after the custom-vocabulary booster, then after the optional second-pass check. A change can therefore be attributed to the stage it actually affected.
Normalised before scoring
Both sides are lowercased and stripped of punctuation, with apostrophes kept inside words. The score measures which words came out, not capitalisation or commas.
A known caveat
Numerals and spelled-out numbers count as different words, so “2” against “two” scores as an error. That is documented in the scorer rather than quietly smoothed away.
Punctuation is excluded deliberately: the cleanup stage is allowed to repunctuate freely, so scoring commas would penalise the thing that makes the output usable. What the number measures is whether the right words survived.
The honest part
Why there is no percentage at the top of this page
We have measurements. They are from audio we recorded ourselves — one speaker, one microphone, one accent, a recording set that is not published and that you could not obtain to re-score. That makes it a good instrument for answering “did this change help?” and a bad basis for a public claim.
Our own testing makes the point: changing the microphone alone moved the result by around a percentage point, which is larger than the gap between several products advertising against each other. A number that sensitive to conditions we did not publish is not a number you should have to take on trust.
Our comparison pages refuse to state a competitor's figures without a source and a date. It would be a poor trade to apply a looser standard to ourselves. When we can publish a recording set that a stranger could run and re-score, the result goes here — including the cases where Dictro does worse.
Reading other people's claims
Why two accuracy numbers are rarely comparable
A WER figure means nothing without the audio and the scoring rules behind it. All of these move the number, often by more than the difference between products:
| Variable | Why it moves the number |
|---|---|
| The recordings | Clean read-aloud prose scores far better than real speech with false starts and thinking-out-loud |
| Speakers and accents | A single-speaker set measures how well a model handles one voice, not how well it handles yours |
| Microphone | In our own testing, worth about a percentage point on its own |
| Normalisation | Whether numerals, punctuation, casing and contractions count as errors changes the result without changing the model |
| Which stage is scored | Raw recognition and post-cleanup output are different numbers from the same run |
A published figure is worth something when the dataset and the method come with it and can be reproduced — some vendors do exactly that, and it deserves credit. A percentage on a landing page with no method behind it is a design element, not a measurement.
What to do instead
What actually makes dictation accurate for you
The honest answer is that for most people the model is not the deciding factor, because the leading models are close together and the conditions around them are not.
Your microphone
The largest single lever most people have, and the cheapest. A basic external mic reduces errors more than switching apps usually does.
Your vocabulary
Names, product names, project jargon and library names are words no general model has seen. Dictro learns the ones you correct, which is where the errors that annoy you actually live.
The cleanup step
A transcript containing every “um” can be perfectly accurate and still unusable. Usable text is a different measure than a correct one.
Which is why the only test that settles it is your own voice, your own microphone and your own words, for a week. That costs nothing on the free plan — 10,000 words a week, no card. For how the pipeline works, see on-device dictation, or the sourced comparison of Mac dictation apps.
Frequently asked questions
What people ask about Dictro
What is word error rate?
Word error rate, or WER, is the standard measure of speech recognition accuracy. You take a recording, write down exactly what was said as a reference, run the recognizer, and count how many words it substituted, deleted or inserted. Divide that by the number of words in the reference and you get the rate. A WER of 5% means one word in twenty came out wrong. Lower is better.
How does Dictro measure its own accuracy?
With a harness built into the app itself, run as 'Dictro eval'. It scores word error rate at each pipeline stage separately — raw recognition, then after the custom-vocabulary booster, then after the optional second-pass check — so a change can be judged by which stage it actually improved. References and output are both normalized before comparison: lowercased, punctuation stripped, apostrophes kept. That means the score measures which words came out, not capitalisation or commas, which the cleanup step is allowed to change anyway.
Why does this page not state a single accuracy percentage?
Because our current measurements come from audio we recorded ourselves, and a number from a private single-speaker recording set is not something you can check, reproduce, or fairly compare against another product's number. It is a good instrument for deciding whether a change helped. It is not evidence for a marketing claim, and publishing it as one would be exactly the kind of unsourced assertion we refuse to make about competitors on our comparison pages.
Why can't I compare accuracy numbers between dictation apps?
Because a WER figure is meaningless without the audio and the scoring rules behind it. Different recordings, different speakers and accents, different microphones, and different normalization all move the number by more than the gap between most products. Our own testing found that changing the microphone alone shifted the result by roughly a percentage point. Unless two numbers come from the same audio scored the same way, comparing them tells you nothing.
What actually makes dictation accurate in daily use?
For most people it is not the model, it is three things around it. A decent microphone, because input quality moves the result more than most software choices. Vocabulary handling, because the words that matter most to you — names, jargon, project terms — are the ones a general model has never seen. And a cleanup step, because a transcript full of filler words and false starts is technically accurate and still not usable text.