Most dictation today is a round trip to someone else's GPU. That is fine until your voice is the part the model was never trained on — an accent, a speech difference, a bad microphone in a noisy room — and you have no way to make it adapt.
That is the problem Vokarys goes after. It is a desktop app that turns speech into text with recognition running locally on your computer, with punctuation and capitalisation added as you talk. It learns from your own voice over time, and recordings stay on the machine unless you explicitly choose to contribute them. There is a voice chat mode wired to Gemini as well, and a browser version if you would rather not install anything.
If you have given up on dictation because the generic model never quite got you, it is worth a look.