The harness · Models

Any speech model, one app.

Hukum is a harness for speech models. It runs 11 offline models on your own computer, free, including NVIDIA Parakeet, OpenAI Whisper and Moonshine, and connects to 5 cloud models with your own API key: Sarvam Saaras, Deepgram Nova-3, Whisper on Groq, OpenAI Whisper and ElevenLabs Scribe. Pick the one that fits your language and your computer, and switch whenever a better one comes out.

Last updated 29 September 2026

Quick answer

Which model should you pick?

If you wantPick
English or a European language, on any computerParakeet V3. Fast, accurate, 25 European languages, and it runs on the CPU, so no graphics card or Apple silicon is needed.
The smallest download, or an older machineMoonshine (31–192 MB, English). Very fast, and it handles accents well.
A language Parakeet doesn't coverWhisper. Small is fastest, Large is most accurate; Turbo sits in between. A graphics card helps with the larger sizes.
Indian English, Hinglish or Indian languagesSaaras v3 from Sarvam, in the cloud with your own key. See the Indian-English test.
The cheapest cloud optionWhisper large-v3-turbo on Groq, about $0.04 per hour of audio on Groq's price list.
Russian, or Taiwanese MandarinGigaAM v3 for Russian; Breeze ASR for Taiwanese Mandarin with code-switching.

Offline · free · on your computer

11 models that never upload your voice.

ModelByLanguagesDownloadNotes
Parakeet V3NVIDIA25 European languages456 MBFast and accurate. Runs on the CPU, no graphics card needed. The recommended default.
Parakeet V2NVIDIAEnglish451 MBEnglish only. Fast on the CPU.
Whisper SmallOpenAIMany languages465 MBFast and fairly accurate.
Whisper MediumOpenAIMany languages469 MBGood accuracy, medium speed.
Whisper TurboOpenAIMany languages1.5 GBBalanced accuracy and speed.
Whisper LargeOpenAIMany languages1.0 GBThe most accurate Whisper. Slower.
MoonshineUseful SensorsEnglish31–192 MBVery fast and very small. Handles accents well.
CanaryNVIDIASeveral languagesVariesAvailable in 180M Flash and 1B v2 sizes.
SenseVoiceAlibabaSeveral languagesVariesCompact multilingual model.
GigaAM v3Open modelRussianVariesSpecialised for Russian.
Breeze ASROpen modelTaiwanese MandarinVariesOptimised for Taiwanese Mandarin, with code-switching.

Cloud · optional · your own key

5 cloud models, billed to you by the provider.

ModelProviderLanguagesNotes
Saaras v3SarvamIndian English, Hindi-English and Indian languagesKeeps Indian English and code-mixed words as you said them.
Nova-3DeepgramMany languagesDeepgram's speech model.
Whisper large-v3-turboGroqMany languagesWhisper, hosted by Groq.
WhisperOpenAIMany languagesOpenAI's speech-to-text.
ScribeElevenLabsMany languagesElevenLabs' speech-to-text.

Your audio, their server, nobody in between

With a cloud model, audio goes straight from your computer to the provider you picked, using your key. Hukum has no server in the middle. Offline models send nothing. Privacy in detail.

Optional AI clean-up

Off by default. When on, it sends the transcribed text (never audio) to a language model you choose to fix punctuation and remove filler words: OpenAI, Anthropic, Groq, Cerebras, OpenRouter, Z.AI, AWS Bedrock, Apple Intelligence, Sarvam.

Questions

People also ask.

What is the best speech-to-text model for dictation?

For most people writing in English or a European language, NVIDIA's Parakeet V3: it is fast and accurate and runs offline on an ordinary CPU. For many other languages, use Whisper; for Indian English and Hinglish, use Sarvam's Saaras v3.

Parakeet or Whisper?

Parakeet V3 is faster on a CPU and covers 25 European languages. Whisper covers many more languages and comes in four sizes, but the larger sizes are slower and run best with a graphics card. In Hukum you can switch between them at any time.

Do offline models cost anything?

No. Offline models download once (from 31 MB to about 1.5 GB) and then run on your computer for free, with no internet connection.

How does bring-your-own-key work?

You create an API key with the provider (for example Sarvam, Deepgram or Groq), paste it into Hukum, and pick that model. Your audio goes straight from your computer to that provider, which bills you directly. There is no Hukum server in between.