The harness · Models
Any speech model, one app.
Hukum is a harness for speech models. It runs 11 offline models on your own computer, free, including NVIDIA Parakeet, OpenAI Whisper and Moonshine, and connects to 5 cloud models with your own API key: Sarvam Saaras, Deepgram Nova-3, Whisper on Groq, OpenAI Whisper and ElevenLabs Scribe. Pick the one that fits your language and your computer, and switch whenever a better one comes out.
Last updated 29 September 2026
Quick answer
Which model should you pick?
| If you want | Pick |
|---|---|
| English or a European language, on any computer | Parakeet V3. Fast, accurate, 25 European languages, and it runs on the CPU, so no graphics card or Apple silicon is needed. |
| The smallest download, or an older machine | Moonshine (31–192 MB, English). Very fast, and it handles accents well. |
| A language Parakeet doesn't cover | Whisper. Small is fastest, Large is most accurate; Turbo sits in between. A graphics card helps with the larger sizes. |
| Indian English, Hinglish or Indian languages | Saaras v3 from Sarvam, in the cloud with your own key. See the Indian-English test. |
| The cheapest cloud option | Whisper large-v3-turbo on Groq, about $0.04 per hour of audio on Groq's price list. |
| Russian, or Taiwanese Mandarin | GigaAM v3 for Russian; Breeze ASR for Taiwanese Mandarin with code-switching. |
Offline · free · on your computer
11 models that never upload your voice.
| Model | By | Languages | Download | Notes |
|---|---|---|---|---|
| Parakeet V3 | NVIDIA | 25 European languages | 456 MB | Fast and accurate. Runs on the CPU, no graphics card needed. The recommended default. |
| Parakeet V2 | NVIDIA | English | 451 MB | English only. Fast on the CPU. |
| Whisper Small | OpenAI | Many languages | 465 MB | Fast and fairly accurate. |
| Whisper Medium | OpenAI | Many languages | 469 MB | Good accuracy, medium speed. |
| Whisper Turbo | OpenAI | Many languages | 1.5 GB | Balanced accuracy and speed. |
| Whisper Large | OpenAI | Many languages | 1.0 GB | The most accurate Whisper. Slower. |
| Moonshine | Useful Sensors | English | 31–192 MB | Very fast and very small. Handles accents well. |
| Canary | NVIDIA | Several languages | Varies | Available in 180M Flash and 1B v2 sizes. |
| SenseVoice | Alibaba | Several languages | Varies | Compact multilingual model. |
| GigaAM v3 | Open model | Russian | Varies | Specialised for Russian. |
| Breeze ASR | Open model | Taiwanese Mandarin | Varies | Optimised for Taiwanese Mandarin, with code-switching. |
Cloud · optional · your own key
5 cloud models, billed to you by the provider.
| Model | Provider | Languages | Notes |
|---|---|---|---|
| Saaras v3 | Sarvam | Indian English, Hindi-English and Indian languages | Keeps Indian English and code-mixed words as you said them. |
| Nova-3 | Deepgram | Many languages | Deepgram's speech model. |
| Whisper large-v3-turbo | Groq | Many languages | Whisper, hosted by Groq. |
| Whisper | OpenAI | Many languages | OpenAI's speech-to-text. |
| Scribe | ElevenLabs | Many languages | ElevenLabs' speech-to-text. |
Your audio, their server, nobody in between
With a cloud model, audio goes straight from your computer to the provider you picked, using your key. Hukum has no server in the middle. Offline models send nothing. Privacy in detail.
Optional AI clean-up
Off by default. When on, it sends the transcribed text (never audio) to a language model you choose to fix punctuation and remove filler words: OpenAI, Anthropic, Groq, Cerebras, OpenRouter, Z.AI, AWS Bedrock, Apple Intelligence, Sarvam.
Questions
People also ask.
What is the best speech-to-text model for dictation?
For most people writing in English or a European language, NVIDIA's Parakeet V3: it is fast and accurate and runs offline on an ordinary CPU. For many other languages, use Whisper; for Indian English and Hinglish, use Sarvam's Saaras v3.
Parakeet or Whisper?
Parakeet V3 is faster on a CPU and covers 25 European languages. Whisper covers many more languages and comes in four sizes, but the larger sizes are slower and run best with a graphics card. In Hukum you can switch between them at any time.
Do offline models cost anything?
No. Offline models download once (from 31 MB to about 1.5 GB) and then run on your computer for free, with no internet connection.
How does bring-your-own-key work?
You create an API key with the provider (for example Sarvam, Deepgram or Groq), paste it into Hukum, and pick that model. Your audio goes straight from your computer to that provider, which bills you directly. There is no Hukum server in between.