Recommended Models
A practical starting point for choosing transcription and enhancement models in VoiceInk.
Choose transcription and enhancement models separately:
- Transcription turns speech into text.
- Enhancement cleans up, formats, or rewrites the transcript.
Recommended Models for Transcription
Local transcription
Use Parakeet first. It is the fastest local path in VoiceInk, has the best speed/accuracy balance, works offline, and supports realtime transcription.
| Model | Best use |
|---|---|
| Parakeet V3 | Best first choice for most users. Fast local realtime transcription with multilingual support. |
| Parakeet V2 | Fast English-only fallback if V3 is not the right fit for your audio. |
Cloud transcription
Use cloud transcription when you want better accuracy from a cloud transcription service and are okay with provider pricing after free credits.
| Provider | Recommended model | Why use it |
|---|---|---|
| AssemblyAI | universal-3-5-pro / Universal-3.5 Pro realtime | Best first cloud option for realtime transcription. Sign up for AssemblyAI using this link to get $100 in free credits. |
| Deepgram | nova-3 realtime | Fast realtime cloud transcription. Deepgram currently advertises a $200 credit to start. |
| ElevenLabs | scribe_v2 / Scribe V2 realtime | Fast, accurate realtime transcription from ElevenLabs. |
Start with Parakeet for privacy, offline use, and low latency.
Use AssemblyAI, Deepgram, or ElevenLabs when you want fast, low-latency cloud transcription. Choose realtime models for the best experience.
Recommended Models for Enhancement
Enhancement runs after transcription. You usually do not need the best flagship AI model for cleanup, formatting, punctuation, vocabulary correction, or Mode prompts.
| Provider | Models |
|---|---|
| Groq | openai/gpt-oss-120b |
| Cerebras | gpt-oss-120b, zai-glm-4.7 / GLM 4.7 |
| Gemini | gemini-3.5-flash |
| OpenRouter | openai/gpt-oss-120b |
Groq and Cerebras are good first choices because they are fast and have free tiers. Free-tier requests can be slower or rate-limited. For lower latency, add billing or use pay-as-you-go.
Use OpenRouter if you want one API key and one billing setup for many models. For enhancement, choose a fast and reasonably priced model unless your prompt needs deeper reasoning.
OpenAI, Anthropic, Mistral, Gemini, Ollama, Local CLI, and custom OpenAI-compatible providers can also work if you are already using one of them.
Keep enhancement fast
If enhancement regularly takes longer than two seconds, switch provider or model. Fast inference keeps dictation feeling natural.
Choosing the Better Setup
Start with Parakeet
Download Parakeet V3. If local transcription does not fit your needs, choose one of the suggested cloud transcription models.
Pick a fast enhancement model
Start with Groq openai/gpt-oss-120b, Cerebras gpt-oss-120b, or Gemini gemini-3.5-flash.
Watch the real timing
If dictation feels slow, compare actual timings in VoiceInk instead of guessing.
Compare Enhancement Models
To compare intelligence level and choose the best model for transcript enhancement, use Artificial Analysis.
Use VoiceInk Insights
For a single recording, open History, select one or more transcripts, and click Analyze in the bottom toolbar. For broader model performance, go to Dashboard and click View Insights.