VoiceInk Docs
Getting Started

Recommended Models

A practical starting point for choosing transcription and enhancement models in VoiceInk.

Choose transcription and enhancement models separately:

  • Transcription turns speech into text.
  • Enhancement cleans up, formats, or rewrites the transcript.

Local transcription

Use Parakeet V3 or Parakeet V2 first. They are the fastest local options in VoiceInk, offer an excellent speed/accuracy balance, work offline, and support realtime transcription.

ModelBest use
Parakeet V3Best first choice for most users. Fast local realtime transcription with multilingual support.
Parakeet V2Fast local realtime transcription for English.
SenseVoice SmallBest local choice for CJK languages—Chinese, Japanese, and Korean—as well as English. It is fast, private, and lightweight.
Cohere TranscribeAccurate, private, multilingual local transcription. It is a little slower than Parakeet V2 and Parakeet V3.

Cloud transcription

Use cloud transcription when you want better accuracy from a cloud transcription service and are okay with provider pricing after free credits.

The table below includes every built-in cloud transcription model currently available in VoiceInk. Batch sends the completed recording for transcription; realtime returns partial text while you speak. An empty cell means that mode is not available.

Provider · modelBatch / dayRealtime / day
Groq · Whisper Large V3 Turbo$0.08
Deepgram · Nova 3$0.52–$0.62$0.58–$0.70
Deepgram · Nova 3 Medical$0.52$0.58
AssemblyAI · Universal-3.5 Pro$0.42$0.90
AssemblyAI · Universal-2$0.30
ElevenLabs · Scribe V2$0.44$0.78
Mistral · Voxtral Mini$0.36$0.72
Gemini · Gemini 3.5 Transcribe≈$0.60≈$1.08
Soniox · Soniox V5$0.20$0.24
Speechmatics · Enhanced$0.80$0.86
xAI · Grok STT$0.20$0.40
Cartesia · Ink 2≈$1.63

How the daily price is calculated

Daily estimates assume two hours of transcription per day at public pay-as-you-go list prices, before free credits, taxes, add-ons, or volume discounts. Deepgram ranges reflect monolingual and multilingual pricing. Gemini publishes estimated per-minute blended rates, so its totals are approximate. †Cartesia sells model credits through monthly plans; two hours every day is about 60 hours per month and requires its current $49/month Startup plan, shown here as an average over 30 days. Prices were checked on September 13, 2026 and can change.

For the best realtime experience, start with Deepgram Nova 3. AssemblyAI Universal-3.5 Pro and ElevenLabs Scribe V2 are strong alternatives. Deepgram currently advertises a $200 starting credit; sign up for AssemblyAI using this link for the currently offered referral credits.

Enhancement runs after transcription. You usually do not need the best flagship AI model for cleanup, formatting, punctuation, vocabulary correction, or Mode prompts.

ProviderModels
Groqqwen/qwen3.8-27b
Cerebrasqwen-3.8-27b
Geminigemini-3.8-flash, gemini-3.7-flash, or the latest Flash model available
OpenRouteropenai/gpt-oss-120b, a Qwen 3.8 27B variant, or the latest fast model available
OpenAIgpt-5.6-luna

For both Groq and Cerebras, start with Qwen 3.8 27B. It is the stronger overall enhancement model compared with GPT-OSS while remaining fast enough for everyday dictation. Both providers have free tiers, although free-tier requests can be slower or rate-limited. For lower latency, add billing or use pay-as-you-go.

Use OpenRouter if you want one API key and one billing setup for many models. For enhancement, choose a fast and reasonably priced model unless your prompt needs deeper reasoning.

OpenAI, Anthropic, Mistral, Gemini, Ollama, Local CLI, and custom OpenAI-compatible providers can also work if you are already using one of them.

Keep enhancement fast

If enhancement regularly takes longer than two seconds, switch provider or model. Fast inference keeps dictation feeling natural.

Choosing the Better Setup

Start with Parakeet

Download Parakeet V3 or Parakeet V2. For Chinese, Japanese, or Korean, choose SenseVoice Small. Choose Cohere Transcribe when multilingual accuracy matters more than getting Parakeet-level speed. If local transcription does not fit your needs, start with Deepgram.

Pick a fast enhancement model

Start with Groq qwen/qwen3.8-27b, Cerebras qwen-3.8-27b, or Gemini gemini-3.8-flash.

Watch the real timing

If dictation feels slow, compare actual timings in VoiceInk instead of guessing.

Compare Enhancement Models

To compare intelligence level and choose the best model for transcript enhancement, use Artificial Analysis.

Use VoiceInk Insights

For a single recording, open History, select one or more transcripts, and click Analyze in the bottom toolbar. For broader model performance, go to Dashboard and click View Insights.