VoiceInk Docs
Getting Started

Recommended Models

A practical starting point for choosing transcription and enhancement models in VoiceInk.

Choose transcription and enhancement models separately:

  • Transcription turns speech into text.
  • Enhancement cleans up, formats, or rewrites the transcript.

Local transcription

Use Parakeet first. It is the fastest local path in VoiceInk, has the best speed/accuracy balance, works offline, and supports realtime transcription.

ModelBest use
Parakeet V3Best first choice for most users. Fast local realtime transcription with multilingual support.
Parakeet V2Fast English-only fallback if V3 is not the right fit for your audio.

Cloud transcription

Use cloud transcription when you want better accuracy from a cloud transcription service and are okay with provider pricing after free credits.

ProviderRecommended modelWhy use it
AssemblyAIuniversal-3-5-pro / Universal-3.5 Pro realtimeBest first cloud option for realtime transcription. Sign up for AssemblyAI using this link to get $100 in free credits.
Deepgramnova-3 realtimeFast realtime cloud transcription. Deepgram currently advertises a $200 credit to start.
ElevenLabsscribe_v2 / Scribe V2 realtimeFast, accurate realtime transcription from ElevenLabs.

Start with Parakeet for privacy, offline use, and low latency.

Use AssemblyAI, Deepgram, or ElevenLabs when you want fast, low-latency cloud transcription. Choose realtime models for the best experience.

Enhancement runs after transcription. You usually do not need the best flagship AI model for cleanup, formatting, punctuation, vocabulary correction, or Mode prompts.

ProviderModels
Groqopenai/gpt-oss-120b
Cerebrasgpt-oss-120b, zai-glm-4.7 / GLM 4.7
Geminigemini-3.5-flash
OpenRouteropenai/gpt-oss-120b

Groq and Cerebras are good first choices because they are fast and have free tiers. Free-tier requests can be slower or rate-limited. For lower latency, add billing or use pay-as-you-go.

Use OpenRouter if you want one API key and one billing setup for many models. For enhancement, choose a fast and reasonably priced model unless your prompt needs deeper reasoning.

OpenAI, Anthropic, Mistral, Gemini, Ollama, Local CLI, and custom OpenAI-compatible providers can also work if you are already using one of them.

Keep enhancement fast

If enhancement regularly takes longer than two seconds, switch provider or model. Fast inference keeps dictation feeling natural.

Choosing the Better Setup

Start with Parakeet

Download Parakeet V3. If local transcription does not fit your needs, choose one of the suggested cloud transcription models.

Pick a fast enhancement model

Start with Groq openai/gpt-oss-120b, Cerebras gpt-oss-120b, or Gemini gemini-3.5-flash.

Watch the real timing

If dictation feels slow, compare actual timings in VoiceInk instead of guessing.

Compare Enhancement Models

To compare intelligence level and choose the best model for transcript enhancement, use Artificial Analysis.

Use VoiceInk Insights

For a single recording, open History, select one or more transcripts, and click Analyze in the bottom toolbar. For broader model performance, go to Dashboard and click View Insights.