# Introduction to VoiceInk (/docs/introduction) ## Welcome to VoiceInk [#welcome-to-voiceink] VoiceInk is a powerful, native macOS application designed to transcribe your speech into text with exceptional speed and accuracy. It is built with privacy and efficiency at its core, so you can use local models on your Mac and choose cloud providers only when you want them. Whether you're a writer, developer, student, or anyone who wants to type faster and more naturally, VoiceInk offers a suite of features to streamline your workflow. ## Core Features [#core-features] VoiceInk is more than a transcription tool. These are the core pieces you will use most: * **Accurate, instant transcription**: Local and cloud models turn speech into text quickly, with strong accuracy for everyday dictation. * **Privacy first**: Local models keep voice processing on your Mac, while cloud providers are optional and configured by you. * **[Modes](/docs/modes)**: Apply different transcription, AI enhancement, context, output, and shortcut settings based on the app or website you are using. * **[Context-aware AI](/docs/context-awareness)**: Let VoiceInk use selected text, clipboard text, or visible screen text to improve AI-enhanced output. * **Global shortcuts**: Configure system-wide shortcuts for toggle recording, push-to-talk, hybrid recording, retry, cancel, and paste actions. * **Personal dictionary**: Teach VoiceInk names, technical terms, phrases, and deterministic word replacements you use often. * **AI enhancement**: Clean up dictated text, rewrite selected text, draft emails, or adapt output to a specific workflow. * **AI Assistant**: Ask questions, summarize text, or give commands with your voice. ## How to Use VoiceInk [#how-to-use-voiceink] 1. **Launch**: Open VoiceInk from your `Applications` folder. You will see a VoiceInk icon in your menu bar. 2. **Set up**: Grant the required permissions and complete onboarding. 3. **Record**: Use your configured keyboard shortcut to start and stop recording. 4. **Transcribe**: VoiceInk turns your speech into text using the selected configurations based on your selected mode. 5. **Insert**: The final text is pasted at your cursor, auto-sent, or shown as an assistant response depending on the Mode. These articles walk through the settings and workflows that help VoiceInk feel natural day to day. # Installation (/docs/installation) ## Install the App [#install-the-app] 1. Download VoiceInk from [tryvoiceink.com](https://tryvoiceink.com). 2. Open the downloaded app. 3. Drag VoiceInk to your Applications folder. 4. Launch VoiceInk. If macOS blocks the first launch, open **System Settings -> Privacy & Security**, approve VoiceInk, then launch it again. ## Required Permissions [#required-permissions] VoiceInk asks for these macOS permissions during setup: | Permission | Why VoiceInk needs it | | ---------------- | ---------------------------------------------------------------------------------------------------- | | Microphone | Records your voice for transcription. | | Accessibility | Sends paste and auto-send keystrokes, handles global shortcuts, and reads selected text for context. | | Screen Recording | Reads visible window text for Mode context when Screen context is enabled. | Screen Recording changes may require restarting VoiceInk before macOS reports the permission as granted. ## Complete Onboarding [#complete-onboarding] After the required permissions are granted, follow the onboarding prompts in VoiceInk and complete the setup process. Onboarding helps you choose an initial mode, configure the basics, and get ready to start dictating. You can manage modes and other settings later in the app. # Recommended Models (/docs/recommended-models) Choose transcription and enhancement models separately: * **Transcription** turns speech into text. * **Enhancement** cleans up, formats, or rewrites the transcript. ## Recommended Models for Transcription [#recommended-models-for-transcription] ### Local transcription [#local-transcription] Use **Parakeet** first. It is the fastest local path in VoiceInk, has the best speed/accuracy balance, works offline, and supports realtime transcription. | Model | Best use | | --------------- | ---------------------------------------------------------------------------------------------- | | **Parakeet V3** | Best first choice for most users. Fast local realtime transcription with multilingual support. | | **Parakeet V2** | Fast English-only fallback if V3 is not the right fit for your audio. | ### Cloud transcription [#cloud-transcription] Use cloud transcription when you want better accuracy from a cloud transcription service and are okay with provider pricing after free credits. | Provider | Recommended model | Why use it | | -------------- | ------------------------------------------------ | --------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | **AssemblyAI** | `universal-3-5-pro` / Universal-3.5 Pro realtime | Best first cloud option for realtime transcription. [Sign up for AssemblyAI using this link](https://assemblyai.cello.so/2H2stYqmF4X) to get **$100 in free credits**. | | **Deepgram** | `nova-3` realtime | Fast realtime cloud transcription. Deepgram currently advertises a **$200 credit** to start. | | **ElevenLabs** | `scribe_v2` / Scribe V2 realtime | Fast, accurate realtime transcription from ElevenLabs. | Start with Parakeet for privacy, offline use, and low latency. Use AssemblyAI, Deepgram, or ElevenLabs when you want fast, low-latency cloud transcription. Choose realtime models for the best experience. ## Recommended Models for Enhancement [#recommended-models-for-enhancement] Enhancement runs after transcription. You usually do **not** need the best flagship AI model for cleanup, formatting, punctuation, vocabulary correction, or Mode prompts. | Provider | Models | | -------------- | --------------------------------------- | | **Groq** | `openai/gpt-oss-120b` | | **Cerebras** | `gpt-oss-120b`, `zai-glm-4.7` / GLM 4.7 | | **Gemini** | `gemini-3.5-flash` | | **OpenRouter** | `openai/gpt-oss-120b` | Groq and Cerebras are good first choices because they are fast and have free tiers. Free-tier requests can be slower or rate-limited. For lower latency, add billing or use pay-as-you-go. Use OpenRouter if you want one API key and one billing setup for many models. For enhancement, choose a fast and reasonably priced model unless your prompt needs deeper reasoning. OpenAI, Anthropic, Mistral, Gemini, Ollama, Local CLI, and custom OpenAI-compatible providers can also work if you are already using one of them. If enhancement regularly takes longer than two seconds, switch provider or model. Fast inference keeps dictation feeling natural. ## Choosing the Better Setup [#choosing-the-better-setup] ### Start with Parakeet [#start-with-parakeet] Download **Parakeet V3**. If local transcription does not fit your needs, choose one of the suggested cloud transcription models. ### Pick a fast enhancement model [#pick-a-fast-enhancement-model] Start with **Groq** `openai/gpt-oss-120b`, **Cerebras** `gpt-oss-120b`, or **Gemini** `gemini-3.5-flash`. ### Watch the real timing [#watch-the-real-timing] If dictation feels slow, compare actual timings in VoiceInk instead of guessing. ## Compare Enhancement Models [#compare-enhancement-models] To compare intelligence level and choose the best model for transcript enhancement, use [Artificial Analysis](https://artificialanalysis.ai/). For a single recording, open **History**, select one or more transcripts, and click **Analyze** in the bottom toolbar. For broader model performance, go to **Dashboard** and click **View Insights**. # License Key Management (/docs/license-management) ## License Key Not Received After Purchase [#license-key-not-received-after-purchase] After purchasing, your license key is sent to the email used at checkout. **If you haven't received it:** 1. Check your spam or junk folder 2. If you used Apple's "Hide My Email" during checkout, the email may not have forwarded — check your iCloud inbox or use your real Apple ID email on the portal 3. Retrieve it yourself from the [Polar Portal](https://polar.sh/beingpax/portal/request) — enter your purchase email and you'll get a verification code to access your license Still not finding it? You may have used the wrong email address or made a typo during checkout. [Contact me](/contact) with your receipt and we'll resend it. ## Activating Your License [#activating-your-license] 1. Open VoiceInk 2. When prompted, paste your license key 3. Follow the on-screen instructions to complete activation ## License Activation Failed [#license-activation-failed] If you see "Invalid License" or "Trial Expired" after entering your key: * **Copied incorrectly** — make sure you copied the full key with no extra spaces * **Device limit reached** — each license covers a set number of Macs. If you've activated on another device, deactivate it first via the [Polar Portal](https://polar.sh/beingpax/portal/request) ## Still Having Issues? [#still-having-issues] [Contact me](/contact) with your purchase email and the specific error you're seeing. # VoiceInk Modes (/docs/modes) ## What a Mode Does [#what-a-mode-does] A Mode tells VoiceInk how to behave for a dictation session. It can control transcription, [AI enhancement](/docs/mode-settings#ai-enhancement), [context](/docs/context-awareness), [output behavior](/docs/mode-settings#output-modes), [triggers](/docs/mode-triggers), and the shortcut that starts recording directly with that Mode. Create different Modes when different tasks need different behavior. Examples: * A simple dictation Mode for fast everyday notes. * An email Mode that turns rough speech into polished email text. * A rewrite Mode that uses selected text as context. * An assistant Mode that answers inside the recorder instead of pasting. ## When to Create a Mode [#when-to-create-a-mode] Create a Mode when a workflow needs its own behavior. That might be a different prompt, a different transcription model, a different language, or a different output style. You can switch Modes manually, assign a Mode shortcut, or let VoiceInk choose a Mode automatically with [triggers](/docs/mode-triggers). ## Starter Modes [#starter-modes] VoiceInk can create starter Modes during setup: | Mode | Default behavior | | ----------- | ----------------------------------------------------------------------------------------------------------- | | Dictation | Fast transcription with no AI enhancement. Pastes the transcript and is the default Mode. | | Enhancement | Cleans up dictated text with AI, using selected text and screen context when available. | | Email | Drafts email text with AI. It can use selected text, clipboard text, and screen context. | | Rewrite | Rewrites selected or dictated text with AI. It uses selected text as the main context. | | Assistant | Answers in the recorder with [Respond output](/docs/assistant-mode) instead of pasting into the active app. | ## What a Mode Can Control [#what-a-mode-can-control] | Area | What It Controls | | -------------- | ---------------------------------------------------------------------------------------------------------------------------- | | Triggers | Apps, websites, spoken word triggers, suggested trigger groups, and keyboard shortcut. | | Transcription | Model, realtime behavior, language, and formatting. | | AI Enhancement | Provider, model, prompt, and whether enhancement is enabled. | | Context | Selected text, clipboard text, and screen text used by AI enhancement. | | Output | Decide whether the final text is pasted, shown as a recorder response, or sent to a [custom command](/docs/custom-commands). | | Auto Send | Press Return, Shift+Return, or Command+Return after pasted text. | | Default | Use this Mode when no app or website trigger matches. | | Enabled | Turn a Mode on or off without deleting it. | ## Related [#related] * [Mode Triggers](/docs/mode-triggers) * [Mode Settings](/docs/mode-settings) * [Custom Commands](/docs/custom-commands) * [Context Awareness](/docs/context-awareness) * [Assistant Mode](/docs/assistant-mode) # Mode Triggers (/docs/mode-triggers) ## Trigger Types [#trigger-types] VoiceInk supports four ways to choose a Mode: | Trigger type | What it matches | | ----------------- | ------------------------------------------------------------------------ | | App trigger | A specific macOS app bundle, such as Mail, Slack, or Cursor. | | Website trigger | A cleaned URL or domain, such as `mail.google.com` or `docs.google.com`. | | Word trigger | A phrase you say at the beginning or end of a dictation. | | Keyboard shortcut | A shortcut that starts recording directly with that Mode. | App and website triggers are automatic. VoiceInk checks the active app when recording starts. If the active app is a supported browser, VoiceInk can also read the current URL and use the website-specific [Mode](/docs/modes). Word triggers are spoken mode switches. After transcription, VoiceInk checks whether your dictated text starts or ends with a configured trigger phrase. If it matches, VoiceInk switches to that Mode, removes the trigger phrase, and processes the remaining text with the triggered Mode. Keyboard shortcuts are direct manual triggers. They skip app and website matching and start recording with the shortcut's Mode immediately. ## Matching Order [#matching-order] When you start recording normally, VoiceInk first uses the active app or the default Mode. If the active app is a browser and the current URL matches a website trigger, the website Mode can take over. After the transcript is ready, a word trigger can still switch the Mode for that dictation. This is useful when the app does not tell VoiceInk enough about the task, or when one app needs several writing styles. ## Suggested Groups [#suggested-groups] The trigger picker includes suggested groups for common workflows: | Group | What it includes | | --------- | ----------------------------------------------- | | AI | AI apps, coding apps, and AI websites. | | Email | Email apps and webmail sites. | | Messaging | Chat and messaging apps or websites. | | Writing | Notes, documents, and writing apps or websites. | You can add a group as a starting point and then remove individual apps or websites you do not want. ## Add Triggers [#add-triggers] 1. Open **Modes**. 2. Edit a Mode. 3. In **Triggers**, click the plus button. 4. Choose a suggested group, search installed apps, enter a website, or enter a trigger word. For websites, use the meaningful part of the address, such as `mail.google.com`, `docs.google.com`, or `notion.so`. VoiceInk cleans the URL before saving it. For word triggers, use a short phrase you would naturally say, such as `email mode`, `rewrite this`, or `assistant`. Each word trigger can belong to only one Mode. ## Duplicate Protection [#duplicate-protection] VoiceInk prevents duplicate Mode names, duplicate app triggers, and duplicate website triggers when saving. The trigger picker also prevents adding a trigger word that already belongs to another Mode. ## Default Mode [#default-mode] Use **Set as default** in a Mode's [Advanced settings](/docs/mode-settings#advanced) when you want it to apply anywhere no app or website trigger matches. [Respond](/docs/mode-settings#output-modes) Modes cannot be set as the default. # Mode Settings (/docs/mode-settings) ## Triggers [#triggers] The top of the Mode editor controls when the Mode activates. For matching behavior, see [Mode Triggers](/docs/mode-triggers). | Setting | What it does | | ----------------- | ----------------------------------------------------------------------- | | Triggers | Adds apps, websites, suggested trigger groups, or spoken word triggers. | | Keyboard Shortcut | Starts recording directly with this Mode. | ## Transcription [#transcription] For transcription, you can manage all the available local and cloud models from the [AI Models catalog](/docs/ai-models). | Setting | What it does | | ---------------- | ----------------------------------------------------------------------------- | | Model | Chooses the transcription model for this Mode. | | Real-time | Shows partial transcript while recording when the model supports streaming. | | Language | Chooses a supported language. Some models use **Autodetected** automatically. | | Paragraph breaks | Applies intelligent paragraph formatting to longer transcripts. | ## AI Enhancement [#ai-enhancement] When **AI Enhancement** is enabled, the Mode can clean up text, rewrite selected text, draft emails, or answer questions. The result depends on the selected [prompt](/docs/prompt-management). | Setting | What it does | | ----------------- | -------------------------------------------------------------------------------------------------------------------------------- | | AI Provider | Chooses a connected enhancement provider. | | AI Model | Chooses the provider model. Local CLI uses its default model setting. | | Prompt | Chooses a prompt. You can edit the selected prompt or add a new one from the Mode editor. | | Context Awareness | Lets the prompt use selected text, clipboard text, and/or visible screen text. See [Context Awareness](/docs/context-awareness). | Turning AI Enhancement on automatically selects an available provider, model, and prompt when the Mode does not already have them. ## Advanced [#advanced] | Setting | What it does | | -------------- | --------------------------------------------------------------------------------------------- | | Output | Chooses where the final text goes after transcription and optional AI enhancement. | | Set as default | Uses this Mode when no app or website trigger matches. Respond Modes cannot be default. | | Auto Send | Presses Return, Shift+Return, or Command+Return after pasted text. | | Custom Command | Runs a local shell command with the final text. See [Custom Commands](/docs/custom-commands). | ## Output Modes [#output-modes] Output controls delivery. The Mode still uses its transcription, formatting, prompt, and context settings first. Then VoiceInk delivers the final text using the selected output mode. | Mode | Best for | What happens | | -------------- | ------------------------------------------- | ------------------------------------------------------------------------------------------------------------------------------ | | Paste | Normal dictation, chat, email, notes, forms | VoiceInk pastes the final text into the active app. Auto Send can press Return, Shift+Return, or Command+Return after pasting. | | Respond | Assistant-style questions and follow-ups | VoiceInk keeps the AI response inside the recorder instead of pasting it. See [Assistant Mode](/docs/assistant-mode). | | Custom Command | Local workflows and automations | VoiceInk runs your shell command and gives it the final text. See [Custom Commands](/docs/custom-commands). | **Respond** is available only when AI Enhancement has a connected provider and selected prompt. A Respond Mode cannot be the default Mode. See [Assistant Mode](/docs/assistant-mode) for the full workflow. **Custom Command** runs locally with your user permissions. VoiceInk sends the final text on stdin and also exposes it as `VOICEINK_TRANSCRIPT`. This is useful for workflows like copying a modified transcript, appending to a journal, opening a search, or handing text to another script. See [Custom Commands](/docs/custom-commands) for examples and technical details. # Custom Commands (/docs/custom-commands) ## What Custom Commands Do [#what-custom-commands-do] Custom Command is an [output mode](/docs/mode-settings#output-modes). It controls what happens after VoiceInk has the final text. Before the command runs, the Mode can still use transcription settings, paragraph formatting, word replacements, [AI Enhancement](/docs/mode-settings#ai-enhancement), [prompts](/docs/prompt-management), and [context](/docs/context-awareness). Custom Command receives the final text that VoiceInk is ready to deliver. Use it when pasting is not enough and you want VoiceInk to hand the text to your own local workflow. ## What You Can Build [#what-you-can-build] Custom Commands can do anything a local non-interactive shell command can do within your macOS user permissions and the command timeout. Common uses: | Use case | What the command does | | --------------------- | ----------------------------------------------------------------------------------------------------- | | Custom paste behavior | Transform the final text, copy it to the clipboard, paste it, press another key, or move focus. | | Notes and journals | Append dictation to a Markdown file, daily note, task list, log, or project file. | | Search and lookup | Open a web search, documentation search, map, issue tracker, or internal URL using the dictated text. | | Local scripts | Pass the text to your own shell, Python, Node, Ruby, or other local script. | | App automation | Use tools such as `open`, `osascript`, `pbcopy`, or installed CLIs to control apps. | | Webhooks and APIs | Use tools such as `curl` to send the text to a service you choose. | If your command sends text to a web service, that service receives the text. VoiceInk only runs the command you configured. ## What VoiceInk Sends [#what-voiceink-sends] VoiceInk sends the same final text in two ways: | Input | Use it when | | ---------------------- | ----------------------------------------------------------------- | | Standard input | The next command reads text from stdin. | | `$VOICEINK_TRANSCRIPT` | You want the text inside a shell variable, argument, or pipeline. | Examples: ```sh cat > "$HOME/Desktop/voiceink.txt" ``` ```sh printf "%s" "$VOICEINK_TRANSCRIPT" | pbcopy ``` The text is UTF-8. Quote `$VOICEINK_TRANSCRIPT` so spaces, punctuation, and line breaks are preserved. ## What Final Text Means [#what-final-text-means] Custom Command receives the Mode's final deliverable text: * If AI Enhancement is off, it receives the transcript after VoiceInk formatting and word replacements. * If AI Enhancement runs successfully, it receives the enhanced text. * If AI Enhancement is skipped or fails, it receives the non-enhanced final transcript. Custom Command does not receive the audio file, raw model tokens, prompt messages, selected text, clipboard text, or screen text directly. If you need context to affect the command output, use a [prompt](/docs/prompt-management) and [Context Awareness](/docs/context-awareness) so VoiceInk includes that context while producing the final text. ## How It Runs [#how-it-runs] VoiceInk runs your command locally. | Detail | Behavior | | --------------------- | ---------------------------------------------------------------------- | | Shell | `/bin/zsh -lc` | | Timeout | 10 seconds | | Permissions | Same macOS user permissions as VoiceInk. | | Recorder | The recorder is dismissed before the command starts. | | Standard input | Final text is written to stdin, then stdin is closed. | | Environment | VoiceInk passes `$VOICEINK_TRANSCRIPT` plus common shell variables. | | Exit code | `0` means success. A non-zero exit code is treated as command failure. | | `stdout` and `stderr` | Captured for diagnostics. They are not pasted into the active app. | VoiceInk tries to use your shell `PATH`. If a command is not found, use an absolute path such as `/opt/homebrew/bin/tool` or `/Users/you/bin/script`. ## Available Variables [#available-variables] VoiceInk passes these environment values when available: | Variable | Meaning | | ------------------------------- | ------------------------------------------------------------- | | `$VOICEINK_TRANSCRIPT` | The final text from the Mode. | | `$HOME` | Your home folder. | | `$USER` | Your macOS user name. | | `$LOGNAME` | Your login name. | | `$SHELL` | Your shell path, defaulting to `/bin/zsh`. | | `$TMPDIR` | A temporary directory. | | `$PATH` | Command search path discovered from your shell when possible. | | `$LANG`, `$LC_ALL`, `$LC_CTYPE` | Locale values when available. | ## Important Limits [#important-limits] * Custom Command is an output mode. It decides delivery; it does not change the saved transcript by returning text on stdout. * The command runs only after VoiceInk has a completed transcription result to deliver. * Auto Send applies only to **Paste** output, not Custom Command output. * The command must not require interactive input. * The command must finish within 10 seconds. If it times out, VoiceInk stops the command process. For longer work, start your own background process and exit quickly. * If you want the result pasted, your command must copy or paste it itself. * If the command is empty, the Mode cannot be saved. ## Check Whether a Command Worked [#check-whether-a-command-worked] VoiceInk does not show a success or failure message after running a Custom Command. To see Custom Command logs from the past 10 minutes, open Terminal and run: ```sh /usr/bin/log show --last 10m --style compact \ --predicate 'subsystem == "com.prakashjoshipax.voiceink" AND category == "CustomCommandDeliveryRunner"' ``` The log will show whether the command succeeded, failed, or exceeded the 10-second limit. If the command prints additional information, VoiceInk includes it in these logs. Logs may therefore contain dictated or command-generated text. ## Built-In Templates [#built-in-templates] The command menu includes starter templates: | Template | What it does | | -------------------------------- | --------------------------------------------------------- | | Paste and Press Tab | Copies the final text, pastes it, then presses Tab. | | Lowercase and Paste | Lowercases the final text, copies it, then pastes it. | | Remove Trailing Period and Paste | Removes one final period before pasting. | | Append to Journal | Adds the final text to `~/Documents/VoiceInk/journal.md`. | | Search Web | Opens a Google search for the final text. | Templates are examples. You can replace them with any command that follows the rules above. ## Examples [#examples] Copy the final text: ```sh printf "%s" "$VOICEINK_TRANSCRIPT" | pbcopy ``` Lowercase, copy, and paste: ```sh printf "%s" "$VOICEINK_TRANSCRIPT" | tr '[:upper:]' '[:lower:]' | pbcopy osascript <<'APPLESCRIPT' tell application "System Events" keystroke "v" using command down end tell APPLESCRIPT ``` Append to a note: ```sh mkdir -p "$HOME/Documents/VoiceInk" printf -- "- %s\n" "$VOICEINK_TRANSCRIPT" >> "$HOME/Documents/VoiceInk/notes.md" ``` Open a search: ```sh query=$(printf "%s" "$VOICEINK_TRANSCRIPT" | LC_ALL=C od -An -tx1 -v | tr -d ' \n' | sed 's/../%&/g') open "https://www.google.com/search?q=$query" ``` Pass text to your own script as an argument: ```sh /Users/you/bin/handle-voiceink "$VOICEINK_TRANSCRIPT" ``` Pass text to your own script through stdin: ```sh /Users/you/bin/handle-voiceink-from-stdin ``` Send text to an API you control: ```sh printf "%s" "$VOICEINK_TRANSCRIPT" | curl -sS -X POST "https://example.com/voiceink" \ -H "Content-Type: text/plain; charset=utf-8" \ --data-binary @- ``` ## Good Practices [#good-practices] * Start with a built-in template, then edit it. * Use absolute paths for your own scripts and uncommon command-line tools. * Quote variables. * Keep commands fast and non-interactive. * Test with harmless text before using it in an important Mode. * Remember that your command controls the side effect. VoiceInk only supplies the final text and runs the command. ## Related [#related] * [Mode Settings](/docs/mode-settings) * [Mode Triggers](/docs/mode-triggers) * [Prompt Management](/docs/prompt-management) # Prompt Management (/docs/prompt-management) ## What a Prompt Does [#what-a-prompt-does] A prompt tells [AI Enhancement](/docs/mode-settings#ai-enhancement) what to do with the dictated text and any enabled [context](/docs/context-awareness). A [Mode](/docs/modes) chooses which prompt to use. VoiceInk includes prompt templates for Default, Chat, Email, Rewrite, and Assistant. You can use those as-is, edit them, or create your own [custom prompt](/docs/custom-prompts). ## Use System Template [#use-system-template] When **Use System Template** is enabled, VoiceInk wraps your prompt with shared dictation cleanup instructions. This is the default for new prompts and works well when the prompt is mainly polishing dictated speech. See [Creating a Custom Prompt](/docs/custom-prompts) for the input tags and prompt examples. With the system template enabled, your prompt can stay focused on the Mode-specific behavior: chat style, email style, support reply, meeting notes, or another writing shape. Disable **Use System Template** only when you want full control over the AI's behavior, such as: * Rewriting selected text according to custom instructions. * Answering questions in Assistant Mode. * Running a workflow that should not behave like dictation cleanup. When disabled, your prompt text is sent directly to the AI without the system template wrapping. ## Creating a Prompt [#creating-a-prompt] 1. Open the **Modes** editor. 2. In the AI Enhancement section, choose **Prompt**. 3. Click **Add New** or edit an existing prompt. 4. Give it a name and write your instructions. 5. Choose whether to use the system template. 6. Save the prompt. ## Related [#related] * [Creating a Custom Prompt](/docs/custom-prompts) * [Mode Settings](/docs/mode-settings) * [Assistant Mode](/docs/assistant-mode) # Creating a Custom Prompt (/docs/custom-prompts) ## How VoiceInk Sends Text to AI [#how-voiceink-sends-text-to-ai] When [AI Enhancement](/docs/mode-settings#ai-enhancement) runs, VoiceInk sends your dictated text and any enabled [context](/docs/context-awareness) as tagged text. | Tag | What it contains | | --------------------------- | --------------------------------------------------------------------------------------- | | `` | The raw dictated speech from the recording. This is usually the main text to transform. | | `` | Selected text from the active app, if Selected Text context is enabled and available. | | `` | Clipboard text, if Clipboard context is enabled and available. | | `` | Text extracted from the active window, if Screen context is enabled and available. | | `` | Vocabulary terms, names, acronyms, or technical words saved in VoiceInk. | ## Use System Template [#use-system-template] For most dictation cleanup prompts, keep **Use System Template** enabled. VoiceInk wraps your prompt with shared instructions. With the system template enabled, your custom prompt can work well by describing only the task-specific behavior: ```text Polish the dictated speech in into a concise Slack message. # Rules - Keep it friendly and direct. - Use short lines when helpful. - Do not add a greeting or sign-off unless one was dictated. ``` Disable **Use System Template** only when you want full control over the AI's behavior, such as Assistant-style answering, complex rewriting, or a prompt that should not behave like dictation cleanup. ## Referencing Context [#referencing-context] Use [context](/docs/context-awareness) only when it helps. The prompt should make it clear whether the AI should transform the dictated speech, selected text, or both. For rewriting selected text: ```text # Goal Rewrite the text in using the instruction in . # Rules - Preserve the meaning and important details. - Follow any tone, length, or format requested in . - Use or only to clarify references. ``` For replying to an email thread: ```text Polish the dictated speech in into a reply email. # Rules - Use , , or only to understand the thread and recipient. - Preserve the user's asks, decisions, dates, numbers, and constraints. - Do not invent a subject line, deadline, promise, or missing detail. ``` ## Writing a Good Prompt [#writing-a-good-prompt] Keep prompts short and specific. VoiceInk already supplies the general dictation rules when the system template is enabled, so your prompt should focus on what makes this mode different. Good prompts usually include: * The output style you want, such as email, chat, notes, summary, or support reply. * Any tone or audience, such as friendly, concise, professional, casual, or technical. * How context should be used. * What not to invent. * The expected output shape if it is special. ## Related [#related] * [Prompt Management](/docs/prompt-management) * [Context Awareness](/docs/context-awareness) * [Mode Settings](/docs/mode-settings) # Context Awareness (/docs/context-awareness) ## What Context Awareness Does [#what-context-awareness-does] Context Awareness lets an AI-enhanced [Mode](/docs/modes) use surrounding information when shaping the final output. The Mode can include: | Context source | How VoiceInk gets it | | -------------- | -------------------------------------------------------------- | | Selected Text | Reads the current selection through macOS Accessibility. | | Clipboard | Reads text currently on the clipboard at recording start. | | Screen | Captures the active window and extracts visible text with OCR. | Context is captured around recording time and sent as text to the enhancement provider only when the Mode has that source enabled. Context is used only by [AI Enhancement](/docs/mode-settings#ai-enhancement). Plain Dictation Modes do not send selected text, clipboard text, or screen text to an AI provider. ## Enable Context for a Mode [#enable-context-for-a-mode] 1. Open **Modes**. 2. Edit a Mode with **AI Enhancement** enabled. 3. Expand **Context Awareness**. 4. Turn on **Selected Text**, **Clipboard**, or **Screen**. 5. Save the Mode. ## Permissions [#permissions] | Context source | Required permission | | -------------- | ------------------- | | Selected Text | Accessibility | | Clipboard | No extra permission | | Screen | Screen Recording | If Accessibility is missing, selected text is skipped. If Screen Recording is missing, screen context is skipped. ## Screen Context [#screen-context] For Screen context, VoiceInk captures the active window, runs OCR locally with Apple's Vision framework, and sends extracted text to the AI provider as part of the prompt context. The screenshot image itself is not the context payload. ## Prompt Tags [#prompt-tags] When you enable a context source for a Mode and that context is available, your [custom prompts](/docs/custom-prompts) can refer to it by tag: | Tag | Context | | --------------------------- | -------------------------------------- | | `` | Selected text. | | `` | Clipboard text. | | `` | Text extracted from the active window. | For the full input format, see [Creating a Custom Prompt](/docs/custom-prompts). ## Good Uses [#good-uses] * Rewrite selected text by speaking an instruction. * Draft a reply using a selected email or message thread. * Ask the Assistant about visible content in the active window. * Preserve names, tasks, or dates from nearby text. # Assistant Mode (/docs/assistant-mode) ## What Assistant Mode Is [#what-assistant-mode-is] Assistant Mode is a [Mode](/docs/modes) configured with [AI Enhancement](/docs/mode-settings#ai-enhancement) and [Output -> Respond](/docs/mode-settings#output-modes). Instead of pasting text into the active app, VoiceInk shows the AI response inside the recorder. The starter Assistant Mode uses the Assistant prompt and is designed for short questions, summaries, and follow-ups. Respond output is available only when the Mode has AI Enhancement enabled, a selected prompt, and a connected AI provider. Respond Modes cannot be set as the default Mode. ## Follow-ups [#follow-ups] When an assistant response is visible and the recorder is idle, starting another recording continues the same assistant session. You can also type a follow-up in the recorder's assistant panel. Closing the recorder resets the assistant session. ## History [#history] Assistant turns are saved in History. The detail panel can show the transcription model, enhancement model, prompt, Mode, timing, and AI request data when available. ## Difference From Enhancement [#difference-from-enhancement] Enhancement rewrites your dictated text and usually pastes it. Assistant Mode treats your speech as a question or instruction and returns an answer in the recorder. # Transcription History (/docs/transcription-history) ## Open History [#open-history] Open **History** from the sidebar. VoiceInk shows saved transcriptions newest-first, loading 20 at a time. Click the **History Settings** button beside the search field to manage transcript and audio retention. ## Automatic Cleanup [#automatic-cleanup] History Settings has two mutually exclusive cleanup modes: | Setting | Options | What it deletes | | ------------------------------ | --------------------------------------------- | ----------------------------------------------------------- | | Auto-delete Transcript History | Immediately, 1 hour, 1 day, 3 days, or 7 days | Transcript records and their associated saved audio files. | | Auto-delete Audio Files | 1 day, 3 days, 7 days, 14 days, or 30 days | Saved audio files only. Transcript text remains in History. | When **Auto-delete Transcript History** is enabled, the separate audio cleanup option is hidden because deleting a transcript also deletes its associated audio file. The retention period is based on each transcript's creation time. VoiceInk checks after transcription completes and when the app launches. Use **Run Cleanup Now** to immediately delete items older than the selected retention period. It does not delete newer items. For example, with **1 day** selected, transcripts less than 24 hours old remain. VoiceInk asks for confirmation before enabling transcript auto-delete and before running cleanup manually. Changing the retention period updates the policy without immediately running cleanup. ### Local storage and deletion [#local-storage-and-deletion] VoiceInk stores History locally in its app-support database. Local processing means transcription can happen without sending audio to a cloud provider; it does not mean no local history is saved. Cloud models still send data to the provider you selected. Deleting History removes those records from the app and deletes associated saved audio. The underlying database uses SQLite, whose temporary write-ahead log and unused database pages can retain old bytes until SQLite reuses or compacts them. Cleanup should therefore be understood as application-level deletion, not guaranteed forensic erasure of every previous byte on disk. ## Search [#search] The search field matches both original transcription text and enhanced text. Search results can be selected, exported, analyzed, or deleted. ## Expand a Transcription [#expand-a-transcription] Click a row to expand it. Expanded rows show: * Original text. * Enhanced text, if available. * Copy button. * Audio playback when the saved audio file exists. * Info button for metadata. If a transcription has both original and enhanced text, use the tabs to switch between them. ## Audio Playback [#audio-playback] History includes a waveform player for saved audio. The player can: * Play or pause audio. * Seek through the waveform. * Show hover time. * Cycle playback speed between 1x, 1.5x, and 2x. If audio cleanup deleted the file, playback and retry are unavailable for that item. ## Info Panel [#info-panel] The Info panel can show: * Date * Audio duration * Transcription model * Transcription time * Enhancement model * Enhancement time * Prompt * Mode * AI system prompt and user message, when enhancement request data was saved The AI request section includes a copy button for the full request text. ## Bulk Actions [#bulk-actions] Select one or more rows to reveal the selection bar. | Action | What it does | | ---------- | ------------------------------------------------------------------- | | Analyze | Opens performance analysis for the selected transcriptions. | | Export | Exports selected rows to CSV. | | Delete | Deletes selected transcriptions and their saved audio files. | | Select All | Selects all rows matching the current search, not just loaded rows. | ## Performance Analysis [#performance-analysis] The Analyze panel summarizes: * Total selected transcripts. * Count with transcription timing data. * Total audio duration. * Count with enhancement timing data. * Transcription model performance. * Enhancement model performance. * Mac model, processor, and memory. Use this when comparing model speed or diagnosing slow output. # Transcribe Audio Files (/docs/transcribe-audio-files) ## Overview [#overview] Use Transcribe Audio Files when you want VoiceInk to process an existing audio or video file instead of dictating in the moment. Files use the same Mode-based transcription, formatting, word replacement, and enhancement pipeline, but they are processed from a queue. ## Add Files [#add-files] You can: * Drag audio or video files into the view. * Click **Choose Files** when the queue is empty. * Click **Add** when the queue already contains files. Supported extensions: `wav`, `mp3`, `m4a`, `aiff`, `mp4`, `mov`, `aac`, `flac`, `caf`, `amr`, `ogg`, `oga`, `opus`, `3gp` ## Choose a Mode [#choose-a-mode] Use the Mode picker in the top bar before starting. The selected Mode controls: * Transcription model * Language * Realtime availability where applicable * Formatting * Word Replacements * AI enhancement provider, model, prompt, and context settings If no enabled Mode exists, processing cannot start. ## Process the Queue [#process-the-queue] Click **Start** to process pending files. VoiceInk runs them sequentially. Use **Cancel** to stop queue processing. In-progress items return to pending. ## Completed Results [#completed-results] Completed rows can be expanded. If enhancement ran, use the **Original** and **Enhanced** tabs. You can copy or save the visible text. Each completed file is also saved in [History](/docs/transcription-history) with the audio file, model metadata, prompt name, timing data, and Mode details. ## Clear the Queue [#clear-the-queue] Click **Clear** to cancel work and remove all queue items from the Transcribe view. Completed transcriptions already saved in History remain there. # AI Models Catalog (/docs/ai-models) ## What the Catalog Is For [#what-the-catalog-is-for] The AI Models catalog is where you make models available to VoiceInk. After a model or provider is available, choose it inside a [Mode](/docs/modes). This keeps model setup separate from writing behavior: the catalog manages what VoiceInk can use, and Modes decide when to use it. ## Model Types [#model-types] | Type | Use It For | | --------------- | ---------------------------------------------------------------------------------- | | Local models | Offline transcription and local enhancement providers such as Ollama or Local CLI. | | Cloud providers | Hosted transcription and enhancement models connected with your own API key. | | Custom models | OpenAI-compatible transcription or chat-completion endpoints. | ## Local Models [#local-models] Local models run on your Mac. They are the best first choice when privacy, speed, and offline use matter. VoiceInk supports Parakeet, Whisper, Apple Speech, imported Whisper `.bin` files, Ollama, and Local CLI providers. See [Local Models](/docs/local-models). ## Cloud Providers [#cloud-providers] Cloud providers are useful when you want a hosted transcription model, a hosted enhancement model, or better performance on a Mac that is not a good fit for local processing. Provider keys are stored securely, and the provider's models become available inside Modes after connection. See [Cloud Providers](/docs/cloud-providers). ## Custom Models [#custom-models] Custom models let you connect your own OpenAI-compatible transcription or enhancement endpoint. Use this when you have a private server, local gateway, or provider endpoint that is not listed directly in VoiceInk. See [Custom Models](/docs/custom-models). # Local Models (/docs/local-models) ## Parakeet Models [#parakeet-models] VoiceInk includes NVIDIA Parakeet models through the Local tab: | Model | Notes | | ----------- | ------------------------------------------- | | Parakeet V3 | Realtime, multilingual, used by onboarding. | | Parakeet V2 | Realtime, English-only. | Click **Download** on the model card. Once downloaded, the model is available in Mode transcription settings. Downloaded Parakeet models can be deleted or revealed in Finder from the model card menu. ## Apple Speech [#apple-speech] Apple Speech uses the native macOS Speech framework. It is: * Built in. * On-device. * Multilingual. * Available only on macOS 26 or later. For Apple Speech language selection, VoiceInk can show whether a language asset is downloaded. If macOS allows it, VoiceInk can start downloading the selected language asset. Apple Speech can reserve a limited number of languages, so you may need to remove one before adding another. ## Whisper Models [#whisper-models] VoiceInk includes these predefined Whisper models: * Tiny * Tiny (English) * Base * Base (English) * Large v2 * Large v3 * Large v3 Turbo * Large v3 Turbo (Quantized) Download a Whisper model from its card. If you no longer need it, delete it from the card menu. ## Import Local Whisper Models [#import-local-whisper-models] Use **Import Local Model...** to add a custom whisper.cpp-compatible `.bin` file. Requirements: * The file must have a `.bin` extension. * VoiceInk treats imported models as Whisper models. * Imported models are multilingual by default in VoiceInk. Steps: 1. Open **AI Models -> Local**. 2. Click **Import Local Model...**. 3. Select your `.bin` file. 4. Open **Modes** and choose the imported model in a Mode's **Transcription -> Model** picker. ## Local Enhancement Providers [#local-enhancement-providers] The bottom of the Local tab also includes: | Provider | What it does | | --------- | ---------------------------------------------------------------------------------------------------- | | Ollama | Connects to a local Ollama server, defaulting to `http://localhost:11434`. | | Local CLI | Sends enhancement prompts to a command template such as Claude, Codex, a script, or any CLI command. | Local CLI exposes `VOICEINK_SYSTEM_PROMPT`, `VOICEINK_USER_PROMPT`, and `VOICEINK_FULL_PROMPT` environment variables. It also writes the full prompt to stdin. # Cloud Providers (/docs/cloud-providers) ## Provider List [#provider-list] Open **AI Models -> Cloud** to connect providers. The current cloud provider list includes: | Provider | Transcription | Enhancement | | ------------ | ------------: | ----------: | | Groq | Yes | Yes | | Cerebras | No | Yes | | Gemini | Yes | Yes | | OpenAI | No | Yes | | OpenRouter | No | Yes | | Anthropic | No | Yes | | Mistral | Yes | Yes | | Deepgram | Yes | No | | ElevenLabs | Yes | No | | Soniox | Yes | No | | Speechmatics | Yes | No | | AssemblyAI | Yes | No | | xAI | Yes | No | | Cartesia | Yes | No | Providers with both capabilities can be used for transcription and enhancement in different Mode settings. ## Connect a Provider [#connect-a-provider] 1. Open **AI Models -> Cloud**. 2. Select a provider. 3. Paste the API key. 4. Click **Verify**. VoiceInk stores verified API keys in Keychain. The provider detail panel also links to the provider's API key page when available. ## Choose Models in Modes [#choose-models-in-modes] Connecting a provider does not automatically change your current workflow. After the key is saved: 1. Open **Modes**. 2. Edit the Mode. 3. Choose a transcription model under **Transcription** or an enhancement provider/model under **AI Enhancement**. 4. Save the Mode. ## OpenRouter Models [#openrouter-models] OpenRouter enhancement models are loaded separately. In the provider detail panel or Mode editor, use **Refresh Models** to fetch the latest available OpenRouter model list. ## Streaming Models [#streaming-models] Some cloud transcription models support realtime partial transcript display. Streaming-only providers require realtime behavior inside Modes. # Custom Models (/docs/custom-models) ## Custom Transcription Models [#custom-transcription-models] Custom transcription models use an OpenAI-compatible audio transcription endpoint. Required fields: | Field | Example | | ------------------ | ------------------------------------------------ | | Display Name | My Custom Transcription | | API Endpoint | `https://api.openai.com/v1/audio/transcriptions` | | API Key | Your provider key | | Model Name | `gpt-4o-mini-transcribe` | | Multilingual Model | On or off | After saving, open **Modes** and choose the custom model from **Transcription -> Model**. ## Custom Enhancement Models [#custom-enhancement-models] Custom enhancement models use an OpenAI-compatible chat completions endpoint. Required fields: | Field | Example | | ------------ | -------------------------------------------- | | Display Name | My Enhancement Model | | Base URL | `https://api.openai.com/v1/chat/completions` | | API Key | Your provider key | | Model Name | `gpt-5.5` | VoiceInk verifies the key before adding a new custom enhancement model. After saving, select it in a Mode under **AI Enhancement**. ## Editing and Deleting [#editing-and-deleting] Use the menu on a custom model row to edit or delete it. Deleting a custom model removes the saved model entry and its associated Keychain API key. # Word Replacements (/docs/word-replacements) ## What Word Replacements Do [#what-word-replacements-do] Word Replacements are deterministic substitutions. VoiceInk applies them after transcription and paragraph formatting, before final cleanup and AI enhancement. Use them for: * Correcting repeated mishears, such as `voice ink` -> `VoiceInk`. * Expanding short phrases, such as `my website` -> `https://tryvoiceink.com`. * Inserting boilerplate. * Normalizing names, products, commands, or punctuation-sensitive text. ## Add a Replacement [#add-a-replacement] 1. Open **Dictionary**. 2. Select **Word Replacements**. 3. Enter the original text in **Original text**. 4. Enter the replacement in **Replacement text**. 5. Press Return or click the plus button. You can enter multiple originals separated by commas: `Voicing, Voice ink, Voiceing` -> `VoiceInk` ## Matching Behavior [#matching-behavior] * Matching is case-insensitive. * For spaced languages, VoiceInk uses word boundaries so partial words are not replaced accidentally. * For non-spaced scripts such as CJK, Thai, and Korean, VoiceInk falls back to substring replacement. * Longer replacement groups are processed first so specific phrases can win over shorter ones. ## Edit, Sort, Delete [#edit-sort-delete] The Word Replacements table can sort by Original or Replacement. Use the edit button on a row to update an existing replacement. Use the remove button to delete it. VoiceInk prevents duplicate original tokens across replacement entries. ## How It Interacts With AI Enhancement [#how-it-interacts-with-ai-enhancement] Because replacements run before AI enhancement, the AI sees the corrected text. This helps enhancement preserve the replacement in the final output. ## Quick Add [#quick-add] Open **Dictionary -> gear -> Quick Add to Dictionary** to assign a shortcut. The Quick Add panel can add either Vocabulary or Word Replacement entries without opening the main window. # Vocabulary (/docs/vocabulary) ## What Vocabulary Does [#what-vocabulary-does] Vocabulary is used as AI enhancement context. It helps the enhancement model preserve names, technical terms, product names, and unusual spellings in the final output. Vocabulary does not directly change the raw transcription. For deterministic substitutions, use [Word Replacements](/docs/word-replacements). ## Add Vocabulary [#add-vocabulary] 1. Open **Dictionary**. 2. Select **Vocabulary**. 3. Type a word or phrase. 4. Press Return or click the plus button. You can add multiple terms at once by separating them with commas. ## Sort and Remove [#sort-and-remove] Vocabulary terms appear as chips. Click the sort label to switch alphabetical order. Click a chip's remove button to delete the term. ## When Vocabulary Is Used [#when-vocabulary-is-used] Vocabulary is included only when AI enhancement runs. If a Mode has AI Enhancement disabled, or enhancement is skipped because the text is below the short-transcription threshold, Vocabulary will not affect the output. ## Good Examples [#good-examples] * Personal names * Company names * Product names * Acronyms * Technical terms * Domain-specific spellings ## Related [#related] * [Mode Settings](/docs/mode-settings) * [Word Replacements](/docs/word-replacements) * [Model Settings](/docs/model-settings) # Filler Words (/docs/filler-words) ## Where Filler Words Live [#where-filler-words-live] Filler words are managed in **AI Models -> gear -> Transcription -> Remove Filler Words**. They are documented here because they affect dictionary-style cleanup, but the control is in AI Models settings. ## Default Filler Words [#default-filler-words] VoiceInk starts with: `uh`, `um`, `uhm`, `umm`, `uhh`, `uhhh`, `hmm`, `hm`, `mmm`, `mm`, `mh`, `ehh` ## How Cleanup Works [#how-cleanup-works] Before formatting and AI enhancement, VoiceInk removes configured filler words from transcription text. Rules: * Matching is case-insensitive. * Only standalone filler words are removed. * A comma or period directly after the filler word can be removed with it. * If no filler words are configured, this cleanup is skipped. ## Add or Remove Filler Words [#add-or-remove-filler-words] 1. Open **AI Models**. 2. Click the gear button. 3. Select **Transcription**. 4. In **Remove Filler Words**, click the plus button to add a word. 5. Remove chips you no longer want cleaned. ## Notes [#notes] Adding common words such as `like`, `right`, or `okay` can remove legitimate content. Add only words you are comfortable losing when they appear as standalone words. # Shortcuts (/docs/shortcuts) ## Recording Shortcuts [#recording-shortcuts] VoiceInk supports a primary recording shortcut and an optional secondary shortcut. Each shortcut can use one of the dedicated modifier keys or a custom key combination. Available shortcut keys include Right Option, Left Option, Right Command, Right Control, Left Control, Right Shift, Fn, and custom combinations. ## Recording Modes [#recording-modes] Each recording shortcut has its own mode: | Mode | Behavior | | ------------ | ---------------------------------------------------------------------------------------------------- | | Toggle | Press once to start recording. Press again to stop. | | Push to Talk | Hold the shortcut to record. Release it to stop. | | Hybrid | A short press starts or stops recording. Holding for at least 0.5 seconds records until you release. | Hybrid is useful if you switch between short dictations and press-and-hold voice notes. ## Second Shortcut [#second-shortcut] Open **Settings -> Shortcuts** and enable **Add Second Shortcut** to configure another global recording shortcut. The primary and secondary shortcuts can use different keys and different recording modes. ## Additional Shortcuts [#additional-shortcuts] The Additional Shortcuts section controls actions that can run without starting a new recording. | Action | What It Does | | --------------------------------- | --------------------------------------------------------------------------------------- | | Paste Last Transcription | Pastes the raw text from your most recent recording. | | Paste Last Enhanced Transcription | Pastes the enhanced text from your most recent recording when enhancement produced one. | | Retry Last Transcription | Runs the last recording again with the current Mode and model settings. | | Cancel Recording | Cancels an active recording. | | Open History Window | Opens the transcription history window. | | Quick Add to Dictionary | Opens a quick dialog to add a word to the dictionary. | Retry Last Transcription is helpful when you want to test a different transcription model, language, or AI enhancement prompt on the same audio. ## Middle-Click Recording [#middle-click-recording] VoiceInk can also start recording from the middle mouse button. 1. Open **Settings -> Shortcuts**. 2. Enable **Middle-Click Recording**. 3. Choose an activation delay in milliseconds. If you release the middle button before the delay finishes, VoiceInk cancels the start action. This helps prevent accidental recordings while scrolling. ## Recorder Panel Shortcuts [#recorder-panel-shortcuts] When the recorder panel is visible, these shortcuts are available: | Shortcut | What It Does | | ------------ | ------------------------------------------------------------------------------ | | Escape | Cancels recording if no cancel shortcut is configured. Press twice to confirm. | | Option + 1–0 | Selects the first ten enabled modes. | ## Mode Shortcuts [#mode-shortcuts] Individual Modes can have their own shortcuts. Configure them from **Modes -> select a Mode -> Shortcut**. A Mode shortcut activates that Mode and starts recording with its saved transcription, AI enhancement, context, output, and auto-send settings. See [Mode Settings](/docs/mode-settings). # Audio Input (/docs/audio-input) Use it when VoiceInk is recording from the wrong microphone, when you switch between a headset and built-in mic, or when you want a fallback order for multiple devices. ## Selected Microphone [#selected-microphone] Selected Microphone pins VoiceInk to one input device. Options include: * **System Default**: Uses the microphone currently selected in macOS. * **Current device**: Records from a specific available microphone. Click **Refresh Microphones** if you connected a new device and it does not appear yet. ## Priority Order [#priority-order] Priority Order lets you build a ranked list of microphones. VoiceInk tries the first available device in the list, then falls back to the next one if the earlier device is disconnected. Use Priority Order when you regularly move between devices, such as: * A desk microphone. * A headset. * The built-in Mac microphone. The Audio view shows which devices are active or unavailable so you can quickly see what VoiceInk will use. ## macOS Input Level [#macos-input-level] If VoiceInk detects the right microphone but accuracy is poor, also check macOS: 1. Open **System Settings -> Sound -> Input**. 2. Select the same microphone. 3. Speak and confirm the input level moves. For best results, use the Mac built-in microphone or a dedicated external microphone. Many wireless headset microphones are optimized for calls and can produce lower-quality transcription audio. ## Recording Feedback [#recording-feedback] Recording feedback controls what you hear and what happens to other audio while VoiceInk records. ### Recording Sounds [#recording-sounds] VoiceInk can play sounds when recording starts and stops. Each sound can be: * A built-in VoiceInk sound. * A custom audio file. * Disabled. Use the test controls to preview sounds, and reset them if you want to return to the defaults. ### Recording Behavior [#recording-behavior] | Setting | What It Does | | --------------------------- | --------------------------------------------------------------------------------------- | | Mute Audio While Recording | Temporarily mutes system audio so app audio is less likely to leak into the microphone. | | Pause Media While Recording | Pauses playing media while recording and resumes it afterward. | | Resume Delay | Waits 0 to 5 seconds before audio or media resumes after recording stops. | These settings are useful when you dictate while music, videos, meetings, or notification sounds may be playing. # General Settings (/docs/settings-general) ## Interface [#interface] **Recorder Style** controls how the recorder appears while you dictate. | Style | Behavior | | ----- | -------------------------------------------------- | | Notch | Shows a compact recorder at the top of the screen. | | Mini | Shows a small floating recorder window. | Both styles can show realtime partial text when the selected transcription model supports streaming. Assistant Mode responses appear in the recorder and stay available for follow-up prompts until you close it. ## General [#general] The General section contains app-wide behavior: | Setting | What It Does | | ------------------ | ----------------------------------------------------------- | | Hide Dock Icon | Runs VoiceInk as a menu bar app without a Dock icon. | | Launch at Login | Starts VoiceInk automatically when you sign in to your Mac. | | Auto-check Updates | Checks for new app versions in the background. | | Show Announcements | Shows in-app announcements for important changes. | | Check for Updates | Runs an immediate update check. | ## History and Privacy [#history-and-privacy] Transcript and audio retention controls are not in General Settings. Open **History** from the sidebar, then click the **History Settings** button beside the search field. See [Transcription History](/docs/transcription-history#automatic-cleanup) for cleanup options, what each option deletes, and how to run cleanup manually. ## Backup [#backup] Backup lets you export and import VoiceInk settings. Use **Export Settings** before moving to a new Mac, testing a large configuration change, or sharing a setup. Use **Import Settings** to restore a saved configuration. The backup includes app settings, Modes, dictionary entries, model configuration, custom providers, and other saved preferences. ## Diagnostics [#diagnostics] Use **Export Logs** when troubleshooting with support. Logs help diagnose model downloads, provider connection issues, permissions, audio input behavior, paste behavior, and other app-level problems. # Common Issues (/docs/common-issues) ## Start Here [#start-here] Most VoiceInk issues come from one of five places: * A missing macOS permission. * The wrong microphone being selected. * A Mode using settings you did not expect. * A local model that is too slow for the Mac. * Paste or clipboard timing in the target app. ## Permissions [#permissions] VoiceInk needs Microphone permission to record. It also needs Accessibility permission to paste reliably, read selected text for context, and control some app interactions. Screen Recording permission is only needed when you use screen context in a Mode. See [Installation](/docs/installation) for the setup flow. ## Recording and Transcription [#recording-and-transcription] If you get no output or poor accuracy, start with [Transcription Failing](/docs/transcription-failing). Check the **Audio** view, confirm the macOS input level moves, and make sure the selected Mode has a working transcription model. If output is slow, see [Transcription Taking Too Long](/docs/transcription-taking-too-long). Local models can be fast on Apple Silicon, but Intel Macs usually work better with cloud transcription. If output repeats text or invents words, see [Repeated Text or Hallucinations](/docs/repeated-text-or-hallucinations). ## Modes and Automatic Changes [#modes-and-automatic-changes] Modes can automatically switch VoiceInk settings based on the active app, website, trigger group, default Mode, or a Mode shortcut. If your model, language, prompt, context sources, output behavior, or auto-send setting changes unexpectedly, see [Model or Language Changing Automatically](/docs/model-or-language-changing-automatically). ## Pasting [#pasting] VoiceInk pastes by placing the final text on the clipboard and sending Command+V. If the text transcribes but does not paste, or if your old clipboard content appears instead, see [Clipboard Issues](/docs/clipboard-issues). Paste controls are in **Settings -> Pasting**. ## License and Activation [#license-and-activation] If you need your macOS license key, open the Polar portal: [https://polar.sh/beingpax/portal/request](https://polar.sh/beingpax/portal/request) Use the same email address you used when purchasing VoiceInk. For more detail, see [License Management](/docs/license-management). macOS and iOS licenses are separate. See [iOS and macOS Licenses Are Separate](/docs/ios-macos-license-confusion). ## Still Need Help [#still-need-help] Use **Dashboard -> Copy System Info** or **Settings -> Diagnostics -> Export Logs** before contacting support. Those details make it much easier to understand permissions, model setup, provider configuration, and recent app state. # Transcription Taking Too Long (/docs/transcription-taking-too-long) ## Where Time Is Spent [#where-time-is-spent] A recording can spend time in several stages: 1. Loading or warming the transcription model. 2. Transcribing the audio. 3. Cleaning and formatting the text. 4. Running AI enhancement. 5. Pasting the result or generating an Assistant response. Use **Dashboard -> Model Performance** to compare model speed over the last 7 days, last 30 days, this year, or all time. Use **History -> Analyze** for recording-level timing. ## Local Models [#local-models] Local transcription is usually fastest on Apple Silicon Macs. If you use an Intel Mac, cloud transcription providers are often much faster. To improve local speed: * Use **Parakeet V3** or **Parakeet V2** for fast local dictation. * Try a smaller Whisper model if a large model is slow. * Enable **Prewarm Model** in **AI Models -> gear -> Transcription**. * Keep **Voice Activity Detection** enabled so silence is skipped more efficiently. Large local models may take longer the first time they are prepared or loaded. ## Cloud Transcription [#cloud-transcription] Cloud transcription speed depends on provider latency, audio length, and your network connection. Groq, Deepgram, Mistral, Gemini, ElevenLabs, Soniox, Speechmatics, AssemblyAI, xAI, and Cartesia are available as cloud transcription options depending on the provider capability. See [Cloud Providers](/docs/cloud-providers). ## AI Enhancement [#ai-enhancement] AI enhancement adds a second model call after transcription. If the raw transcription is fast but the final output is slow, check the Mode's AI enhancement provider and model. To reduce enhancement delay: * Disable AI enhancement for simple Dictation Modes. * Enable **Skip short transcriptions** in **AI Models -> gear -> Enhancement**. * Increase or decrease **Timeout Duration** based on your tolerance. * Keep **Retry on timeout** enabled if quality matters more than speed. * Try a faster enhancement provider or model. See [Model Settings](/docs/model-settings) and [Mode Settings](/docs/mode-settings). ## File Transcription [#file-transcription] Long audio and video files naturally take longer than live dictation. The Transcribe view shows queue states such as Loading model, Processing audio, Transcribing, Enhancing, Completed, and Failed, so you can see where the delay is happening. See [Transcribe Audio Files](/docs/transcribe-audio-files). # Model or Language Changing Automatically (/docs/model-or-language-changing-automatically) ## Why Settings Change [#why-settings-change] VoiceInk records through Modes. A Mode can save its own transcription model, realtime setting, language, formatting behavior, AI provider, AI model, prompt, context sources, output behavior, auto-send key, and shortcut. When a Mode becomes active, those saved settings are used for the recording. ## Check the Active Mode [#check-the-active-mode] Open **Modes** and look at your enabled Modes. Settings may change because: * A Mode matches the active app. * A Mode matches the active website. * A trigger group matches the current app. * A Mode is marked as the default. * You used a Mode shortcut. See [Mode Triggers](/docs/mode-triggers) for how matching works. ## The Default Mode Applies Everywhere [#the-default-mode-applies-everywhere] A default Mode is used when no more specific app, website, or group trigger matches. If the default Mode uses a different model or language, VoiceInk can feel like it is changing settings on its own. To change this: 1. Open **Modes**. 2. Select the Mode marked as default. 3. Turn off **Set as default**, or make a different Mode the default. ## Disable or Reorder Modes [#disable-or-reorder-modes] If a Mode is matching too often, you can disable it without deleting it. You can also reorder Modes so your most important workflows are easier to inspect and manage. If two Modes feel similar, open each one and compare the transcription, AI enhancement, context, output, and auto-send sections. See [Mode Settings](/docs/mode-settings). ## Language Changes [#language-changes] Language is configured per Mode. Open the active Mode and check its transcription language setting. Some models and providers handle language selection differently. If a provider ignores a language choice or performs poorly in a language, try another transcription model from [AI Models](/docs/ai-models). # Repeated Text or Hallucinations (/docs/repeated-text-or-hallucinations) ## What This Looks Like [#what-this-looks-like] Hallucinations are words, phrases, tags, or repeated text that were not clearly present in the audio. They can happen when the recording is too quiet, too noisy, too short, or when a model guesses during silence. VoiceInk already removes some common transcription artifacts, but model choice and audio quality still matter. ## Turn On Voice Activity Detection [#turn-on-voice-activity-detection] Voice Activity Detection helps the transcription model ignore silence and low-confidence audio. 1. Open **AI Models**. 2. Click the gear button. 3. Select **Transcription**. 4. Enable **Voice Activity Detection**. ## Try a Different Transcription Model [#try-a-different-transcription-model] Open the Mode you are using and change its transcription model. Good options to try: * **Parakeet V3** for fast local dictation on supported Macs. * **Parakeet V2** if V3 is not the best fit for your audio. * **Whisper Base** or **Whisper Tiny** if a larger Whisper model is over-generating. * A cloud transcription provider if local models struggle with your microphone or accent. See [AI Models](/docs/ai-models) and [Recommended Models](/docs/recommended-models). ## Check Filler Words [#check-filler-words] Open **AI Models -> gear -> Transcription -> Remove Filler Words**. Removing filler words such as `uh` and `um` can clean up starts and pauses. Avoid adding common words you may actually dictate, because they will be removed whenever they appear as standalone words. See [Filler Words](/docs/filler-words). ## Use Retry Last Transcription [#use-retry-last-transcription] If a recording produced repeated or imaginary text, change the Mode or model settings, then use **Retry Last Transcription** from **Settings -> Shortcuts**. This runs the same audio again with your current configuration, so you can compare models without recording again. ## Improve the Audio Source [#improve-the-audio-source] Poor input audio can make hallucinations worse. Use **Audio** to select the best microphone, and check the input level in **System Settings -> Sound -> Input**. Built-in Mac microphones and dedicated external microphones usually work better than wireless headset microphones. # Transcription Failing or No Output (/docs/transcription-failing) ## No Text Appears [#no-text-appears] If VoiceInk records but no text appears, check these first. ## 1. Check Permissions [#1-check-permissions] VoiceInk needs Microphone permission to record audio. It also needs Accessibility permission for reliable paste delivery and selected-text context. Screen Recording permission is needed only when a Mode uses screen context. Open **System Settings -> Privacy & Security** and review: * Microphone. * Accessibility. * Screen Recording, if you use screen context. After changing permissions, restart VoiceInk. ## 2. Check Audio Input [#2-check-audio-input] Open **Audio** in the VoiceInk sidebar. * Use **Selected Microphone** if you want one specific device. * Use **Priority Order** if you want VoiceInk to fall back through multiple microphones. * Click **Refresh Microphones** after connecting a new device. Also open **System Settings -> Sound -> Input** and confirm the input level moves when you speak. ## 3. Check the Active Mode [#3-check-the-active-mode] Open **Modes** and select the Mode you are recording with. Confirm: * The Mode is enabled. * It has a valid transcription model selected. * The selected provider or local model is available. * The language setting matches what you are speaking. * AI enhancement is not the only place you expect output from. If a trigger is selecting the wrong Mode, see [Model or Language Changing Automatically](/docs/model-or-language-changing-automatically). ## 4. Check AI Models [#4-check-ai-models] Open **AI Models**. * For local models, confirm the model is downloaded and available. * For cloud transcription providers, confirm the API key is configured and the provider supports transcription. * For custom models, confirm the endpoint and model name are correct. If you are using an imported local Whisper `.bin` model, make sure it is still present and readable. ## Inaccurate Transcription [#inaccurate-transcription] If VoiceInk produces text but it is inaccurate: * Try the Mac built-in microphone or a dedicated external mic instead of a wireless headset mic. * Move closer to the microphone and reduce background noise. * Try a different transcription model in the active Mode. * Add deterministic fixes in [Word Replacements](/docs/word-replacements). * Add important names and terminology in [Vocabulary](/docs/vocabulary) when the Mode uses AI enhancement. ## Still Not Working [#still-not-working] Open **Dashboard -> Copy System Info** and **Settings -> Diagnostics -> Export Logs**, then contact support with those details. # Clipboard Issues (/docs/clipboard-issues) ## How VoiceInk Pastes Text [#how-voiceink-pastes-text] VoiceInk pastes by temporarily placing the final text on your clipboard and sending Command+V to the active app. If **Keep Clipboard Content** is enabled, VoiceInk restores your previous clipboard content after the paste. Paste settings live in **Settings -> Pasting**. ## Old Clipboard Content Gets Pasted [#old-clipboard-content-gets-pasted] This usually means the clipboard was restored before the target app finished pasting. Fix it: 1. Open **Settings -> Pasting**. 2. Keep **Keep Clipboard Content** enabled. 3. Increase **Restore Delay**. Start with 500ms or 1s. If the target app is slow, try 2s or longer. ## Transcription Replaces Your Clipboard [#transcription-replaces-your-clipboard] If your previous clipboard content disappears after each recording, **Keep Clipboard Content** is off. Open **Settings -> Pasting** and enable **Keep Clipboard Content**. ## Text Is Transcribed But Nothing Pastes [#text-is-transcribed-but-nothing-pastes] Try these in order: 1. Click into the target app or text field before recording. 2. Confirm VoiceInk has Accessibility permission in macOS System Settings. 3. Open **Settings -> Pasting** and switch **Paste Method** from Default to AppleScript. 4. If you use a custom keyboard layout, keep AppleScript paste enabled. If the text appears in History but not in the target app, transcription worked and the issue is with focus, permissions, or paste delivery. ## Extra Space After Paste [#extra-space-after-paste] If VoiceInk adds or does not add a trailing space the way you expect, open **AI Models -> gear -> Transcription** and change **Add Space After Paste**. # iOS and macOS Licenses Are Separate (/docs/ios-macos-license-confusion) ## My macOS License Does Not Work on iPhone [#my-macos-license-does-not-work-on-iphone] A VoiceInk for macOS license activates the Mac app. It does not activate VoiceInk for iPhone. The iOS app has its own distribution and purchase flow. If the iOS app asks for payment or account information, follow the iOS app's instructions rather than entering your macOS license key. ## I Bought a Multi-Device macOS License [#i-bought-a-multi-device-macos-license] A multi-device macOS license covers multiple Macs. It does not count an iPhone as one of those Mac devices. If you purchased the wrong license by mistake, [contact support](/contact) with your purchase email. ## Retrieve a macOS License [#retrieve-a-macos-license] Use the Polar portal to retrieve or manage your Mac license: [https://polar.sh/beingpax/portal/request](https://polar.sh/beingpax/portal/request)