VoxaFlow

User guide

The same guide that ships with the app. Inside VoxaFlow: menu bar › Guide.

VoxaFlow is a dictation app for macOS. Hold a key, speak and release. VoxaFlow types the cleaned-up text wherever your cursor is, in any app. Speech recognition and text cleanup run entirely on your Mac.

The same content is available in the app under menu bar › Guide …, with practice fields and a guided practice.

Contents

  1. Installation and setup
  2. Dictation
  3. Hands-free mode
  4. Command mode
  5. Corrections, lists and formatting
  6. Dictionary
  7. Snippets
  8. Style per app
  9. Overview, history and shortcuts
  10. Sync via iCloud
  11. Settings
  12. Privacy
  13. Troubleshooting

Installation and setup

Requirements

Download the beta from voxaflow.ai/download: open the DMG and in the installer window drag VoxaFlow onto the Applications folder, as the arrow shows. The app is signed with Developer ID and notarized by Apple; from version 0.6.4 on, updates arrive automatically.

First launch

On first launch the welcome window opens. After the welcome it guides you through three steps; use “Continue” and “Back” at the bottom:

  1. Grant permissions. VoxaFlow needs two:

    Permission Purpose
    Microphone to record your dictation
    Accessibility to detect the dictation key in every app, insert text at the cursor and read selected text

    VoxaFlow does not need Input Monitoring.

    Click “Allow” for each and switch VoxaFlow on in System Settings. VoxaFlow picks up the permissions without a restart. Once “Dictation key active” shows a green checkmark, everything is ready.

  2. Check the dictation key. Choose the dictation key here (fn, right ⌥ or right ⌘). When you press it, the waves in the background swell. VoxaFlow also checks whether the fn/🌐 key is also assigned to another function. If so, a notice with an “Open Keyboard settings” button appears; set “Press 🌐 key to” to “Do Nothing” there. “Test key” checks whether the key reaches VoxaFlow.

  3. Load the models. On first launch VoxaFlow downloads the speech recognition model (about 470 MB) and the language model (about 2.3 GB; on Macs with 8 GB of memory the smaller one, about 1.7 GB). Depending on your connection this takes a few minutes. Afterwards VoxaFlow works offline. The same step has “Open VoxaFlow at login”, switched on by default, so the dictation key is ready right after every restart. You can change this anytime under Settings › General.

“Start practice” opens the guided practice. It shows all main features in five steps and checks off each step as soon as it worked.

Dictation

  1. Click into a text field.
  2. Hold fn and speak.
  3. Release fn. Short sentences appear after about half a second.

While you speak, a pill appears at the bottom of the screen:

Display Meaning
Red dot with level meter, “Listening” VoxaFlow is recording
“Writing …” the text is being cleaned up
“Nothing recognized” nothing was said, nothing is inserted
Orange notice the fn key is also assigned to something else

Esc cancels a recording. If you press fn briefly without speaking, VoxaFlow inserts nothing.

What VoxaFlow improves

You say VoxaFlow writes
“hi sarah um thanks for your email I’ll get back to you by friday” Hi Sarah, thanks for your email. I’ll get back to you by Friday.
“how much does the license for ten users cost” How much does the license for ten users cost?

Which language model?

Under Settings › AI › Model you choose which model cleans up the text. Qwen3 4B is recommended, and Qwen3.5 2B on Macs with 8 GB of memory: the large model needs about 2.3 GB of memory, and if macOS has to swap, every dictation gets noticeably slower. If you choose the large model anyway, VoxaFlow shows a notice.

Model Short dictation Accuracy in test
Qwen3 4B, local 0.2–0.6 s 99 %
Qwen3.5 2B, local about half as long as Qwen3 4B 93 %
Apple Intelligence (macOS 26 or later) 0.6–2 s 87 %
Off (rules only) instant no corrections, lists or style

Measured on a Mac with M4 Pro using the same test recordings, Qwen3.5 2B compared on an M4. Times may differ on other Macs. Long dictations take longer because more text is generated; with Qwen3 4B a 25-second dictation takes about 1.5 s (M4 Pro) to 3 s (M4). Qwen3.5 2B leaves filler words in more often and corrects dictionary terms less reliably. In Low Power Mode cleanup takes about twice as long.

Long dictations

From about 8 seconds of recording on, VoxaFlow recognizes and cleans up finished sentences while you are still speaking. After you release the key only the last part is left: in a test a 95-second dictation appeared after 1.2 seconds (Qwen3 4B, M4) instead of about 5 seconds. No pauses are needed; VoxaFlow only splits at sentence ends. Corrections across a sentence end (“… on Monday. No, on Tuesday.”) and continued enumerations are cleaned up together.

Hands-free mode

For longer texts you don’t have to hold fn.

  1. Tap fn twice in quick succession.
  2. Speak as long as you like. The pill shows “Hands-free · press key to stop”.
  3. Tap fn once to stop.

If Apple’s own dictation is also set to “Press 🌐 twice”, both start at the same time. VoxaFlow warns about this under Settings › Keyboard.

Command mode

Command mode rewrites existing text with a spoken instruction.

  1. Select the text.
  2. Hold ⌃ Control and fn. The order doesn’t matter: if fn comes first, you have one second to add ⌃ Control. Shift or other keys are not needed.
  3. Say what should happen to the text and release.

VoxaFlow replaces the selection with the rewritten version. The pill turns purple and shows “Command”. If nothing is selected, VoxaFlow dictates as usual.

Examples: “Make this shorter.”, “Make it friendlier.”, “Translate this into German.”, “Turn this into a bulleted list.”, “Fix the spelling.”

Corrections, lists and formatting

Self-corrections: signal words such as “no”, “sorry”, “I mean” or “actually” tell VoxaFlow that the previous detail is replaced.

You say VoxaFlow writes
“let’s meet on thursday at 3 pm no sorry at 4 pm” Let’s meet on Thursday at 4 pm.

Lists: “first, second, third” produce a numbered list.

You say VoxaFlow writes
“for the project we need first a logo second a website and third a flyer” For the project we need:
1. a logo
2. a website
3. a flyer

Paragraphs and punctuation: “new paragraph”, “new line”, “period”, “comma” and “question mark” are applied. Usually this isn’t necessary because VoxaFlow adds punctuation and paragraphs on its own.

Dictionary

Speech recognition sometimes mishears names, technical terms and abbreviations. Under Settings › Dictionary you use + to define how a term is spelled and add typical mistakes under “Misheard as”, separated by commas. Double-click a term to edit it; if you add a term that already exists, VoxaFlow merges both.

Term Misheard as You say VoxaFlow writes
VoxaFlow Voxa Flow, Box a Flow “I’m testing box a flow” I’m testing VoxaFlow.

Misheard variants are replaced directly. In addition, the language model receives the terms and also recognizes similar-sounding variants.

Learning from corrections: when you correct a word right after inserting, e.g. “MLS” to “MLX”, a suggestion appears at the top of the dictionary; accept it with “Add” or ignore it. VoxaFlow reads only the text field it just wrote into, never a password field. Faster with the shortcut “Add selection to dictionary”: select the wrong word, press the shortcut, type the correct spelling.

Import and export: “Import and export” saves the dictionary as JSON or as a CSV file for Excel and Numbers (columns “Term;Misheard as”). Before an import VoxaFlow shows how many terms are new, changed or the same, and you choose whether the file adds to or replaces the dictionary.

Snippets

Snippets are text blocks. Under Settings › Snippets you create a trigger and a text. When you say the trigger, VoxaFlow inserts the text verbatim, without rewriting by the language model.

Trigger Inserted text
my address 12 Example Street, Springfield 12345
my calendar link https://cal.example.com/meeting

Pick triggers you would not say by accident in normal speech. Snippets too can be edited by double-click and imported or exported as JSON or CSV (columns “Trigger;Text”).

Style per app

VoxaFlow detects which app you are writing in and adapts the tone. Under Settings › AI › Style per app you choose Formal, Casual or Very casual for each category.

Category Examples Default
Email Mail, Outlook, Spark Formal
Chat Slack, Messages, WhatsApp, Teams Casual
Code & terminal Xcode, VS Code, Terminal, Claude Casual, technical terms stay unchanged
Documents & notes Notes, Pages, Word, Notion Casual
All other apps Casual

Style changes punctuation, capitalization and line breaks, not the content and not your wording. VoxaFlow does not add greetings or sentences you did not say.

Overview, history and shortcuts

Overview: shows for today, 7 days, 30 days, the year or all time how many words you dictated, how much time you save compared with typing, your speaking speed and the corrections, split into AI cleanup, dictionary and “by you afterwards”. For the time saved VoxaFlow assumes a typing speed of 40 words per minute; change it right in the tile. There are also words per day, when you dictate and your top apps. Only numbers are counted, never the text.

History: the search field searches the final text and the original, the filter limits the list to a period or an app. Copy by double-click, with ⌘ C or with the icon that appears in the row on hover. On the right are “Copy”, “Copy original” and, under “…”, “Clean up again” in another style. Copy several entries together or export them as text, CSV or JSON. How long dictations are kept is set under iCloud & Data.

Menu bar: “Paste last dictation again”, “Copy last dictation” and recent dictations as a submenu. If no text field is active while you dictate, e.g. on the desktop, the text stays on the clipboard; the pill then shows “No text field”.

Shortcuts are set under Settings › Keyboard:

Action What it does
Open quick history window with search across recent dictations; ↩ inserts into the app you are writing in, ⌘ C copies, ⌘ ↩ inserts the original
Paste last dictation again inserts the newest dictation at the cursor once more
Process last recording again recognizes the last recording again, e.g. after an error; the audio is only kept in memory
Add selection to dictionary select a misspelled word, press the shortcut, type the correct spelling

Sync via iCloud

Under Settings › iCloud & Data you sync dictionary, snippets, settings and statistics between your Macs, and optionally the history. The data lives in the “VoxaFlow” folder in your iCloud Drive and is end-to-end encrypted (AES-256): only your devices have the key, neither Apple nor VoxaFlow can read the data.

  1. Turn on “Sync with iCloud” on your first Mac. VoxaFlow creates the key.
  2. Open “Recovery code …” and keep the code safe, for example in your password manager.
  3. Turn on sync on every further Mac and enter the code there.

Every Mac writes only its own files, changes apply per entry, and deleted entries disappear on all Macs. If iCloud Drive is off, VoxaFlow says so. “Delete all data in iCloud” removes the synced data from iCloud; the data on the Macs stays.

Settings

Open the VoxaFlow window from the VoxaFlow icon in the menu bar › Settings … or with ⌘ ,. On the left is a sidebar like in System Settings; the search field at the top also finds settings by keywords such as “clipboard”. While the window is open, VoxaFlow shows in the Dock and in ⌘ ⇥.

Section What you set there
Overview words, time saved, speaking speed, corrections, top apps
History your dictations with search, copy, cleaning up again and export
General app language (system, Deutsch, English), open at login
Keyboard dictation key (fn, right ⌥, right ⌘), key check with key test, shortcuts
Inserting restore clipboard, keep text without a text field, learn from corrections
Audio microphone, fallback microphones, sound output, mute other audio, start sound
AI dictation language, model, style per app, status of the local models
Dictionary terms, recognition mistakes, suggestions, import and export
Snippets text blocks, import and export
iCloud & Data sync via iCloud, history retention, statistics, backup
Permissions microphone and Accessibility, repair permissions
About version, updates (check automatically, check now), guide, models used

Many settings have a question mark next to them. Clicking it explains the setting with an example.

Audio devices: “System default” follows the device set in macOS. A device you choose stays selected; if it is not connected, VoxaFlow tries the fallback microphones in order and then the system default until the device is back. If the device changes during a recording, VoxaFlow keeps recording.

Tip for AirPods: choose the Mac’s microphone. Otherwise macOS switches the AirPods to headset mode and audio quality drops.

Backup: under iCloud & Data you put settings, styles, dictionary and snippets into one file, optionally with the history, and restore it on another Mac.

Privacy

Troubleshooting

Nothing happens when I press fn. Check under Settings › Permissions that both permissions are green. Under Settings › Keyboard, “Test key” shows whether the key arrives. Many external keyboards do not send fn to the Mac; set the dictation key to the right ⌥ Option instead.

fn opens emoji or Apple’s dictation. In System Settings › Keyboard, set “Press 🌐 key to” to “Do Nothing”. If Keyboard › Dictation has the shortcut “Press 🌐 twice”, change it or turn Apple’s dictation off.

Permission granted, but VoxaFlow says it is missing. macOS ties the entries to the app’s signature. If VoxaFlow was installed with a different signature (for example a development build instead of a release), they no longer apply. Under Settings › Permissions, click “Repair permissions” and then switch VoxaFlow on again under Accessibility. Releases and updates always have the same signature.

Microphone switched or AirPods disconnected. VoxaFlow detects device changes by itself. If a microphone delivers no audio after pressing the key, VoxaFlow restarts the recording once and otherwise shows “Microphone not responding”.

The text appears in the wrong language. With very short sentences the language detection can be wrong. Set the dictation language to English or German under Settings › AI.

A word is always misrecognized. Add it to the dictionary with its typical mistakes.

The start sound is missing or annoying. Under Settings › Audio › Sounds choose a different start sound or “No start sound” and check the sound output.

Cleanup takes a long time. On first launch VoxaFlow downloads the models; the menu bar shows the progress. Check the model under Settings › AI: Qwen3.5 2B is about twice as fast as Qwen3 4B, Apple Intelligence considerably slower. In Low Power Mode cleanup takes about twice as long.

The text is not inserted. Check that the cursor is in a text field. Password fields block pasting on purpose. If no text field is active at all, e.g. on the desktop, the text is on the clipboard and the pill shows “No text field”. The last text is also available in the menu bar under “Paste last dictation again” and “Copy last dictation” and in the history.

VoxaFlow isn’t running after a restart. Turn on “Open VoxaFlow at login” under Settings › General. If it says macOS still has to allow this, click “Open Login Items” and switch VoxaFlow on under System Settings › General › Login Items.

None of this helps. Choose “Report a Problem …” in the menu. VoxaFlow creates a diagnostic file (version, Mac model, settings, log of the last 24 hours, no dictated text, no audio, no history) and opens an email to hello@voxaflow.ai with the file attached. Nothing is sent until you click Send yourself.