How to transcribe audio to text in Audacity
Audacity transcribes audio to text for free, offline, through the Whisper Transcription effect in the OpenVINO AI Tools plugin. It runs in Audacity 3.7.x and does not run in Audacity 4 yet. The words land on a label track under your recording, and you export them as plain text, SRT or WebVTT subtitles.
This guide covers what to install, how to prepare the recording, which Whisper model to pick, how to fix the label track and export a transcript or subtitles, and the effect’s limits.
Tested on Audacity 3.7.9 on Windows 11, with the plugin installed through MuseHub; the macOS build (v3.7.1-R4.2-beta-3) is a beta.
Transcription in Audacity
What you need
- Audacity 3.7.x. The plugin needs Audacity 3.7.4 or a later 3.x release, and the Audacity team says it is not yet compatible with Audacity 4. In MuseHub, that is the app listed as Audacity 3; the one listed as Audacity installs Audacity 4. Both are free and install side by side. Check your version under Help ▸ About Audacity.
- OpenVINO AI Tools. A free, open-source plugin pack from Muse Group and Intel, listed under that name on MuseHub; its own installer calls it OpenVINO AI Plugins for Audacity. It adds five AI effects to Audacity 3; Whisper Transcription is the one this guide uses, and how to edit audio with AI in Audacity covers the other four.
- A speech recording. An interview, a podcast episode, a lecture, a voice memo. The effect is built for spoken words; it also runs on sung vocals, with rougher results.
- Time and disk space. The models run on your own processor, so a long file takes real minutes, and the model files are large. Windows has a stable build; the macOS build is a beta for macOS 12 and later. On Linux the plugin is compiled from source; MuseHub runs on Windows and macOS only.
Step 1: Install the plugin and check it loaded
Install Audacity 3 from the Apps section of MuseHub, then open Plugins and install OpenVINO AI Tools. MuseHub launches the plugin’s own installer, Setup - OpenVINO AI Plugins for Audacity, which asks for a destination folder and for the models; there is no choice of effects. The Whisper models on offer are Base, Small, Small (English-only) with experimental speaker segmentation, Medium, and Large v1, v2 and v3. With Install recommended models, the only Whisper model ticked is Base, next to the models for the other effects (at least 4.04 GB in total), so tick Small as well, because it is the model this guide recommends. Tick every model you expect to use now, because the effect lists only the models installed with it. The model files are what take the disk space.

Restart Audacity and open the Effect menu. Since Audacity 3.7.4, OpenVINO Whisper Transcription sits in its OpenVINO AI Effects group with the other OpenVINO effects, not under Analyze as older guides say. A MuseHub install sets the plugin’s module, mod-openvino, to Enabled on its own. If the effect is still missing, open Edit ▸ Preferences ▸ Modules (on a Mac, Audacity ▸ Preferences ▸ Modules), check that mod-openvino is set to Enabled, click OK and restart.
Step 2: Prepare the recording
Ten minutes of cleanup saves an hour of correcting half-heard words later.
- Open the file with File ▸ Open. Audacity reads WAV, MP3, FLAC and the other common formats.
- Remove the noise. Air conditioning, traffic and room hum all blur consonants. Select the whole track with Select ▸ All (Ctrl+A, ⌘A on a Mac) and run Effect ▸ OpenVINO AI Effects ▸ OpenVINO Noise Suppression, which is tuned for speech and needs no noise sample. Its Preview button stays greyed out, so apply it, listen, and undo with Ctrl+Z (⌘Z) if the voice turns hollow. The built-in route is Effect ▸ Noise Removal and Repair ▸ Noise Reduction, which wants a second of silence to learn the noise first; noise reduction in Audacity walks through both.
- Even out the level if one speaker is much quieter than the other. Effect ▸ Volume and Compression ▸ Normalize raises the whole track to a set peak; a compressor narrows the gap between loud and quiet passages.
- One track. The effect processes a mono or a stereo track. If your project has several tracks (a host on one, a guest on another), select them and use Tracks ▸ Mix ▸ Mix and Render to make one track first. The effect also runs on several selected tracks at once. It transcribes them one after another and puts every label on one label track, with nothing to show which track a line came from, so mix first.
- Select what you want transcribed. Everything with Ctrl+A, or drag across the part you need. A ten-minute test region shows whether the model and language are right before you commit a two-hour file.
Step 3: Run Whisper Transcription
Open Effect ▸ OpenVINO AI Effects ▸ OpenVINO Whisper Transcription. Four settings matter; the rest you will rarely touch.

Whisper Model is the speed-versus-accuracy choice. Each model is a separate file, and only the ones you installed appear in the list.
| Model | Best for | Time per hour of audio on a CPU |
|---|---|---|
| base | Clean English, a quick draft | about 20 minutes |
| small | Accents, noisier audio, most languages | about 40 minutes |
| small.en-tdrz | English with experimental speaker-change detection | about 40 minutes |
| medium | Reliable work in many languages | about 90 minutes |
| large-v3 | Final subtitles, hard recordings | 3 to 5 hours |
The times come from the Audacity team’s own CPU estimates (base at roughly 0.3× the audio length, large-v3 at 3–5×) and vary with your processor. On the Windows 11 machine this guide was tested on, small took about 17 seconds on the CPU for a 1:40 clip, first-run compile included, far under the table’s figure; how close your own runs come to either depends on the CPU. Start with small. In a ten-recording benchmark posted on the Audacity forum, small matched large-v3 on accuracy in a fifth of the time, and medium took longer than small for a lower score. Move up to large-v3 only for a recording small gets wrong.
Mode is transcribe or translate. Translate outputs English whatever language was spoken, which is a fast way to get an English draft of a foreign-language interview.
Source Language defaults to auto, which works on long files and stumbles on short clips. Whisper handles 99 languages; set the language yourself when you know it.
OpenVINO Inference Device picks the processor that runs the model. CPU works everywhere; on an Intel Core Ultra laptop, the NPU option leaves the CPU free.
Under Advanced Options, Initial Prompt is the one to remember. Type the names, product names and jargon that appear in the recording, spelled the way you want them, and the model uses that spelling. Max Segment Length sets the longest label in characters; 1 gives you one word per label, which is useful for word-level captions and unreadable for anything else.
Click Apply. The first run on a device takes an extra 10–30 seconds while the model compiles; later runs skip that, and a progress bar shows the rest.
Step 4: Read and edit the label track
When the effect finishes, a new label track named after the model, such as Transcription(small), appears under the audio. Each label is one segment of speech with its start and end time and the words the model heard. Play the recording and the labels pass in time with the voice, so every mistake is anchored to the moment it was spoken.

To fix a word, click inside the label and type; the label turns white while it is open for editing. To remove a label, right-click it and choose Delete Label. Drag a region label’s chevron handles to move its start or end. For a long transcript, Edit ▸ Labels ▸ Label Editor opens the Edit Labels dialog. It lists every label in a table with its track, start time, end time and text, and you double-click a cell (or press F2) to edit it, which is faster than scrolling the timeline.

Label tracks save with the project, so File ▸ Save Project keeps the transcript with the audio for the next session.
Step 5: Export the transcript or subtitles
Choose File ▸ Export Other ▸ Export Labels, name the file (it starts as the label track’s name), and pick the file type under Save as type. The list holds four, Text files (.txt), SubRip text file (.srt), WebVTT file (.vtt) and Podcast Chapters (.json).

| File type | What the file contains | Use it for |
|---|---|---|
| .txt | One line per label, prefaced by the start and end time in seconds, then the text, tab-separated | show notes, a searchable log, a spreadsheet |
| .srt | Numbered subtitle blocks with timecodes | captions for YouTube and most video editors |
| .vtt | WebVTT subtitle blocks with timecodes | HTML5 video players and web platforms |
| .json | A Podcasting 2.0 chapters file: each label’s start time and text, no end time | chapter markers for an episode, from a label track of chapter titles rather than a full transcript |
The .txt export carries timestamps on every line, which surprises people who wanted a clean transcript. To strip them, open the file in a spreadsheet with tab as the delimiter, delete the two time columns and copy what is left.
For captions, the .srt or .vtt file drops straight into the video editor or the upload form. Read it against the video once. The model segments on pauses rather than meaning, so a subtitle occasionally breaks mid-phrase.
What it will not do
- Name the speakers. The transcript has no Speaker 1 and Speaker 2. The small.en-tdrz model detects speaker turns and writes alternating labels to two label tracks, and it is marked experimental and English-only. Anything more needs a service or your own pass.
- Punctuate like an editor. Whisper adds periods and commas from context, and it guesses. Expect to fix run-on sentences and the odd question mark.
- Transcribe music well. Lyrics under a full mix, several people talking at once and heavy accents on a poor microphone all lower accuracy. If you need the notes rather than the words, the best AI tools for musicians covers music transcription, which is a different job.
- Run fast on any machine. A two-hour file with large-v3 on a five-year-old laptop can take most of a day. Pick the model for the hardware.
Common mistakes
- Installing Audacity 4. The effects do not appear because the plugin does not support it yet. Install Audacity 3 alongside it and use that for transcription.
- Skipping the models during install. The effect opens with an empty model list and a greyed-out Apply button. The plugin’s developers add models by running its installer again over the existing install, without uninstalling, and ticking only the missing models.
- The module is disabled. Rare after a MuseHub install, which enables it. Preferences ▸ Modules, set mod-openvino to Enabled, restart.
- Nothing selected. The effect works on the selection. Ctrl+A, then run it.
- Exporting audio instead of labels. File ▸ Export ▸ Export Audio writes a sound file. The transcript is under Export Other ▸ Export Labels.
- Running large-v3 first. Test with small on a short region; only then decide whether the bigger model earns its hours.
Tips for a better transcript
- Clean first, transcribe second. Noise Suppression before Whisper improves the words more than any model change.
- Feed it the names. Put every proper noun in Initial Prompt. Most corrections on a real transcript are names and brands.
- Record for the model. A microphone close to the mouth, one voice at a time and no music bed under the speech make a bigger difference than any setting. How to edit a podcast in Audacity covers the chain from raw recording to a clean episode.
When a paid service makes sense
Web transcription services such as Otter.ai and Descript add what Whisper in Audacity lacks. They name the speakers, tie a text editor to the audio, and run faster because the work happens on their servers. The cost model is the trade. A free tier gives a set number of minutes per month, after which you pay per month or per hour of audio, and every file is uploaded. For a confidential interview, an unreleased episode or a hundred hours of archive, Whisper on your own machine costs nothing and the audio never leaves it. For a weekly show with three guests who need to be named, a service may pay for itself in editing time.
FAQ
Next steps
The other four AI effects in the same plugin, including the Noise Suppression used above, are in how to edit audio with AI in Audacity. For the episode itself, how to edit a podcast in Audacity takes the recording from import to export, and the plugin section has more voice tools for Audacity. Get Audacity 3 and OpenVINO AI Tools in one place — download MuseHub free.