Voice tracking
How TalkPatch follows your voice, which models there are and how you tune the reading line.
How it works
TalkPatch listens through the selected microphone, recognizes the words on your Mac and matches them against the prompter text. Case and accents don’t matter, and compound words are handled. The moment you move on, the reading line follows. Skipped or reworded passages are tolerated in a window of roughly eight words back and sixty forward. With no match, the text waits, so you can wander off and step back in whenever you like.
Chapter headings and stage directions stay out of it: their words are never expected and never marked.
Models
Pick one under voice tracking chip → arrow menu → Recognition (the same menu opens Voice tracking options …). Every model runs on your Mac. The first time you switch one on it downloads to ~/Library/Application Support/FluidAudio/Models/ and stays there.
| Model | Languages | Type | Note |
|---|---|---|---|
| Nemotron multilingual 0.6B | German, English | streaming, steps 560/1120/2240 ms | the default. 1120 ms balances speed against accuracy |
| Nemotron English 0.6B | English | streaming | |
| Parakeet TDT v3 (sliding window) | German, English + 23 | sliding window | |
| Parakeet TDT v2 / TDT-CTC 110M | English | sliding window | light, fast |
| Parakeet EOU 120M | English | streaming, 160/320/1280 ms | very low latency |
| Parakeet Unified 0.6B | English | streaming | |
| Cohere Transcribe (beta) | German, English + 12 | block | 1.8 GB, high quality, higher latency |
| Apple SpeechAnalyzer / classic | German, English | Apple | needs the speech recognition permission |
The step of a streaming model is the length of the audio blocks: shorter reacts faster, longer recognizes more reliably.
Tuning the reading line
Under voice tracking chip → arrow menu → Voice tracking options …:
- Marking: the current line, the single word, the progress so far, a combination of those or nothing at all, with color and opacity.
- Lead-in in words and where scrolling starts in the line: how far ahead of your voice the reading line sits.
- Smoothing: how softly the text follows.
- Microphone and level (real dBFS) sit on the studio bar, speech sensitivity in the gear beside it.
- Countdown before the start can be switched off.
You can also drag the reading line itself at the edge of the stage.
Lines and jumps
↑/↓ move one real line back or forward, even with voice tracking running: TalkPatch jumps at once and re-aligns recognition to the new spot. The time left in the chip starts from 132 words per minute and, after the first sentences, keeps recalculating with your measured pace, like a sat nav. The estimated total also shows in the script list and the editor.