Manual

Voice tracking

How TalkPatch follows your voice, which models there are and how you tune the reading line.

How it works

TalkPatch listens through the selected microphone, recognizes the words on your Mac and matches them against the prompter text. Case and accents don’t matter, and compound words are handled. The moment you move on, the reading line follows. Skipped or reworded passages are tolerated in a window of roughly eight words back and sixty forward. With no match, the text waits, so you can wander off and step back in whenever you like.

Chapter headings and stage directions stay out of it: their words are never expected and never marked.

Models

Pick one under voice tracking chip → arrow menu → Recognition (the same menu opens Voice tracking options …). Every model runs on your Mac. The first time you switch one on it downloads to ~/Library/Application Support/FluidAudio/Models/ and stays there.

Model Languages Type Note
Nemotron multilingual 0.6B German, English streaming, steps 560/1120/2240 ms the default. 1120 ms balances speed against accuracy
Nemotron English 0.6B English streaming
Parakeet TDT v3 (sliding window) German, English + 23 sliding window
Parakeet TDT v2 / TDT-CTC 110M English sliding window light, fast
Parakeet EOU 120M English streaming, 160/320/1280 ms very low latency
Parakeet Unified 0.6B English streaming
Cohere Transcribe (beta) German, English + 12 block 1.8 GB, high quality, higher latency
Apple SpeechAnalyzer / classic German, English Apple needs the speech recognition permission

The step of a streaming model is the length of the audio blocks: shorter reacts faster, longer recognizes more reliably.

Tuning the reading line

Under voice tracking chip → arrow menu → Voice tracking options …:

  • Marking: the current line, the single word, the progress so far, a combination of those or nothing at all, with color and opacity.
  • Lead-in in words and where scrolling starts in the line: how far ahead of your voice the reading line sits.
  • Smoothing: how softly the text follows.
  • Microphone and level (real dBFS) sit on the studio bar, speech sensitivity in the gear beside it.
  • Countdown before the start can be switched off.

You can also drag the reading line itself at the edge of the stage.

Lines and jumps

↑/↓ move one real line back or forward, even with voice tracking running: TalkPatch jumps at once and re-aligns recognition to the new spot. The time left in the chip starts from 132 words per minute and, after the first sentences, keeps recalculating with your measured pace, like a sat nav. The estimated total also shows in the script list and the editor.