Text from Narration

A narration becomes the words on screen, each one lit as it is read.

‹ ClapperBoard Help

Give ClapperBoard a recording of someone reading and it can make the pictures itself: it transcribes the narration and fills the image lane with its text, timed to the voice, with the word being read highlighted as it plays. You choose how the text looks and moves. The book's own text, when you have it, corrects what the transcription heard.

What it needs

A narration on an audio lane. Speech, in a language your Mac can transcribe. Tools › Text from Narration is available as soon as an audio lane holds a recording.

Pages are optional. Without them, the words are the transcription's. With the book's pages on the image lane, or a PDF of it, the words take the book's spelling and punctuation — see correcting the words with the book.

The transcription runs on your Mac. Nothing is uploaded.

Creating the text

Choose Tools › Text from Narration, or Text from Narration in an audio item's context menu, which works on that item's lane. Every choice is in the one sheet, beside a preview that plays your choices on a sample of the narration. Create Text then transcribes the narration, reads the book if you asked it to, and lays the text out, under a progress bar you can cancel.

The text replaces what was on the image lane, as one undoable step: Undo brings it all back. To replace only the pages with text, keep the pictures.

Transcripts are kept. A second run — to try another layout, say — goes straight to the layout, and an audiobook split into chapters is transcribed once, however many items it has become.

The narration's language

Transcribing in the wrong language gives nonsense rather than an error, so the sheet shows the language it will use and says where the choice came from: the project, the book's text, the audio file, your last choice, or your Mac's language, in that order. Change it if it is wrong. The project remembers it, and Sync to Audio uses the same one.

Screens or scrolling

Screens show the text a screenful at a time. Each screen changes to the next halfway through the pause between them, so it arrives just before its first word is read and stays until just after its last. Each screen holds as many sentences as fit, or a set number of sentences, with the type shrunk when they would not fit.

Scrolling text runs as one column that glides up, keeping the words being read near the same line.

Alignment sets the text left, centred or justified.

The spoken word

Four looks mark the word being read, alone or together:

Movement decides how the mark travels: Jump to each word, or Slide from one to the next.

Type and colours

Font, face and size, margins and line spacing set the type. Four colours set the text, the background, the marker and the highlight. The size is relative to a 1080-line frame, so the text sits the same way at every export size.

Correcting the words with the book

The transcription hears the words, but not always how they are written: names, old spellings, punctuation, a word misheard. Correct with puts that right from the book's own text — the pages on the visual lane, or a PDF file chosen just for its text. A chosen file is only read, never shown or added to the project.

Where the book agrees with the narration, the words take its spelling and punctuation. Where the transcription misheard a word, the book's replaces it, but only when it looks like what was heard: a scan is often another edition than the one being read, and its different wording never overwrites what the narrator says. What the narrator says that the book lacks — a recording's introduction, an announcement — stays. What the book prints that nobody reads — front matter, captions, footnotes — is left out.

When the book matches nothing — the wrong file, or the wrong language — ClapperBoard says so, and the words stay as they were heard.

Words the dictionary does not know

A scan's text has errors of its own: a smudge read as "Sa", "the" read as "ths". So a book word only replaces what was heard when the dictionary knows it, or when the narrator says it somewhere. Accept words the dictionary doesn't know also lets in words the book prints more than once — invented words, rare names — and with them a noisy scan's repeated errors.

Keeping the pictures

With the book's pages on the image lane, Keep the pictures replaces only the pages with text. Pages with little or no text — illustrations, plates, a cover — stay, each where it falls in the book: after the text before it, at the end of a sentence. Blank pages go, and so does anything else on the lane.

Whether a page counts as text is judged against the book's typical page, so a picture book whose every picture carries a verse keeps the verses as text.

Each picture takes its time from the start of the text after it, up to the length set in the sheet and never more than half of that screen. The words read meanwhile show as read when the text returns.

The pictures before the text — a cover, a frontispiece — open the video, and those after it close it, as long as the narration reaches that end of the book. A narration of one chapter, with the whole book on the lane, does not bring in the rest of the book's pictures.

How the text stays in time

Every word is tied to the moment it is heard in the recording, not to its text item. Move or trim the narration and the highlight follows. Move a text item away from its narration and it shows its words without a highlight. Split a text item and its words are divided at the cut.

Sync to Audio is unavailable while the image lane holds text: it retimes pages, and the text keeps its own time.

Changing the style later

Open Text from Narration again to change how the text looks. Apply Style keeps the words, edits included. Colours and the highlight change in place; a new font, size, spacing or paging lays the words out afresh in the time the text already had, and leaves pictures and other items where they are.

Regenerate from Narration starts over: it transcribes again and replaces the text, edits included, keeping the pictures among it if you ask it to. Undo brings the old text back.

Editing the words

Choose Edit Text in a text item's context menu, or double-click the item, to correct its words by hand. Words you keep stay timed where the narrator says them, even with new capitals or punctuation. A word you change takes the time of the one it replaces, and a new word is fitted into the pause before the next. ⌘Return saves.

Exporting

Text exports like everything else, in videos and GIFs. A sliding or scrolling highlight moves on every frame, so while it moves the export renders frame by frame, and a long narration takes noticeably longer than with a jumping highlight. Final Cut Pro receives each text item as a still, without the highlight.

Last updated September 26, 2026