Voice guidelines written as adjectives produce copy that matches the adjectives and sounds like nobody. Voice built from transcripts sounds like the person.
Most voice guides are a list of words to avoid. That removes the obvious tells and leaves generic competent prose, which is not the same as sounding like a specific person.
Working from recordings gives you the positive half: the structures they actually use, the way they qualify a claim, what they say instead of the word they avoid, the rhythm of how they build to a point. That is what makes a caption recognisable.
A verbal tic that lands warmly in conversation reads as a stumble in text. A digression that works when you can hear the person's tone becomes confusing on a page. So each pattern is marked as spoken-only or as carrying into writing, and only the second set is used.
You rarely invoke this directly. Every post-type skill loads it first, which is why a caption, a newsletter and a case study post all sound like the same person despite being written by different workflows.
Skills compound. These are the ones we usually install alongside it.