If you have asked an untrained AI tool to caption a product video, you have already seen the problem: radiant, glowing, effortless, game-changing. Every brand gets the same four adjectives because the tool has nothing else to go on. Feed it 20 of your real captions first and the output starts to sound like the person who runs your account.
The gap between generic and on-brand
Take a serum launch. The generic AI caption reads: "Unlock radiant skin with our game-changing serum!" Any skincare account launching this week could post it. The trained caption reads: "ten days in. skin's calmer than it's been all year." Lowercase, no exclamation point, one specific number.
Both lines describe the same result. The second one gets a second look because a follower recognizes the voice before the sentence ends.
Brand voice is a spec
On most teams the voice lives in one person's head, which is why briefs to freelancers, agencies, and AI tools all come back wrong. Write it down as five variables anyone can execute against:
- Word bank: the 10-15 words you actually use ("calmer," "routine," "results") and the words you ban ("unlock," "elevate," "game-changing").
- Sentence length: short and clipped, or long and conversational.
- Capitalization: lowercase and casual, or full sentences and proper case.
- Emoji policy: none, one per caption, or a heavy hand.
- Joke frequency: dry humor every third caption, or never.
A new hire or an AI tool can follow that on day one. Without it, both spend six months reverse-engineering your old posts.
What untrained AI gets wrong beyond the words
Word choice is the loudest tell, but three quieter patterns expose untrained AI just as fast:
- Caption length. If your brand posts two clipped lines, a mid-length paragraph looks wrong. Untrained AI writes that mid-length paragraph every time.
- CTA style. Maybe you always close with a question, or a product detail, or nothing. Untrained AI closes with a generic ask.
- Hashtag habits. Some brands run 8-10 tags, some run zero. Untrained AI averages toward 3-5 generic tags, wrong in both directions.
Any single post can pass; the pattern shows across ten or twenty, which is why the test step below matters as much as the training step.
Sameness is the actual cost
Strip the product name out of most AI-written captions and you cannot tell which account posted them: "Join the glow up." "You deserve this."
Two brands in adjacent categories can and should sound nothing alike. Picture a beauty brand whose captions read quiet and results-first next to a fashion accessories brand that reads playful and visual. Swap their captions and both accounts read as off, even with similar products, price points, and platforms.