Brand Voice in the AI Era

How to Train AI on Your Brand Voice (Without Fine-Tuning)

A practical method to train AI on brand voice using curated examples, an operational style guide, and a feedback loop, no fine-tuning required.

You don't need a custom model to make a general AI model sound like your brand. You need to give it the right context at the moment you ask it to write. That means a handful of curated examples, a style guide it can actually apply, task-specific reference material, and a feedback loop that tightens the output over time. Do that consistently and a general model gets close enough to your best writer's voice that the gap stops mattering. This is the practical follow-up to Brand Voice in the AI Era, which covers what voice is and why it's worth protecting. Here, we're focused on the mechanics of how to train AI on brand voice using context instead of retraining. The approach below comes out of running this setup with content teams at several client companies, mostly SaaS and fintech, who write daily under a shared style and needed something that held up across dozens of writers rather than one careful editor.

Start with a weak draft

Say you run a mid-size SaaS company and you ask a general model to write a product update announcement. No examples, no guide, just the request. You'll get something like this:

"We're thrilled to announce an exciting new update to our platform! This game-changing feature empowers users to seamlessly streamline their workflows and unlock next-level productivity. Our team has worked tirelessly to deliver this cutting-edge solution."

That paragraph could belong to almost any company. It's not wrong, exactly. It's just generic in a way that erases the brand rather than expressing it. The vocabulary is marketing filler ("game-changing," "seamlessly," "unlock"), the tone is performed enthusiasm rather than earned confidence, and there's nothing here a reader could point to and say "only this company writes like that."

Now give the same model three things: two short passages of on-voice writing, a rule that bans hype adjectives and requires plain verbs, and a one-line brief on what the feature actually does. Ask again:

"We shipped a way to reconcile invoices automatically overnight, so your team stops doing it by hand on Monday mornings. It runs on the same schedule as your existing sync, no setup required."

Same underlying announcement. Completely different register. Nothing was retrained. The model didn't learn your voice in any lasting sense, it was simply given enough of the right material to imitate it convincingly for this one task. That's the entire premise of context-based voice work, and it's why the rest of this piece is about what to feed the model, not how to change it.

What makes voice portable

Voice becomes portable to a model when it's shown as pattern rather than described as personality. Telling a model to "sound confident but approachable" gives it an adjective to interpret however it likes. Showing it three paragraphs that are confident and approachable gives it something concrete to copy, and models are far better at pattern-matching than at executing abstract instructions.

The inputs that make this work reliably fall into a few categories. A small, curated set of on-voice examples matters more than a large mixed one. Five to ten short passages is the range that shows up consistently across teams doing this well. We arrived at that range through our own informal testing across product announcements, support replies, and blog intros for several client accounts, comparing editor-flagged voice drift against the size and mix of the example set each team used. It's not a controlled study, more a pattern that held across a handful of engagements, but it's held consistently enough to treat as a starting point rather than a guess. That range covers a few different formats, so a model learns that voice holds steady even as tone shifts by format. One off-voice sample in that set will drag the average down, so the curation matters as much as the collection.

An operational style guide does different work than the examples. Most brand guides are written for humans and read like mood boards: values, adjectives, vague aspirations. A model needs sentence-level rules it can apply mechanically, point of view, preferred and banned terms, and do/don't pairs that contrast an on-voice phrase against an off-voice one. "Say we built it, not our team leveraged a solution" teaches faster than any adjective could. For a fuller walkthrough of building one, see Defining Your Brand Voice, and for the boundaries that keep output from drifting, Brand Voice Guardrails.

Reference snippets are the third input, and they're task-specific rather than standing. Writing about a particular feature works better when the model has your past description of that feature in front of it, or a paragraph from a related published piece. Examples teach the general sound. Snippets anchor the model to your actual phrasing on the actual subject, which is what stops it from drifting into generic territory the moment the topic gets unfamiliar.

Context beats retraining

Fine-tuning means adjusting a model's underlying weights on your data. It sounds like the obvious answer to a voice problem. It's the wrong default for nearly every brand, and the reason has nothing to do with cost.

Voice isn't fixed. It sharpens as a company matures, flexes for a new product line, gets corrected after a campaign lands wrong. A fine-tuned model, whether that's an OpenAI fine-tune, a LoRA adapter on an open-weights model like Llama, or a hosted fine-tuning job through a platform like Anthropic's, freezes a snapshot of a moment that's already passed. Changing it means gathering new training data, running a job, evaluating the result, and redeploying, which is slow and expensive for a problem that context can solve in seconds.

Context-based methods stay editable. Change a line in the style guide and the next draft reflects it immediately. You can see exactly what's shaping the output at any given moment, which means you can audit why a model wrote what it wrote and adjust the specific input responsible. When a fine-tuned model produces an off-voice sentence, tracing the cause means guessing at what training data might have caused it. When a context-fed model does the same thing, you can usually point to the exact example or rule that misled it, and fix that one thing without touching anything else.

A concrete case, drawn from a support-reply project we ran for a mid-size fintech client over two weeks in the same quarter: their generator kept producing lines like "we sincerely apologize for any inconvenience this may have caused," a phrase their brand guide explicitly banned. Rather than retrain anything, we swapped in three reply examples that used their preferred pattern ("that's on us, here's the fix") and added the banned phrase to the do/don't pairs. An editor reviewed a sample batch of fifty generated replies each week, flagging any draft that needed a voice correction before it shipped. Off-voice phrasing dropped from roughly one in four replies to one in twenty over that two-week window. That's the kind of fix a fine-tuned model would have needed a full retraining pass to absorb, and even then, only after the pattern had already shipped in production for a while.

There's a genuine tradeoff worth naming. Context-based prompting asks something of the person running it every time, someone has to assemble the examples, keep the guide current, and paste in the right snippets for the task at hand. Fine-tuning, once done, removes that per-task effort. For a brand with one fixed voice and no plans to change it, that tradeoff might land differently. For most commercial teams, whose voice shifts as the product and positioning shift, the ongoing assembly cost is smaller than the cost of retraining every time something changes.

This is also where prompt engineering and brand voice work overlap more than people expect. The context bundle described above, examples, rules, snippets, is a system-level instruction set in everything but name. Teams that already think in terms of system prompts and few-shot design tend to pick this up faster, because it's the same discipline applied to a narrower problem: getting a model to sound like one specific writer instead of a generic assistant.

A repeatable setup

None of this works as a one-time setup. It works as a loop that gets tighter the more you run it. A reusable prompt pattern is what makes the loop repeatable rather than reinvented each time.

Here's what that pattern looks like assembled into one block, using the same fintech support-reply case from above:

ROLE: You are writing customer support replies for [Brand], a fintech
company. Voice is direct, warm, and takes ownership of mistakes without
over-apologizing.

EXAMPLES (match this pattern exactly):
1. "That's on us. We've already refunded the fee and flagged the bug
 for the team."
2. "You're right, that shouldn't have happened. Here's what we're
 doing to fix it."

TASK: Write a reply to a customer whose transfer was delayed by two
days due to a bank holiday we didn't flag on the confirmation screen.

SNIPPETS: [paste last month's delay-notice copy, and the current
help-center article on transfer timing]

CONSTRAINTS: No "we sincerely apologize," no "any inconvenience,"
no passive voice for the mistake. Under 80 words. End with a concrete
next step, not a general reassurance.

That's the full context bundle in one place, a role and voice statement, two on-voice examples, the actual task, the reference snippets, and the hard constraints. Running the same shape every time is what makes the output predictable enough to trust without a full review pass on each draft.

Few-shot examples, the two or three on-voice passages shown before the model produces its own, do more work than any other part of the pattern. If a team changes exactly one thing about how it prompts for voice, this is the change worth making. Concrete examples consistently beat abstract description, because the model has something to match against instead of something to interpret.

Two related techniques extend this further once the basic pattern is working. Retrieval-augmented generation automates the snippet-selection step, pulling the most relevant past examples or reference documents into the prompt automatically instead of someone choosing them by hand each time. That matters more as the volume of content grows and no single person can hold every past piece in their head. A versioned prompt library does the same for the pattern itself, storing each iteration of the role statement, examples, and constraints so a team can compare what changed between a version that worked and one that drifted. Alongside these, a simple evaluation rubric, three or four sentence-level criteria an editor checks against, turns "does this sound on-voice" from a gut call into something closer to a repeatable measurement, which is what makes the feedback loop below worth trusting.

A few operational details separate teams that keep this working from teams that let it decay. Version the prompt pattern itself, not just the style guide, so when a change makes output worse, there's a previous version to revert to rather than a guess at what changed. Route generated drafts through the same review step every time, one editor with the authority to reject or fix, rather than leaving it to whoever's free. And agree in advance on what "on-voice" means well enough to check against, the rubric above, applied in under a minute, rather than a gut feeling that varies by reviewer.

The feedback loop closes the system. Generate a draft from the pattern, have someone edit it to genuinely on-voice, and save the before-and-after rather than discarding it. When the same correction shows up more than once, say a model repeatedly reaching for "utilize" where the brand says "use", that's not a one-off fix. It's a rule that belongs in the style guide and possibly a term that belongs on the banned list. Fold recurring corrections back into the guide and the example set on some regular cadence, weekly or monthly depending on volume, and the gap between what the model produces and what the brand actually sounds like keeps shrinking.

There's a quieter payoff to keeping voice this legible. As more discovery moves through AI answer engines rather than search results, content with a recognizable, consistently structured voice is easier for those systems to attribute to a single coherent source. The clarity that makes voice legible to a human editor is often the same clarity that makes it easier to cite. That connection is worth exploring further in Generative Engine Optimization if visibility in AI-driven search is part of the strategy.

Less work, more on-brand content

Austen runs this whole workflow for you: from research to on-brand drafts that get found by Google and AI.

Start free

More in Brand Voice in the AI Era