Part 4 of Marketing in the Age of AI. Start with part 1 at /insights/articles/marketing-value-chain/ or the series overview at /insights/articles/marketing-in-the-age-of-ai/.
Ask any capable model for ten campaign concepts, a positioning statement, a segmentation, a set of email subject lines. You will get them in seconds, and they will read well. Everyone noticed that speed first, though it matters least.
The hard part of applied AI in marketing is knowing whether the output is any good, in a domain where the ground truth is fuzzy and the wrong answer looks exactly as polished as the right one. A model will hand you a confident, well-written positioning line that a competitor could say word for word. It will give you a segmentation with tidy labels and no decision inside it. Generation is free; telling the good from the plausible-but-wrong is the job now.
The thing that does that telling has a name worth borrowing from the people who build these systems: the evaluation function. It is the encoded domain judgement that separates a good output from a plausible wrong one. A model does not have yours unless you give it to it.
The bottleneck moved from making to judging
For most of marketing's history, making the thing was the expensive step. The hours and the budget went into writing the deck, shooting the creative, running the study, and drafting the messaging. Judgement was cheap by comparison, because it was carried around in a few experienced heads and applied at the end.
AI inverts that. The making is now close to free and close to instant. The judgement about whether what got made is right has not changed and cannot be bought off the shelf. The scarce, valuable, defensible thing has moved from production capacity to the ability to look at a fluent output and say "no, that is slop, and here is exactly why."
Most teams are still spending their AI energy on the generation half: better prompts, more variations, faster drafts. The teams pulling ahead are spending it on the judgement half: encoding what "good" means for their category so the machine can be held to it. That encoded judgement is the asset. Everything upstream of it is a commodity.
Defining the evaluation function
The distinction between a prompt and a skill separates a party trick from a system you can trust.
A prompt is judgement you re-type and re-trust every time. You know what good looks like, so you steer the model toward it in the moment: a few instructions, a bit of context, some corrections when it drifts. It works, but it is invisible, unversioned, and inconsistent. The quality lives in your head and gets re-applied by hand on every run. Two people prompting the same task get two different standards. You, on a tired Friday, get a different standard from you on a fresh Monday.
A skill is that same judgement written down once: version-controlled, inspectable, reusable. The standard stops living in your head and starts living in a document the tool consults every time, the same way each time, whoever runs it. When your judgement improves, you improve the document, and every future run inherits the upgrade.
The single most valuable behaviour you can put into that document is a gate: a point where the tool stops and says "I can't proceed - I don't yet have X" instead of producing confident output anyway. A positioning tool that refuses to write a headline until it can name the customer it is for, the competitor it beats, and the demand it captures is doing exactly its job. The refusal is the judgement made visible. A tool that always answers, no matter how thin the input, is only a generator.
The shift is to move your quality bar out of your head and into something that can refuse.
Encoding an evaluation function
You do not need to be an engineer to do this. The recipe is five steps, and it is the same whether you are encoding it into an AI skill or writing a checklist your team applies by hand.
1. Name the failure mode. Describe what "plausible but wrong" looks like for this specific output: a positioning line that any rival could also claim, a segmentation whose segments do not change any decision, a campaign concept that meets the brief but addresses a problem nobody has. Be concrete. You cannot catch a failure you have not described.
2. Write the test that catches it. Turn the failure into the conditions a good output must satisfy, and make them mechanical enough to check. "Distinctive" cannot be checked; "a claim a named competitor could not truthfully make about themselves" can. "Actionable segmentation" cannot be checked; "each segment implies a different budget, message, or channel" can. The more mechanical the condition, the more reliably a person or a model can apply it.
3. Add a gate. Make the tool, or the process, refuse to proceed until the test passes, or at minimum label its own confidence and say what is missing. This is the step most people skip, and it is the one that does the work. An output that arrives with "I proceeded but I could not verify the competitor claim" is worth ten confident outputs that quietly guessed.
4. Keep the reasoning. Have it show its working. A positioning line that arrives with the argument for why it passes the distinctiveness test is auditable. A human can read the reasoning and overrule it. A bare answer can only be accepted or rejected on faith, which means the judgement never got transferred.
5. Version it. When your judgement improves, with a new failure mode you learned the hard way or a sharper test, update the document. Every future run is then better automatically, and the improvement compounds. A checklist you improve twenty times over a year is a different instrument from the one you started with, and none of that learning leaked away between projects.
You already ran evaluation functions before AI
None of this is new. The best marketers have always carried evaluation functions around - they were never written down, so they could not be handed to anyone or anything else. Making them explicit fixes that. The pre-AI version of an explicit evaluation function is the weighted driver model: name the drivers, name the weights, name the rules.
One real example comes from a global professional-services brand study. The question was which factors move a firm from merely known to a client's preferred partner. The drivers that carried weight were strong relationships, at roughly a quarter of the total, and the felt quality of the client experience, and genuine distinctiveness, meaning a claim a rival could not truthfully make. Awareness and clever taglines carried close to zero weight; they do not drive preference.
Use that as a scoring function. When a model hands you a B2B positioning, score it against those weights. High on relationships, experience, and distinctiveness means defensible; high on awareness and wordplay means slop, however nicely it reads. That single scorecard catches the most common way AI positioning fails, which is to sound memorable while being something anyone in the category could say. (I go deeper on this in part 3 of the series.)
The same shape works for choose-where-to-play decisions. A real prioritisation exercise scoring eighty markets weighted competitive intensity at 40 percent, route to market at 20, role of brand at 20, country risk at 10, and right to play at 10. Score each market, weight, sum to one comparable number, then plot value-of-the-prize against probability-of-success and tier the results into must-win, strongly-consider, and lower-priority. The output is a decision you can defend line by line.
The point is the discipline: explicit drivers, explicit weights, explicit rules. Swap the drivers for whatever you are deciding, whether feature, account, channel, or market, and the method transfers intact. An AI cannot supply that discipline on its own; once you supply it, the output becomes trustworthy.
Implications for marketing teams
Prompts are a commodity and getting cheaper. The teams that win the next few years will be the ones who have done the unglamorous work of writing down what good looks like in their category, in tests concrete enough that a machine can be held to them, with gates that make the machine flag a thin brief instead of dressing it up.
Start with your highest-stakes, most-repeated output: the positioning, the segmentation, the quarterly plan. Name how it fails, write the test, and add the gate. You will get a tool that is honest about what it does not know, which is worth far more than a tool that is fluent about everything.
If you want a running start, the free Marketing Skills Pack packages several of these evaluation functions as ready-to-use skills, including the positioning gate described above, so you can see what encoded judgement looks like before you write your own.