Marketing in the Age of AI · AI Strategy

The Marketing Evaluation Function

Generation got cheap almost overnight. The scarce skill now is judging whether an output is any good, and that judgement is something you can write down and make a machine apply.

Anton Dudarenko · 8 min read · 16 July 2026
TL;DR Generation is cheap now. The scarce advantage is the evaluation function: written-down domain judgement that tells a good output from a plausible wrong one.
  • The bottleneck in marketing moved from making the work to judging whether the work is any good, and a model does not hold your judgement unless you supply that judgement yourself.
  • A prompt is judgement you re-type every run; a skill is that judgement written down once and applied the same way whoever runs it.
  • The key behaviour is a gate: a point where the tool stops and names what it is missing before it will produce an answer.
  • Encoding an evaluation function takes five steps: name the failure mode, write the test that catches it, add a gate, keep the reasoning, and version it.
  • The pre-AI version is the weighted driver model, which states drivers, weights and scoring rules in the open and applies to any decision you are scoring.

Part 4 of Marketing in the Age of AI. Start with part 1 at /insights/articles/marketing-value-chain/ or the series overview at /insights/articles/marketing-in-the-age-of-ai/.

Any capable model will produce ten campaign concepts, a positioning statement, a segmentation, or a set of email subject lines on request. You will get them in seconds, and they will read well. That speed is the least important of the changes.

The difficulty in applied AI in marketing is knowing whether the output is any good, in a domain where the ground truth is fuzzy and the wrong answer looks exactly as polished as the right one. A model will hand you a confident, well-written positioning line that a competitor could say word for word. It will give you a segmentation with tidy labels and no decision inside it. Generation is cheap, so the work is telling a good output from a plausible wrong one.

The evaluation function is the name machine-learning engineers give to that judgement. It is the encoded domain judgement that separates a good output from a plausible wrong one.

Judgement as the scarce step

For most of marketing's history, making the thing was the expensive step. The hours and the budget went into writing the deck, shooting the creative, running the study, and drafting the messaging. Judgement was the smaller cost, because a few experienced people held it and applied it at the end.

AI inverts that. The making is now close to free and close to instant. The judgement about whether what got made is right has not changed, and no one sells it. The scarce input has moved from production capacity to the ability to explain why a fluent output is wrong.

The better use of AI effort is on the judgement half: encoding what "good" means for the category so the machine can be held to it. That encoded judgement is the asset; the generation work before it is a commodity.

Writing the evaluation function down as a skill

The distinction between a prompt and a skill separates a one-off result from a system you can trust.

A prompt is judgement you re-type and re-trust every time. You know what good looks like, so you steer the model toward it in the moment: a few instructions, a bit of context, some corrections when it drifts. It works, but it is invisible, unversioned, and inconsistent. You hold the standard yourself and re-apply it by hand every time you prompt. An unwritten standard varies by person and by day.

A skill is that same judgement written down once: version-controlled, inspectable, reusable. The standard moves from your head into a document the tool reads on every run, the same way each time, whoever runs it. When your judgement improves, you improve the document, and every future run inherits the upgrade.

The most useful behaviour you can put into that document is a gate: a point where the tool stops and names the input it is missing instead of producing confident output anyway. A positioning tool that refuses to write a headline until it can name the customer it is for, the competitor it beats, and the demand it captures is doing exactly its job. The refusal is the judgement applied in practice. A tool that always answers, no matter how thin the input, is only a generator.

Encoding an evaluation function

The method needs no engineering background. It is the same whether you are encoding it into an AI skill or writing a checklist your team applies by hand.

1. Name the failure mode. Describe what "plausible but wrong" looks like for this specific output. The positioning and segmentation examples above show it; a campaign concept that meets the brief but addresses a problem nobody has is another. Describe the failure in concrete terms.

2. Write the test that catches it. Write the failure as the conditions a good output must satisfy, and make them mechanical enough to check. "Distinctive" cannot be checked, but "a claim a named competitor could not truthfully make about themselves" can; part 3 of the series applies that test to B2B positioning. "Actionable segmentation" cannot be checked; "each segment implies a different budget, message, or channel" can. The more mechanical the condition, the more reliably a person or a model can apply it.

3. Add a gate. Make the tool, or the process, refuse to proceed until the test passes, or at minimum label its own confidence and say what is missing. The gate is the step that catches the failure. An output that labels its own gaps is more useful than a confident output that quietly guessed.

4. Keep the reasoning. Have it show its working. A positioning line that arrives with the argument for why it passes the distinctiveness test is auditable. A human can read the reasoning and overrule it. A bare answer can only be accepted or rejected on faith, which means the judgement never got transferred.

5. Version it. Update the document when your judgement improves, through a new failure mode you have met or a sharper test. A checklist you improve again and again over a year is a different document from the one you started with, and none of that learning is lost between projects.

Weighted driver models as evaluation functions

None of this is new to experienced marketers, who have always applied evaluation functions. Those functions were never written down, so they could not be handed to anyone or anything else. Making them explicit fixes that. The pre-AI version of an explicit evaluation function is the weighted driver model. In brand analytics the same idea takes the form of a brand equity path model, which weights each perception driver by its effect on a commercial outcome, and a brand equity valuation that uses those weights to value the brand in money.

The method also works for market prioritisation decisions. Suppose a prioritisation across eighty markets that weights competitive intensity at 40 percent, route to market at 20, role of brand at 20, country risk at 10 and right to play at 10. Score each market on each driver, weight and sum to one number for probability of success, plot it against the size of the prize in each market, and tier the results into must-win, strongly consider and lower priority. The output is a decision you can defend line by line.

Swap the drivers for whatever you are deciding, whether feature, account, channel, or market, and the method transfers intact. An AI cannot supply the drivers and weights on its own; once you supply them, each output can be checked against them.

Implications for marketing teams

The durable work is writing down what good looks like in your category, in tests concrete enough that a machine can be held to them, with gates that make the machine flag a thin brief.

Start with your highest-stakes, most-repeated output: the positioning, the segmentation, the quarterly plan. Name how it fails, write the test, and add the gate. The result is a tool that states what it does not know.

The free Marketing Skills Pack includes the positioning gate described above as a ready-to-use skill, so you can see what encoded judgement looks like before you write your own.