Homemade Pasta · Research
The ad is the largest factor in paid media return that an advertiser controls. In most accounts it is also the input given the least expertise: a folder of old assets handed to the media team, or a request for the media team to make the ads themselves.
Why this exists
Homemade Pasta runs paid media and has its own creative team, and we keep seeing the same thing at the start of a campaign. The media plan gets weeks of work. The creative arrives as a folder of old assets. Or the media team is asked to make the ads.
Media buyers are good at a specific job: deciding who sees an ad, when, how often and at what price, then measuring what happened. Writing and designing an ad is a different craft with its own training. When the brief skips it, the account ends up tuning bids, audiences and budgets around ads nobody built for the job.
We wanted to know what that costs, so we read the evidence on what predicts return from paid advertising. This page covers paid advertising only.
Executive summary
Three findings, in the order they change what you should do.
Recommendations
Each change belongs to one part of the team. The split is the point: the people buying the media should not also be writing the ads.
The ad is the largest controllable factor, and its content matters more than its execution. Choosing that content is a creative craft that media buyers are not trained in.
Frees the media team for the delivery and measurement work it is good at.
In the weight tests among 389 randomized TV experiments, extra weight alone did not broadly lift sales. Changes to brand, copy and media strategy made a sales effect more likely.
Moves effort to the input that caps every bid, audience and budget decision downstream.
Across 2,251 ads from 91 brands, content showing how the product delivers its benefit raised elasticity most, ahead of factual claims and then subjective or emotional ones.
Starts every brief from the content type with the strongest evidence behind it.
In lab studies of low-involvement products, repeating one ad lowered product attitudes while varied executions held them. One model of an online campaign estimated that varying creative by each person’s impression history would lift conversions 13.8%.
Extends the working life of an idea without paying the repetition penalty.
Measured on clicks across 30 campaigns on one publisher, wear-out arrived after one or two exposures for some and showed little sign within fifty for others.
Sets refresh timing from your own account instead of the fiscal calendar.
Platform delivery sends each variant to a different mix of people, so an A/B winner can win on audience alone. That is fine for choosing an ad on that platform and unreliable for judging an idea.
Separates which ad performs on a platform from which idea moves sales.
Phasing, seasonality and audience refinement rank low on every list. They still need a correct setting, because a broken default can cost more than the factor is worth.
Frees the hours spent on weekly adjustments for the brief and the testing plan.
Part One
What the evidence says drives return from paid advertising, who should be making the ads, and why the folder of existing assets underperforms.
1.1 The ad
The cleanest evidence on this question is old, randomized and large. Between 1982 and 1988, 389 split-cable experiments assigned households to see different television advertising and tracked what they bought. In the tests that varied weight, increasing it alone did not broadly lift sales. The authors found that changes to brand, copy and media strategy made a sales effect more likely, in categories with many purchase occasions and little in-store merchandising. The same study found the recall and persuasion scores used to pre-test ads only weakly related to sales.1
The modern equivalent of adding weight is raising a budget or loosening a bid cap. The platforms have changed and the principle carries over: delivery mostly amplifies whatever response an ad produces.
More recent work points the same way and says something about which content works. A study of 2,251 ads from 91 consumer packaged goods brands over four years measured how creative strategy changed advertising elasticity. What the ad said mattered more than how it was executed. Content showing how the product’s features combine to deliver what the buyer wants produced the largest gains, followed by factual claims about verifiable features, then subjective or emotional claims that depend on the viewer’s interpretation. Execution played a supporting role.2 A separate meta-analysis of 878 effect sizes from 67 papers found robust positive effects of advertising creativity, stronger for high-involvement products, on how people respond to ads and brands.3
Creative held at about half in both slices the vendor published: 49% across 359 digital campaigns and 48% across 47 TV campaigns. The method is not published in enough detail to reproduce, so read it as an industry claim.
The largest decomposition we could find comes from a measurement vendor that links household ad exposure to purchase data and compares exposed households with matched unexposed ones. Across roughly 450 consumer packaged goods campaigns on TV and digital, it attributes 49% of measured sales lift to the creative, against 14% for reach and 11% for targeting.4
None of these sources is decisive alone. Together they agree in direction: the ad is the largest single factor in whether advertising moves sales, and it is the one an advertiser can change inside a quarter.
1.2 Who makes it
The strongest result in the elasticity study is also the most useful for deciding who should make an ad. What the ad said moved elasticity more than how it was executed. The best-performing content showed how the product delivers what the buyer wants, which means someone first had to decide which part of the product experience matters to this buyer, and how to show it in the seconds a feed allows.2
That decision is creative strategy. It comes before design, and it is the step that disappears in the two most common shortcuts. A folder of existing assets answers the question of what to say with whatever an earlier brief decided, usually for a different audience or channel. A media buyer asked to make the ads on deadline starts from the assets and templates at hand, because that is all the time the job allows.
Media teams do a different job well. Who sees an ad, how often, at what price, and whether it worked are hard problems, and our incrementality research shows how much judgment the measurement half alone takes. Asking the same people to also write and design the ads splits their time across two crafts and hands the input with the most room in it to the people with the least training for it.
The evidence measures what drives the effect, not who made the ad, so that last step is our reading. Content drives the effect, and choosing content is the work a creative team exists to do.
1.3 The folder
Advertising response falls with repetition to the same person. In a 72-day online campaign covering more than 12,000 users across 473 publishers, about 24% of users showed weariness: past a point, additional exposures made them less likely to visit. For those users it set in at roughly three impressions within two days. The authors’ simulation of profiling users and reallocating or capping impressions raised expected visits by as much as 15%.5
No single frequency cap fits all 30. How fast wear-out arrives is a property of the campaign.
How fast that happens varies by campaign. Across 30 natural experiments on the Yahoo! front page, measured on clicks, four campaigns showed significant wear-out after one or two exposures and ten showed little after as many as fifty. Observational methods overstated wear-out in 26 of the 30.6 In a lab experiment, highly creative ads showed little wear-out over repeated exposures, while ads low on creativity followed the classic pattern of rising and then declining response.7
Variation slows it down. In experiments where the product mattered little to viewers, people shown the same ad repeatedly liked the product significantly less than a control group, while people shown varied executions of the same message held at the control level.8 A model of an online display campaign estimated that choosing each person’s next creative from their impression history would raise visits 12.7% and conversions 13.8%.9 The packaged-goods study in 1.1 reaches a similar recommendation from its own data: focus the content on one dimension, match it with consistent executional elements, and vary the composition of the creative over time.2
An old asset is a problem because of what it usually signals. Anything that has been in market for years has been seen, often many times, by the people you are about to pay to reach again. And a folder assembled for other purposes rarely holds enough distinct executions of one idea to rotate. The long-running campaigns held up as effectiveness case studies, such as Aldi’s Christmas carrot character, which returned with a new story every year from 2016 to 2021, kept a consistent idea and produced new executions of it.10 That is different from re-running the same file.
The folder has a second problem. Much of it was made for another channel, like a TV spot cut down for a feed. The platforms say creative built for their format performs better. They also sell the space, so test it on your own account before taking it on trust.
Part Two
Tactics for large, medium and small advertisers, and what to stop spending time on.
2.1 By stage
Stage here means how established the brand is and how much budget it has to test with. The creative job is the same at every stage. What changes is where a better ad pays back first.
Established brand, large budget.
The useful question is whether the next dollar moves anything. Across 288 established brands on US television, the median answer was barely.11
Growing, with real headroom and a budget large enough to measure.
Response to advertising runs highest in the growth stage, and the budget can support a proper test. A better ad has the most room to pay back here.13
Early stage, small budget.
Response to advertising tends to be strongest early in the life cycle, so a good ad has real room to work, and a weak one burns budget you cannot spare.
The common thread across all three is that the media plan decides how efficiently an ad reaches people, and the ad decides what happens when it does. The stages differ in how much room the ad has to move return, and none of them gets the lift from media alone.
2.2 Priorities
A widely presented ranking of profitability drivers comes from a UK effectiveness consultant. It draws on academic papers, industry reports and case histories, including awards entries, covering some 28,000 campaigns, and it gives each factor a multiplier: the ratio between the best and worst returns observed for that factor.10 It is not a peer-reviewed method, and a best-to-worst ratio is not the gain you should expect from working on a factor. It is still a fair map of where the spread sits.
| Rank | Factor | Best to worst |
|---|---|---|
| 1 | Brand size | 20× |
| 2 | Creative quality | 12× |
| 3 | Budget setting across geographies | 5× |
| 4 | Budget setting across portfolios | 3× |
| 5 | Multimedia | 2.5× |
| 6 | Brand versus performance | 2× |
| 7 | Budget setting across variants | 1.7× |
| 8 | Cost and product seasonality | 1.6× |
| 9 | Laydown and phasing | 1.15× |
| 10 | Target audience | 1.1× |
Brand size tops the list, and nobody can change it this quarter. Creative quality is the largest item an advertiser can change. Its multiplier draws partly on comparisons with effectiveness award winners, who are chosen for their results, so trust the rank more than the number.
Read the bottom half as settings. Seasonality, phasing and audience definition each need a correct default, and past that point the hours spent tuning them are worth more in the brief.
The distinction matters because ignoring a factor means leaving it wherever it lands, and a bad default in a minor factor can cost more than the factor is worth. A campaign dark through its peak season is one example. An audience definition so narrow the ad never reaches new buyers is another. Deprioritizing means setting it once, deliberately, checking it on a schedule, and moving the effort to the work with more room in it.
2.3 Practices
The practices the evidence backs, the ones where the risk sits in how they get read, and the ones that cost more than they return.
Supported by the evidence above.
It sets the ceiling on everything the media plan does afterward. In 389 randomized TV experiments, changes to brand, copy and media strategy made a sales effect more likely in frequently bought categories, while extra weight alone did not broadly lift sales.
For low-involvement products, repeating a single ad lowered product attitudes in lab studies, while varied executions of the same message held them. Rotation is only possible if the executions exist.
In one online campaign, exposures past about three impressions in two days made a quarter of users less likely to visit. On one publisher, ten of 30 campaigns showed little wear-out in clicks within fifty. Your own frequency and response data say which case you are in.
Across 2,251 ads from 91 brands, content showing how the product’s features deliver its benefit raised elasticity most, ahead of factual claims and then subjective or emotional ones.
A holdout measures whether the advertising changed behavior. A platform A/B test tells you which execution performs best inside that platform’s delivery. Both are useful for their own question.
Nothing wrong with the practice. The trap is in what gets concluded from it.
Consistency of idea and brand assets is worth keeping. The trap is reading “consistent” as permission to re-run the same file for years to the same people.
Dynamic creative tools can find combinations that perform on that platform. The trap is reading what they pick as a lesson about the idea, when delivery chose who saw each combination.
Pre-testing catches problems. In the split-cable experiments, recall and persuasion scores were only weakly related to sales, so a high score is not a forecast.
These cost more than they return.
The folder hands the media plan assets your audience has likely seen, made for other channels, with too few executions to rotate. Asking the media team to fill the gap gives the most important input to the people with the least training for it. Every optimization afterward works against that ceiling.
Delivery mostly amplifies the response an ad produces. In the weight tests among 389 randomized TV experiments, weight alone did not broadly lift sales.
The platform shows each variant to a different mix of people, so a winner can win on audience alone. In an audit of 181,890 A/B tests run on Meta, 22% of audience balance checks exceeded a common imbalance threshold, against 0.16% in randomized lift tests.
Across 30 campaigns on one publisher, wear-out in clicks arrived anywhere from the first exposure or two to not at all within fifty. A single cap is too tight for some ads and too loose for others.
Appendix
The limits of the evidence, and every source behind the figures above.
A.1 Limits
This page argues from several kinds of evidence, and they carry different weight. The case for creative rests on randomized experiments, an elasticity study, a meta-analysis and a vendor decomposition that agree in direction and each have gaps.
Almost all of the sales evidence on creative comes from consumer packaged goods, and much of it from television. We found no peer-reviewed study measuring creative’s share of return in auction media such as search and paid social, and the vendor figure for digital campaigns is an industry claim.
The wear-out evidence comes from single campaigns, one publisher, lab studies and a model simulation. It establishes that repetition wears ads out and that the rate varies widely. It does not give a frequency cap that transfers between accounts.
In a platform A/B test, the delivery algorithm chooses who sees each variant, and it optimizes each arm separately. The variants end up in front of different people. A 2025 paper in the Journal of Marketing shows the winning ad can win because it was shown to more responsive users, and the same ad can look better or worse depending on who it reached.14 A study of experiments run on Meta’s tools, four of whose five authors are employed by or contract with Meta, found that 22% of audience balance checks across 181,890 A/B tests exceeded a common imbalance threshold, against 0.16% across 3,204 randomized lift tests.15
That makes A/B results a fair guide to which execution will perform on that platform and a poor guide to whether the idea itself works. Lift tests and geo holdouts assign exposure at random and answer the second question.
A.2 Sources