Last updated: 23 September 2026
Quick Answer: The ad testing platforms consumer brands shortlist in 2026 are Alchemic, Kantar, System1, Ipsos, Zappi, YouGov, Swayable, Entropik, quantilope, Qualtrics and SurveyMonkey. Kantar, System1 and Ipsos score an ad against norm databases, self-serve tools cost least per study, and Alchemic interviews viewers to explain each score.
An ad testing platform shows an ad to real consumers before launch and measures how it lands. A brand with three cuts of a 30-second spot and six weeks to launch can now get a scored read in hours from a norm database, a randomized persuasion test within a day, or a few hundred recorded interviews within a week. Those are different instruments sold under one label, and the instrument matters more than the vendor.
The evidence that pretests work at all is thinner than the category's marketing suggests. The study its authors called the only public test of fresh commercials against both copy tests and split-cable sales used five pairs of ads, and its best predictors measured liking.
Key Takeaways
- Three tribes, not one market. Norm-database pretests (Kantar, System1, Ipsos, Zappi), panel and experiment platforms (YouGov, Swayable, quantilope, Qualtrics, SurveyMonkey) and interview or signal-led testing (Alchemic, Entropik) answer different questions.
- Norms are the reason to buy the big three. Kantar cites 260,000 tested ads and System1 more than 120,000.
- Liking predicted sales winners 87 percent of the time in the Advertising Research Foundation's validity project, ahead of most persuasion measures.
- Platform A/B tests are not a substitute for a pretest. Delivery algorithms serve each variant to a different mix of users.
- Interview-led testing fits when the reason behind a score matters. Alchemic interviews viewers on video, WhatsApp or AI phone, which also reaches people who never join a browser panel.
What an Ad Testing Platform Measures Before Launch
An ad testing platform shows finished or rough creative to people who resemble the target buyer and measures four things: whether they notice it, whether they remember the brand, whether it shifts their view of the product, and why. Every vendor below covers the first three; they differ most on the fourth and on who the respondents are.
The Advertising Research Foundation's Copy Research Validity Project, reported in the Journal of Advertising Research, was completed in 1990 and is still the reference point. It tested ten packaged-goods commercials in five pairs, each with a known split-cable sales winner, through six copy-testing methods and 12,000 to 15,000 interviews. Average liking picked the sales winner 87 percent of the time, and the average overall brand rating 84 percent.
Liking is not a soft metric. Five pairs is still a small base, though, so treat any single validation claim, a vendor's included, as provisional. The guide to creative testing before launch covers the validation-versus-diagnosis split and sample sizing.
How This Guide Evaluates Ad Testing Platforms
Each platform was read on its own website on 23 September 2026. None was tested hands-on or paid for inclusion. Every vendor was judged on six checkable criteria:
- Service model: self-serve, serviced or full research team.
- Respondent source: own panel, marketplace sample, the brand's list, or managed recruitment.
- What it reports: recall, persuasion, attention, diagnostics, verbatims.
- Norms: whether a stated database lets an ad be ranked against others.
- Stated speed: the fastest result the vendor publishes.
- Channels: how the respondent sees the ad and answers.
Ad-ops tools that rotate live variants are left out; the A/B section explains why. "Not stated" means the vendor's page did not document it.
Comparison at a Glance
Rows are sorted alphabetically by platform name.
| Platform | Service model | Respondents | What it reports | Norms stated | Fastest stated result |
|---|---|---|---|---|---|
| Alchemic | Research team plus self-serve | Managed fieldwork or bring your own, 14 markets including the USA and UK; video, WhatsApp or AI phone | Unprompted recall, element-by-element probes, per-second facial read, intent by variant, separate brand-recall check, verbatims | None published; knowledge base carries across studies | Fielding within 48 hours, report within a week |
| Entropik (Decode) | Platform | Not stated on the ad testing page | Facial coding, eye tracking, attention, performance prediction, variant comparison | Not stated | Not stated |
| Ipsos (Creative|Spark) | Self-serve to full service | Not stated on the page | Attention in a distracted setting, facial coding, sales-validated Creative Effect Index | Creative|Spark benchmarks, database size not stated | 24 hours; AI version 15 minutes |
| Kantar (LINK+) | Self-serve or serviced | Not stated on the page | Brand equity KPIs, attention, facial coding, predicted brand lift | 260,000 ads | 6 hours; LINK AI in minutes |
| Qualtrics | Template in the Research Core license | Own contacts or purchased sample; 300 completes typical | Purchase intent, brand recognition, consideration, change in impression | Not stated | Not stated |
| quantilope | Automated platform plus consulting team | Not stated on the page | A/B pre-roll test, implicit association test, video open-ends | Pre/post benchmarks | 24 to 48 hours, per a client quote |
| SurveyMonkey (LaunchPad) | Self-serve, priced per study | Own contacts or a 335M+ panel in 130+ countries | Appeal, humor and memorability scorecards, key driver analysis | Industry benchmarks | Targeted responses in hours |
| Swayable | Platform | Verified respondents in randomized controlled trials | Persuasion lift on favorability and purchase intent, qualitative feedback | Not stated | 24 hours |
| System1 (Test Your Ad) | Platform plus AI screening | Real viewers across 80+ markets | Star, Spike and Fluency metrics, attention | 120,000+ ads | 24 hours; AI screen in minutes |
| YouGov | Platform plus research team | 30 million+ registered panel members, re-contact | Prompted and unprompted awareness, recall, likability, claim believability | Not stated | Not stated |
| Zappi (Amplify) | Platform plus professional services | Not stated on the page | In-context and forced exposure, recall, ad skipping, AI reports | Brand, category, country and custom norms | Not stated |
How to Read the Norms Column
A stated database lets a new spot be ranked against thousands of past ones. Where the cell reads "Not stated", the study compares only the variants inside it, which is still the most durable read a pretest gives.
For a creative that has to explain itself, or reach viewers a browser panel misses, see how Alchemic runs ad tests from storyboard to post-launch wave.
Ad Creative Testing Platforms for Consumer Product Advertising
Entries after slot 1 run alphabetically. Each states what the platform is best for, what its own site says, and one limit.
1. Alchemic: Best for Knowing Why an Ad Scored the Way It Did
Alchemic runs end-to-end consumer research at scale, and its ad testing replaces the fixed questionnaire with an AI-moderated interview. Each respondent first says what stayed with them a minute after viewing, then the moderator probes the opening, music, end-frame branding and call to action one at a time. On camera-on interviews it reads facial emotion per second across the spot, and a separate prompt later asks which brand the ad was for without showing it again, as the ad testing service page describes.
Studies run 200 to 500 interviews per market across 2 to 10+ variants, with fielding live within 48 hours. Recruitment is managed fieldwork or bring your own across 14 markets including the USA and the UK, and the service publishes 57+ languages including Hindi, Tamil and Telugu.
Limit: Alchemic publishes no normative database, so it cannot tell you how a spot ranks against thousands of category ads.
2. Entropik (Decode): Best for Eye Tracking, Attention and Facial Coding
Decode reads creative through facial coding, eye tracking and attention measurement across TV, digital, print and social formats, uses AI models to forecast campaign performance, and compares variations side by side. It also runs AI-moderated interviews, so a team can pair the signal trace with conversation. Limit: facial and gaze signals need respondents willing to switch a camera on, which shapes who completes.
3. Ipsos (Creative|Spark): Best for Sales-Validated Scores in a Distracted Setting
Creative|Spark captures attention in a distracted environment rather than a forced viewing, applies facial coding as standard, and reports a sales-validated Creative Effect Index. It runs from self-serve to full service with results in as little as 24 hours, and Creative|Spark AI predicts responses in as little as 15 minutes on Ipsos.Digital. Limit: the full study is built for hero campaigns; routine social cutdowns sit more naturally on the AI tier.
4. Kantar (LINK+): Best for Benchmarking Against the Largest Ad Database
LINK+ tests ideas, storyboards and finished film across TV, digital, audio, print and outdoor, backed by what Kantar calls the world's largest ad testing database of 260,000 ads. It layers facial coding, a digital attention framework and predicted brand lift over its brand equity framework, with results in as few as 6 hours. Limit: a fixed questionnaire asks everyone the same questions, so an unanticipated reaction surfaces only if an open-end catches it.
5. Qualtrics: Best for Teams Already Licensed on Qualtrics
The Advertising Creative Testing solution is an expert-built template included with a Research Core license. It measures purchase intent, brand recognition, consideration and change in impression, recommends about 300 completes, and works with uploaded contacts or purchased respondents. Limit: the solution page lists English only, and the team runs the study itself.
6. quantilope: Best for Implicit and A/B Pre-Roll Methods on One Platform
quantilope automates advanced methods: an A/B pre-roll test, a single implicit association test for immediate associations, and inColor video open-ends analyzed automatically. A consulting team supports projects, and its site quotes OMD Germany completing campaign tests in 24 to 48 hours. Limit: the method menu rewards a team that knows which method it needs.
7. SurveyMonkey (LaunchPad): Best for Low-Cost Self-Serve Scorecards
LaunchPad's video and image ad tests score appeal, humor and memorability against industry benchmarks, and run key driver analysis. Brands survey their own contacts at no extra charge or buy responses from a 335M+ panel across 130+ countries. Limit: the scorecard ranks variants well but leaves the diagnosis to the buyer.
8. Swayable: Best for Measuring Persuasion With Randomized Trials
Swayable runs randomized controlled trials on verified respondents, with an exposed group and a control, and reports lift on brand love, favorability and purchase intent with results in 24 hours. Limit: an RCT measures whether an ad moved opinion, not which frame or line did it.
9. System1 (Test Your Ad): Best for Long-Term Brand-Building Prediction
Test Your Ad measures how real viewers feel second by second and turns it into Star, Spike and Fluency metrics that predict long-term growth and short-term sales, with results from 24 hours. It benchmarks against 120,000+ ads, including every new US and UK TV ad, across 80+ markets, and Test Your Ad Screen triages up to 100 assets at a time. Limit: if the decision turns on whether one product claim was understood, that needs a claim-specific question set.
10. YouGov: Best for Panel Reach With Re-Contact
YouGov shows each respondent a single execution to avoid comparison bias, then measures prompted and unprompted awareness, recall, likability and claim believability. Its 30 million+ registered panel members can be re-contacted for the reasons behind an answer, including by video. Limit: an online opt-in panel reaches whoever joins online panels.
11. Zappi (Amplify): Best for a Continuous Ad Research System
Zappi's Amplify system tests early ideas, storyboards, TV and streaming, digital, static and social video, combining in-context exposure that captures skipping and recall with forced exposure for core KPIs. It states it is 60 percent more predictive of in-market outcomes and reports against brand, category and country norms. Limit: standardized paths trade conversational depth for comparability, as the Alchemic and Zappi comparison sets out.
Recall, Persuasion and Diagnostics Compared
Every platform reports a mix of three outputs, and the mix is the buying decision.
Recall and branding. Zappi's in-context exposure captures recall, System1's Fast Fluency measures brand recognition after two seconds, and YouGov measures prompted and unprompted recall. Alchemic separates immediate unprompted recall from a later brand-attribution check, which catches an ad remembered for the wrong brand.
Persuasion. Swayable is the purest case: a randomized exposed-versus-control design. Kantar's predicted brand lift, Qualtrics's purchase intent, SurveyMonkey's scorecards and Alchemic's intent read by variant measure stated persuasion among people who saw the ad.
Diagnostics. Facial coding (Kantar, Ipsos, Entropik and Alchemic's camera-on interviews), second-by-second viewer response (System1) and implicit association (quantilope) locate where a reaction happened. Interview-led testing then asks the viewer why.
Predicted Scores Without Respondents
Kantar LINK AI, System1's Test Your Ad Screen, Ipsos Creative|Spark AI and attention-prediction tools such as Neurons score creative in minutes from models trained on past human data. They are useful for triaging 50 social cutdowns. They inherit the norms they were trained on, so a new format or an audience missing from the training panels is exactly where a prediction is weakest.
Who a Browser Panel Never Shows the Ad To
Most platforms above recruit from online panels, and for a mass-market consumer brand that is a narrower group than the media plan buys.
The US Census Bureau's 2021 American Community Survey report on computer and internet use found 11 percent of households reached the internet only through a smartphone data plan. That share was 16 percent among households earning $25,000 or less, against 5 percent at $150,000 or more. Smartphone-only households are also more likely to be headed by someone 65 or older.
A value-tier brand, a regional grocery line or a campaign aimed at older buyers should ask how many phone-only, lower-income or older viewers its panel actually holds. Alchemic addresses this through channel choice: a WhatsApp interview with no link and no app, or an outbound AI phone call that records consent on the first turn. Voice notes and calls carry no facial signal, so camera-on video stays the mode for a second-by-second read. The guide to reaching respondents without smartphones covers the mechanics.
When a Creative Pretest Beats an In-Market A/B Test
A pretest beats a platform A/B test whenever the question is about the ad rather than about the audience the algorithm found for it. Michael Braun and Eric Schwartz's Marketing Science Institute report on divergent delivery shows that platform A/B tools serve each ad to a different, optimized mix of users even during the test, which can confound the magnitude, and even the sign, of the result.
So a Meta split test that crowns version B may be reporting that the algorithm found better prospects for B. A pretest shows every variant to comparable people. The in-market test still wins when the goal is simply to scale whichever asset converts best on that platform.
Exposure differs too. The Media Rating Council's mobile viewable impression guidelines count a video ad as viewable after 2 continuous seconds with half its pixels on screen, while a forced-exposure pretest plays all 30. Zappi's in-context exposure and Ipsos's distracted-setting measure exist to close that gap.
Tools such as Madgicx, Motion and Marpipe belong to a separate tribe: they analyze and rotate live variants inside ad accounts, optimizing spend rather than researching the audience. For measuring what a live campaign did, the guide to brand lift study providers covers the controlled-exposure designs.
When Another Platform Is the Better Choice
Five situations point away from interview-led testing:
- You need a normed score for a board or an agency contract. Kantar, System1 and Ipsos rank an ad against databases no interview study can match; Alchemic's own service page says to use a norm database when that is the need.
- You need causal persuasion proof. Swayable's randomized design answers "did it move opinion" more cleanly than any interview.
- You are screening 50 social cutdowns this week. An AI screening tier is faster and cheaper per asset.
- The budget is small and the team already uses SurveyMonkey or Qualtrics. A self-serve template ranks variants well enough.
- You want gut-level associations. quantilope's implicit test measures what a stated answer may hide.
How to Run a Two-Platform Trial
- Pick one live decision, such as which of three cuts runs nationally.
- Run the same variants on two platforms of different tribes, one norm-based and one interview-led.
- Match the sample to the media plan, including income, age and phone-first viewers.
- Check brand attribution separately from recall, and later.
- Compare both recommendations with in-market results after launch.
For multi-market work, check how each vendor handles language using the guide to multilingual AI-moderated interviews.
Where Ad Testing Platforms Fall Short
Pretests share limits whichever vendor runs them:
- Validation is thin. The ARF project rested on five commercial pairs, and vendor validation studies are mostly unpublished for outside review.
- Attention is artificial. A forced viewing overstates how much of the ad a scrolling viewer sees; in-context designs narrow the gap without closing it.
- Norms favor the familiar. An unusual execution can score below category averages that were built on conventional work.
- Predicted scores inherit their training data, so new formats and under-sampled audiences are where they are least reliable.
- Guide latitude varies by tool. How far an interview can leave the script to chase an unexpected reaction differs between systems. Once it has the client's brief, Alchemic designs and tailors the discussion guide to handle these risks before fielding, rather than leaving that work to the buyer.
Sources and Methodology
Each external source is linked inline where its claim appears.
- Advertising Research Foundation, Copy Research Validity Project (Haley and Baldinger, Journal of Advertising Research reprint, 2000): how well copy-test measures picked known split-cable sales winners.
- Marketing Science Institute Report 24-122 (Braun and Schwartz, 2024): why platform A/B tests confound ad content with algorithmic audience selection.
- US Census Bureau, Computer and Internet Use in the United States: 2021: smartphone-only households by income and age.
- Media Rating Council, Mobile Viewable Ad Impression Measurement Guidelines (2016): the 2-second viewable video standard.
- Vendor facts: each platform's own website, read on 23 September 2026.

