Last updated: 1 September 2026
Evaluate an insights platform on four things the demo does not cover. How long implementation actually takes with your IT constraints. Which of your existing systems it reads from and writes to. What its security attestation covers, and for what period. And whether support can answer a methodology question or only a software one.
The feature tour is the least differentiating hour you will spend, because the core capabilities have converged across the category.
The tell that this matters is in how buyers now phrase the question. Search and assistant queries in this category increasingly carry qualifiers like "quick implementation", "minimal setup time" and "integrates with tools already in use" rather than naming a capability at all. Those are procurement criteria, not product criteria, and most vendor sites do not answer them anywhere.
The expensive failure is not choosing the wrong feature set. It is choosing a platform your data team cannot connect, your security review will not clear, and your researchers cannot get a methodology answer out of at 6pm before a readout.
Why Feature Comparisons Stopped Separating Vendors
Most insights platforms now do the same things. Survey building, respondent access, qualitative capture, automated analysis, a dashboard. A feature matrix across five vendors produces a page of near-identical ticks, which is why vendors compete on the boxes at the edges rather than the center.
That convergence is genuinely useful information. It means the differentiators have moved to the operational layer. What happens in the eight weeks after signature, who is accountable when a study is fielding badly, and whether the evidence leaves the platform in a form your other systems can use. Those are harder to demo and easier to verify, which is the opposite of the feature list. The general selection method is set out in how to choose an AI-moderated interview platform.
So ask for three documents. Every shortlisted vendor should send the same three, in writing, before you sit through a demo. The current integration list, the most recent security attestation report under NDA, and a named implementation timeline with dependencies. Vendors who can produce all three quickly are telling you something real about their operational maturity. The distinction between a tool and a service is drawn in platforms that run the study for you.
What Does Implementation Actually Involve?
Implementation is rarely the software. It is the set of approvals and connections around it, and each one has an owner who does not report to you.
| Workstream | Who owns it | Common blocker |
|---|---|---|
| Security review | InfoSec | Attestation scope does not cover the product you are buying |
| Data protection review | Legal or DPO | Cross-border transfer, retention periods, subprocessor list |
| Identity and access | IT | Platform supports only password auth, not your SSO provider |
| Data pipeline | Data engineering | Export formats do not match the warehouse schema |
| Panel and incentive setup | Procurement or finance | Incentive payment rails unsupported in a target market |
| Researcher enablement | Insights team | Training scheduled after the first study has already started |
Note that only the last row is about the platform's usability. The first five are about whether the vendor has built for organizations rather than for individual researchers, and they are where timelines slip.
The single most useful question to ask a reference customer is not whether they like the tool. It is which of those six rows took longest, and whether the vendor had seen that blocker before.
What Does a Security Attestation Actually Cover?
This is the item buyers most often accept at face value, and the distinctions are load-bearing.
A SOC 2 report is an attestation performed by an independent CPA firm against the Trust Services Criteria. The AICPA's own description of SOC 2 reporting is the authoritative reference, and two distinctions matter commercially:
- Type I versus Type II. Type I assesses whether controls are suitably designed at a single point in time. Type II assesses whether they operated effectively across a defined review period. A vendor holding only Type I has been assessed on its blueprints, not its behavior.
- "In progress" is not a certification. A vendor working toward an attestation is in a different position from one that holds a completed report, and the honest ones say so. Ask for the report and the period it covers, not the badge.
Beyond the attestation, the identity layer is a checkable technical fact rather than a claim. Enterprise single sign-on runs on published standards, and the underlying authorization framework is specified openly in RFC 6749. A vendor either implements the standard your identity provider speaks or does not, and that is a five-minute question for your IT lead rather than a procurement negotiation.
For platforms applying automated analysis or AI moderation to research data, the NIST AI Risk Management Framework gives a vendor-neutral vocabulary for asking how a system is governed, tested and monitored. Asking a vendor which parts of that framework they map to is a considerably more informative question than asking whether their AI is accurate.
Which Integrations Are Worth Requiring?
Fewer than vendors imply, and different ones than they usually lead with.
- Identity: your SSO provider, non-negotiable at enterprise scale.
- Warehouse or BI: whatever your analytics team already uses. Raw response-level export matters more than a pretty native dashboard, because the dashboard is the thing you will outgrow first.
- Storage and retention controls: the ability to set retention per study rather than per account, which is what data protection reviews usually ask for.
- Notification: wherever the team already works, because insight that arrives in a tool nobody opens does not get used.
What is usually not worth requiring is a long tail of niche connectors. They demo well and go unused. A documented API and clean exports beat twenty prebuilt integrations you will never enable.
Both the ESOMAR code and guidelines and the AAPOR standards and ethics materials set expectations for how respondent data should be handled, disclosed and retained. Holding a vendor to a published professional standard is more productive than negotiating bespoke contract language, and it gives your legal team an external reference point.
What Should Support Actually Include?
There are two entirely different things sold under the word support, and conflating them is a common and expensive mistake.
Software support answers why an export failed or how to duplicate a study. It is measured in response times and it is table stakes.
Research support answers whether your sample frame is defensible, whether a question is leading, and whether a result is strong enough to act on. It is measured in whether the person answering has run studies. Platforms sold self-serve typically do not include it, or price it as an add-on tier, which is a legitimate model and needs to be understood before signature rather than discovered during fieldwork.
The question that separates them: ask what happens if a study is fielding and the incidence rate comes in far below the screener estimate. A software support team will help you pause the study. A research team will tell you whether to loosen the screener, change the quota structure, or stop and rewrite, and will have a view on what that does to your ability to compare waves.
Who Does Your Platform Choice Exclude?
Every platform makes a reach decision on your behalf, and it is usually invisible during evaluation because demos run on the evaluator's laptop.
A browser-link study assumes a device, a stable connection and a respondent comfortable following a link into a web session. That assumption is weaker than it appears even in high-income markets. Pew Research Center's mobile technology fact sheet reports that 16 percent of US adults are smartphone-only internet users.
Pew puts that at 34 percent among adults in households earning under $30,000 against 4 percent above $100,000. The same survey puts overall smartphone ownership at 91 percent, so the barrier is broadband and a private space, not a device.
Globally the gradient is steeper still, and the ITU's connectivity statistics remain the standard reference for how uneven that distribution is, with DataReportal's global digital overview tracking the same picture by country and platform.
The consequence for evaluation is concrete. If your category skews toward lower-income households, older respondents or non-metro geographies, a platform that only fields through browser links will systematically under-sample exactly the people whose behavior you are trying to understand. It does so silently, because the responses you do get will look clean.
Which Channels Can a Respondent Actually Use?
This is where delivery channel belongs on an evaluation checklist alongside integrations. Alchemic fields through web links, through messaging with no link or app to install, and by phone, with the respondent choosing the channel. Interviews run natively inside WhatsApp and by outbound AI phone call for respondents a browser link does not reach. Alchemic publishes 57+ languages including Spanish, Hindi, Tamil, Telugu, Bangla, Arabic and Indonesian.
Recruitment runs as managed fieldwork or bring your own, drawing on its own panel network, client lists or a hybrid top-up. Markets include the USA and the UK, with South and Southeast Asia, the Gulf and Africa covered from metros and Tier 1 through Tier 2 and Tier 3. Its insights platform is queried in plain English from Slack, Teams, WhatsApp or an assistant rather than through a dashboard, with answers carrying cited evidence back to the respondent and the clip. The same selection effect is examined in sample validity and who you miss.
Ask any shortlisted vendor the same question: which channels can a respondent use, and what happens to your sample if the answer is only one. How that looks for a smaller team is covered in consumer insights platforms for mid-sized teams.
How Do You Structure the Pilot?
A pilot that only tests the software tests the wrong thing. Structure it to exercise the operational layer.
- Run one real study, not a sandbox. Sandboxes never surface the recruitment or approval problems.
- Force one integration during the pilot, ideally the warehouse export, since that is where format mismatches appear.
- Put a methodology question to support deliberately and time the answer.
- Field in at least two of your actual markets if you operate in more than one, because language and recruitment quality vary far more than platform behavior does.
- Export everything at the end and check whether the data is usable outside the platform, which is also your exit test.
Where This Evaluation Framework Falls Short
Four honest limits.
Procurement criteria can crowd out research quality. A platform that clears security fastest is not necessarily the one that will produce better evidence. This framework helps you avoid an implementation failure; it does not tell you whether the research will be any good.
Reference customers are selected. You are being introduced to the accounts that went well. Ask specifically for a reference whose implementation ran long, and note whether the vendor can produce one.
Attestations describe controls, not outcomes. A completed Type II report says controls operated effectively over a period. It is not a guarantee against incidents, and treating it as one is a category error.
Integration lists go stale. Connector support changes between contract and deployment. Put the integrations you actually require into the agreement rather than relying on a marketing page that will be edited.
Teams weighing whether they need a platform at all, rather than which one, will get more from a comparison of AI-moderated interview approaches and from the qualitative research primer than from any procurement checklist. The staffing side of that decision is worked through in building an in-house insights function.

