Last updated: 9 September 2026
Quick Answer: AI open-end coding software for survey verbatims in 2026 includes Displayr, Qualtrics Text iQ, Forsta, Ascribe and Thematic, plus Alchemic for interview transcripts. The best published result puts model-to-human agreement at an Adjusted Rand Index of 0.61 against 0.68 between humans. Every tool over-assigns codes unless the code frame carries explicit definitions.
Coding open ends used to be the slowest step in a survey. A coder read a few hundred verbatims, built a code frame, and assigned one or more codes to every answer so the text could be counted.
Software now does the first pass in minutes. What it cannot do on its own is decide whether the codes are right, and the research published in 2025 and 2026 is unusually clear about where it goes wrong. Accuracy is only half of what a brand team needs to know. The other half is where the verbatims came from, and whether that changes what coding can find.
Which AI Tools Code Open-Ended Survey Verbatims in 2026?
Six tools document AI-assisted open-end coding for survey verbatims on their own sites: Displayr, Qualtrics Text iQ, Forsta, Ascribe, Thematic and Relative Insight. Alchemic codes the open ends it collects, which are interview transcripts and voice notes rather than typed survey boxes.
They fall into three groups:
- Survey-native coding. Qualtrics Text iQ and Forsta code open ends inside the survey platform that collected them, so the codes land next to the closed-ended data.
- Standalone analysis tools. Displayr and Ascribe take exported verbatims and build a code frame with AI suggestions that a coder reviews; Displayr's own guidance is to start with around ten themes and review every assignment. Thematic and Relative Insight sit closer to text analytics, with themes, sentiment and comparison across sources.
- Coding inside the interview. Alchemic auto-codes every response into themes as it lands, from an AI-moderated interview on WhatsApp, web or phone, and links each theme to the respondents, the verbatim and the voice note behind it.
Which group a tool belongs to decides what "AI coding" means in its marketing. Survey-native tools code what a survey box captured. Standalone tools code whatever is uploaded. The third group codes a conversation, which is a different input and, as the last section covers, a different output.
Is Open-End Coding the Same as Text Analytics?
No. Open-end coding assigns each verbatim to one or more codes in a frame built for the question, so the result is a set of proportions that can be crossed with the rest of the survey. Text analytics is the broader family of sentiment scoring, topic detection and keyword extraction that runs across any text, from reviews to support tickets, without a question-specific frame.
The distinction matters because the two are scored differently. A code frame is judged on whether a second coder, human or model, would assign the same codes; the standard measure is inter-rater agreement. A sentiment score is judged on whether it tracks an outcome. Different test, different tool.
Displayr's coding guidance captures the discipline. Code the meaning of the response in relation to the question, not the words. Keep an "Other" bucket, and treat more than 10 percent of responses in "Other" as a sign the frame needs work.
The American Association for Public Opinion Research's best-practice note on generative AI, published in November 2024, names the validation step directly. Report how AI-produced findings were checked, which for open ends may mean calculating inter-rater reliability between the AI's coding and a human's. A tool that reports sentiment without a frame cannot be validated that way.
AAPOR's Transparency Initiative goes one step further for anyone publishing results: a content analysis must disclose the coding scheme it used, or state that none was, and whether AI assisted the selection, coding or analysis of the content.
How Do the Named Coding Tools Compare on Coding, Not Just Sentiment?
The table reads each tool's own documentation as of 9 September 2026 and scores coding capability, not the wider analytics each one also sells. "Not stated" means the site does not document it. Rows are alphabetical.
| Tool | Builds a code frame from scratch | Applies a supplied frame | Human review built in | Multilingual verbatims | Survey-native or standalone | Exports to crosstabs |
|---|---|---|---|---|---|---|
| Alchemic | Yes, themes extracted as interviews complete | Yes, guide-driven themes across waves | Yes, drill from theme to verbatim to voice note | 57+ languages including Hindi, Tamil and Telugu, code-mixed speech transcribed and translated | Inside the interview; managed research | Yes, plus the quote bank from the same field |
| Ascribe | Yes | Yes, its coding heritage | Yes, coder workflow | Yes | Standalone coding tool | Yes |
| Displayr | Yes, AI suggests about ten starting themes | Yes | Yes, review and reclassify per theme | Not stated | Standalone, with Q Research Software | Yes |
| Forsta | Yes | Yes | Yes | Yes | Survey-native | Yes |
| Qualtrics Text iQ | Yes, topic suggestions | Yes | Yes | Yes, within the platform's languages | Survey-native | Yes, inside Qualtrics |
| Thematic | Yes, theme discovery | Partial | Yes | Yes | Standalone text analytics | Via dashboards |
Displayr and Thematic own the standalone text-analytics job, and a team with 50,000 review comments should start there. Alchemic's row is narrower and different: it codes what its own moderator collected, so the theme count and the quote bank come from one field, and the input includes speech.
Common Mistakes When Buying Open-End Coding Software
- Buying sentiment when the brief needs a code frame. A polarity score cannot be crossed with a segment the way a coded proportion can.
- Reading "AI-coded" as "validated." Agreement with a human coder is a number the vendor should be able to state for a sample of your own data.
- Skipping the definitions. The 2026 evidence below shows that verbose code definitions with inclusion and exclusion criteria are the single largest accuracy lever.
What Does the Research Say About Automated Coding Accuracy?
The research says a current model can approximate human coding on a well-defined frame. It over-assigns codes when definitions are thin, and varies far more between models and questions than the marketing suggests. Three studies published between June 2025 and August 2026 carry the evidence.
The most direct comparison is an arXiv paper from August 2026 that had five human coders and a GPT-5.4 model inductively code 903 open-ended answers from a European PhD student survey. Agreement between the humans and the model reached an Adjusted Rand Index of 0.61 for coding and 0.54 for theme generation, against 0.68 among the humans themselves and 0.76 for the model's own consistency.
Agreement varied widely by question, and the questions humans disagreed on were the ones the model disagreed with them on.
The over-classification problem is documented in a March 2026 preprint on SocArXiv that benchmarked 21 models across six providers against sociologist-coded ground truth. Every model over-classified, with precision lagging sensitivity by 40 to 50 percentage points, which means default configurations overstate how many people said a thing. The mitigations that worked were verbose category definitions with explicit inclusion and exclusion criteria, unanimous multi-model ensembles, and an automatic "Other" category; ensembles of inexpensive open-weight models beat the best single cloud model.
Model choice is not a detail. It moves the numbers.
A study of German open-ended responses, published in Survey Research Methods in 2025, found performance differed greatly between models. Only a fine-tuned model reached satisfactory accuracy, and unequal performance across categories changed the distribution of answers when fine-tuning was skipped. The coded proportions, the thing a brand team reports, are where the error shows.
Where Do the Verbatims Come From, and Does It Change the Coding?
Where the verbatims come from changes what coding can find, because a typed survey box captures a phrase and an interview captures a reason. Coding a phrase yields a frequency. Coding a probed answer yields a frequency and the reason behind it. The second is rarer.
Coding inside the interview
That is the case for coding inside the interview rather than after it. Alchemic's moderator asks the open question, probes when the answer is thin, catches a contradiction with something said earlier, and codes the exchange into themes as fieldwork runs.
Respondents answer by voice note or text on WhatsApp, by AI phone call, or on the web, in the language they speak, including code-mixed Hindi and English. Transcription and translation are handled automatically. The guide to voice notes as qualitative data covers what a recorded answer carries that a typed one does not.
What one field returns
Two properties follow. Qualitative and quantitative questions sit in the same interview, so the coded theme count and the verbatim quote bank come from one field rather than a survey and a separate qual study.
And every theme is cited. A brand team can click from "packaging confused 41 percent" to the respondents behind it, then to the voice clip. That is the review step the AAPOR note asks for, built into the output rather than added afterward.
Recruitment runs as managed fieldwork or on the brand's own list. The platform publishes 57+ languages including Hindi, Tamil and Telugu, and has fielded in the USA and the UK as well as across India, the Gulf and Southeast Asia. The multilingual interviewing guide covers how coding holds up when the verbatims arrive in three languages in one study, and the piece on turning insights into decisions covers what the coded output is for.
Where Automated Coding Still Fails
Automated coding still fails in five places. Thin definitions, rare codes, sarcasm and negation, mixed-language answers a model was not trained on, and any frame the buyer never validated against a human sample. Each failure has a documented mechanism. Each has a fix.
Failures in the model
Thin definitions. The SocArXiv benchmark's central finding is that models over-assign when a category is described in a phrase. The fix is the one the paper encodes as a default: a definition with inclusion and exclusion criteria, and an "Other" escape valve.
Rare codes. The German study found unequal performance across categories shifted the coded distribution. A code that 3 percent of respondents use is the one most likely to be inflated or missed, and it is often the one the brand cares about.
Ambiguous questions. In the 903-answer comparison, the questions where humans disagreed with each other were the questions where the model disagreed with humans. A vague open end produces vague codes from everyone.
Failures in the input
Speech and mixed language. A model tuned on typed English survey boxes has not seen a Hindi-English voice note about a shampoo. Coding that input needs transcription, translation and a moderator that understood the answer in the first place, which is the bias question the AI moderation guide covers.
AI-written verbatims. The input can be synthetic too. A paper in the Proceedings of the National Academy of Sciences built an autonomous respondent whose open-ended answers were linguistically sophisticated and calibrated to its assigned persona, and which passed 99.8 percent of 6,000 attention-check trials. A coding tool will code those answers as fluently as real ones. The 2025 ICC/ESOMAR Code puts responsibility for the whole chain, sample included, on the commissioning client.
No validation. ESOMAR's 20 questions for buyers of AI-based services and the AAPOR note converge on the same demand: disclose the model, state how the output was checked, and report the agreement figure. A coding tool that cannot produce an inter-rater number on a sample of your verbatims has not been validated for your study, whatever it scored on someone else's.

