Home Feeds Careers Get in Touch

Critical Incident Technique for Research Interviews (2026)

critical incident approach critical incidents technique critical incident technique critical incident technique cit critical incident technique example
Critical incident technique interview notes showing one remembered event broken into context, action and outcome

TL;DR

  • The critical incident technique is a recall protocol.
  • Instead of asking how someone feels about a product or a service, you ask them to describe one specific event they remember, in enough detail that you can judge what happened and what it led to.
  • John Flanagan formalized it in 1954.
  • Its accuracy depends almost entirely on how fast you field it.

Last updated: 17 September 2026

Quick Answer: The critical incident technique asks a person to describe one specific event they remember, not an opinion they hold. John Flanagan formalized it in Psychological Bulletin in 1954 as a set of procedures for collecting direct observations of behavior. Its accuracy depends mostly on how soon after the event you ask.

In the spring of 1949, Flanagan's team began a study at General Motors' Delco-Remy division. Three matched groups of 24 foremen recorded critical incidents about the people they supervised: one daily, one weekly, one only after a fortnight. The daily group produced 315 incidents, the weekly 155, the fortnightly 63.

Only the delay changed. An incident is perishable, and every rule in the method protects the detail that delay destroys. That is the case for asking soon after the event, which a continuous research program makes routine.

What Counts as a Critical Incident?

A critical incident is one bounded event the respondent can narrate start to finish. Flanagan's 1954 paper defines an incident as any observable human activity "sufficiently complete in itself to permit inferences and predictions to be made about the person performing the act", and calls it critical when its consequences are "sufficiently definite to leave little doubt concerning its effects."

Four things must be present before what you collected counts as an incident:

  • What the person was trying to do. Without the intent nobody can judge whether it helped or hurt.
  • What actually happened, as action rather than a characteristic.
  • The circumstances, enough to follow the sequence.
  • What it led to, concretely.

"The setup process is confusing" is not an incident. "Last Tuesday I was connecting our payroll export before a board meeting, the mapping screen listed both subsidiaries under one code, I picked the wrong one, and re-ran the import next morning" is.

Flanagan also gave the method's cheapest quality test: precision is the proxy for accuracy. A vague report is evidence the event is not well remembered, so vagueness is data, not a style problem.

Why You Collect Good and Bad Incidents Together

Collect effective and ineffective incidents in the same study, in roughly equal numbers, from two separate prompts. An incident only reads as critical against its opposite, and complaints alone produce a category scheme with no way to describe what working well looks like. Bitner, Booms and Tetreault worked this way in the Journal of Marketing volume 54, issue 1 (1990), splitting 700 service incidents into favorable and unfavorable sets.

The enhanced technique adds a third prompt, wish list items: supports absent at the time that the person believes would have helped. Nothing happened, so they are categorized separately, and they carry the clearest product implication.

How Fast Does Recall Decay?

Field a critical incident study in days, not quarters. Against the daily baseline above, Flanagan reported that foremen reporting after one week appeared to have forgotten about half the incidents they would otherwise have given, and after two weeks about 80 percent. A study of air route traffic controllers found something worse than loss: incidents reported months later showed selective recall of the dramatic ones, which biases a sample rather than thinning it evenly.

The Agency for Healthcare Research and Quality lists among its disadvantages that reliance on memory can result in data degradation, alongside participant reluctance where the report touches their own performance. Fair objections to a study fielded late, weak ones inside a week.

Fielding speed is also where recruitment becomes a validity question. If one segment takes three weeks to reach, its incidents have decayed while a faster segment's have not, so the sample looks intact but is not comparable across itself. That bites hardest for people a browser link cannot reach, who will answer a voice note but not a form. Alchemic fields AI-moderated interviews inside WhatsApp with no link and no app, as text or voice notes, and by outbound AI phone call, with managed fieldwork or bring your own sample across 14 markets.

Which Collection Method Fits?

Flanagan documented several collection procedures. Where you can instrument the observer, the daily record form beats every interview on the axis that damages this data, and Flanagan prefers direct observation outright. Where incidents are sensitive, the individual interview earns its fifteen minutes, because the account is checked as it is given. The table is ordered alphabetically by the first column; the first row is the modern addition.

Collection method What it returns Recall lag Interviewer time per incident
AI-moderated interview Narrated incidents, adaptive follow-up, transcribed live Fielding speed No published figure
Daily record form Incidents logged against a checklist as they happen Under 24 hours Observer records, n/a
Individual interview Richest detail, criteria applied live Respondent's own 15.7 minutes
Mailed questionnaire Comparable where respondents are motivated Respondent's own Lowest, no probing

How Many Incidents Are Enough?

Stop when new incidents stop producing new categories. Flanagan's test: count how many new critical behaviors each additional hundred incidents adds, and coverage is adequate when another hundred adds only two or three. His bands were tied to the activity's complexity, not the number of people interviewed: 50 to 100 incidents for simple activities, 1,000 to 2,000 for skilled work, 2,000 to 4,000 for supervisory work.

Modern studies land far below those bands. A study of medical registrars in Scotland placed 221 critical incidents into 24 categories and reported exhaustiveness after six participants. It averaged about 18 usable incidents per participant, and that ratio is the planning number: size the study in incidents, then divide. Thirty people at three incidents is weaker than twelve at twenty.

Build categories out of the incidents, never before them, and report participation rates next to counts: ten incidents from one talkative respondent is not ten from ten people. The enhanced technique adds credibility checks, including independent placement against an 80 percent match rate. Keep the incident set: it is the evidence a persona should rest on.

Where the Critical Incident Approach Falls Short

The method has four real limits, and the first two are not fixable by better fieldwork.

  • It only sees what people can narrate. Routine behavior produces no story, so an incident set under-samples the ordinary. For the middle of an experience rather than its edges, observation or instrumented logging is the right tool.
  • Frequencies count incidents, not people. A category holding 40 incidents looks dominant until you find six respondents produced them all.
  • It is retrospective, and the recall evidence above is unforgiving about studies fielded months later.
  • It asks one person to report on another, an ethical constraint rather than a measurement one.

That last one has teeth. In the original workplace form, respondents describe a named colleague's ineffective behavior, which Flanagan treated as the central obstacle to honest reporting. The ICC/ESOMAR International Code, revised in June 2025, and the Insights Association's Code of Standards both stress duty of care and informed consent. That duty reaches the person being described, who consented to nothing.

The technique is narrow, not fragile. It tells you what happened in a set of events and what they had in common, not how common they are. Where an AI moderator runs the elicitation the known bias risks apply, and this overview of qualitative research maps the alternatives.

Frequently Asked Questions

Who invented CIT and when was the method published?
John C. Flanagan, at the American Institute for Research and the University of Pittsburgh. It grew out of the Aviation Psychology Program of the US Army Air Forces, set up in 1941, where an early study analyzed why 1,000 candidates were eliminated from flight training. The formal statement was published in 1954.
What is the critical incident technique in job analysis?
It builds a functional description of a role from observed behavior rather than opinion. Supervisors report specific effective and ineffective acts, those are grouped inductively into critical requirements, and those feed selection tests, training and performance measures.
Is CIT qualitative or quantitative?
Both, deliberately. Incidents are collected as open narrative and categorized inductively, which is qualitative. The categories are then counted, with frequencies and participation rates reported, which is quantitative. It is usually classed as qualitative with a quantitative analysis stage.
How long should a critical incident interview run?
Long enough for three to five complete incidents, usually 30 to 60 minutes. The constraint is completeness rather than clock time, because each incident needs its intent, actions, circumstances and outcome before it counts. Two full incidents beat six half-told ones.
How is it different from root cause analysis?
Root cause analysis works backward from one known failure to find why it occurred. The critical incident approach works forward from many events, good and bad, to find the behaviors separating effective from ineffective performance. One diagnoses a case, the other an activity.

About the Author

Sreenadh Narayanan is the founder of Alchemic, an AI-powered consumer research platform used for ad testing, concept testing and brand tracking. He writes Alchemic's guides on qualitative research and research methods, covering interview design, sample sizes and how teams turn customer conversations into decisions.