How to run concept testing
- Reading time
- 6 minutes
- Assumes
- You have concepts to test and access to an audience
- Updated
- Sep 6, 2026
The two questions that ruin a concept test
"Which do you prefer?" People can always rank things, and their ranking of concepts on a screen predicts very little about behavior. Preference is cheap to express, costs nothing to be wrong about, and reliably favors whatever is most familiar or most polished.
"Would you use this?" Everyone says yes. It's a socially easy answer to a hypothetical, and stated intent for a product that doesn't exist yet is among the weakest signals in research.
Both feel like the natural questions and both produce data that will be cited in a decision and shouldn't be.
What works instead is asking about comprehension, relevance, and what the person currently does — none of which require them to predict their own future behavior.
Ask what they understood, not what they liked
"In your own words, what does this do?" tells you more than any rating scale. If they can't say it back, nothing else you measure about that concept means anything.
Show one concept properly, not five briefly
The temptation is to show everything and ask for a ranking. That gets you a ranking of things nobody understood.
Show one concept, ask everything you want to know about it, then move on. Rotate which one people see first if you're testing several, since order effects are real and consistent.
Two or three concepts per respondent is the practical ceiling. Past that, attention degrades and later concepts get systematically worse evaluations regardless of merit — you'll be measuring fatigue.
If you have eight concepts, use a design where each respondent sees a subset, and accept that you need more respondents to get comparable coverage.
Write concepts at the same level of finish
A polished concept beside a rough one is not a fair test. The polished one wins, and you learn nothing except that polish wins.
Standardize the format: same length, same structure, same level of visual finish, same specificity. If one concept names a price and another doesn't, you're testing pricing rather than concept.
Resist making your favorite better. It happens unconsciously, it's visible in the results, and it means the test can only confirm what you already thought.
Ask about the current behavior first
Before showing anything, ask what the person does today about the problem your concept addresses. This is the most useful data in the whole study and it costs two questions.
You get a baseline that makes the rest interpretable. Somebody who has an elaborate workaround has a real problem and their reaction to your concept is meaningful. Somebody who has never encountered the situation is guessing, and their enthusiasm should be discounted heavily.
It also gives you the comparison your concept has to beat, which is almost never "nothing" and is usually a spreadsheet, a habit, or a colleague they ask.
Probe the surprising reactions
Ratings tell you where to look. The reasons tell you what to do, and you only need them from the people whose reaction was informative.
Someone who rates a concept at the bottom of the scale, or picks an option almost nobody else did, or says something that contradicts an earlier answer — those are the responses worth a follow-up question, asked immediately while they're still in it.
Asking everyone to explain every rating buys you a pile of "n/a" and costs completions.
Test the message and the concept separately
A concept test and a messaging test are different studies, and running them together confounds both.
The concept test asks whether the thing is worth having. Keep the language plain and descriptive so you're measuring the idea rather than the copy.
The messaging test asks which framing of a fixed thing lands. Now the concept is held constant and only the words vary.
Teams that combine them get a winner and can't say whether it won on substance or on phrasing — which matters enormously, because one tells you what to build and the other tells you what to say.
Know what this can and can't tell you
Concept testing is good at elimination and comprehension. It reliably identifies concepts nobody understands, concepts that solve a problem people don't have, and language that confuses.
It's poor at predicting adoption, sizing a market, and ranking two good concepts. If two concepts both test well, the honest conclusion is that both are viable, and the choice should be made on strategy, cost, or feasibility rather than on a two-point difference in a rating.
Say this explicitly when you present. The most common misuse is treating a narrow preference gap as a mandate, and it's usually the researcher's job to stop that.
Common mistake
Testing concepts with an audience recruited for being interested in the category. They understand more, care more, and rate everything higher than the people you'd actually be selling to. Recruit for the real target, including the people who don't currently think about this at all.
Before you field it
0 of 6 checked