A Prototype Test Run Too Early
A team has a vague hypothesis — "we think onboarding is confusing" — but no clear idea what specifically is wrong or why. Instead of exploratory research, they jump straight to building a polished prototype of a redesigned onboarding flow and run a usability test on it. The test goes fine: users complete the new flow without visible struggle. The team ships it. Three months later, the same activation-metric problem that motivated the redesign hasn't moved.
- What generative-vs-evaluative mistake did this team make, and how does it explain the outcome?
- What should the team have done first, and what method(s) would you use for it?
- Is the usability test result ("users completed the new flow without struggle") actually worthless here? Explain what it did and didn't tell them.
1. The mistake
The team ran evaluative research (a usability test on a specific prototype) without first doing generative research to establish what the actual problem was. "We think onboarding is confusing" is an untested hypothesis, not a diagnosis — it names a symptom's general location but not its cause. Building and testing a specific solution before understanding the problem means the usability test could only tell them whether that particular redesign was usable, never whether it addressed the real reason the activation metric was low. If the actual cause was, say, users not understanding why a setup step mattered (a motivation/content problem) rather than struggling with how to complete it (a usability problem), a smoother flow fixes the wrong half of the issue and the metric stays flat — exactly what happened.
2. What should have happened first
A generative research pass before any redesign: analytics/funnel data to see exactly where in the current onboarding the activation drop-off concentrates (quantitative, behavioral), followed by a handful of contextual-inquiry or interview sessions with new users going through the current flow to understand what they're actually thinking and where they hesitate or misunderstand, not just where they technically get stuck. This maps to the Double Diamond's Discover phase — establishing the real problem — before any Develop- or Deliver-stage prototype work begins.
3. Was the usability test worthless?
No — it answered a real, narrower question: "can users physically complete this specific flow without getting stuck." That's valid, useful evaluative information about usability specifically. What it could never answer, because it wasn't designed to, is whether this flow was solving the right problem — a test can only validate what it was scoped to test. The team's error wasn't running a bad usability test; it was treating a passing usability test as proof the underlying, never-diagnosed problem had been fixed.
Share this question