Paths Subjects Questions Quizzes Pricing Search

Usability Testing and Evaluation

Moderated vs. unmoderated testing, heuristic evaluation, task success metrics, the SUS scoring formula, think-aloud facilitation, and knowing usability testing from A/B testing when interviewers push on 'how would you validate this'

Usability Testing and Evaluation

"How would you know if this design actually works?" is the question this subject prepares you to answer, and it's asked in almost every UI/UX interview in some form — sometimes literally, sometimes disguised as "walk me through how you'd validate this before launch" or "how do you know your redesign was actually an improvement?" The weak answer is "I'd do usability testing," delivered as if that were one activity with one correct execution. The strong answer distinguishes which evaluation method fits the question being asked, the constraints in play, and the type of evidence needed — the same "toolbox, not a tool" framing that user-research-methods establishes for research broadly, applied specifically to the Deliver end of the process where a concrete design already exists and needs to be checked.

This subject is deliberately narrow: it's not about generating ideas or understanding user needs (that's Discover-phase research), it's about evaluating something that already exists — a wireframe, a prototype, or a shipped feature — against real usage, real heuristics, or a real metric. Every method here answers a different flavor of "does this work," and the interview skill being tested is knowing which flavor a given question is actually asking about.

It builds directly on two earlier subjects in this track. interaction-design-principles introduced Nielsen's 10 usability heuristics as a checklist for expert review; this subject is where that checklist becomes an evaluation method (heuristic evaluation) and gets compared against testing with real users. user-research-methods introduced the qualitative/quantitative and attitudinal/behavioral axes and Nielsen's 5-user sample-size rule; this subject applies both directly to usability testing and shows where they stop applying once you cross into A/B testing.


Moderated vs. Unmoderated Testing

Both put a real user in front of a real task. What differs is whether a facilitator is present, live, while it happens.

Moderated testing — a facilitator runs the session in real time, whether in person or over video, watching the participant attempt tasks and asking follow-up questions as things happen: "what did you expect to happen there?", "you paused for a second before clicking — what were you thinking?" The facilitator can redirect a participant who's gone down an irrelevant path, probe an ambiguous reaction, or dig into why a mistake happened instead of just recording that it did. This makes moderated sessions the right tool for exploratory evaluation (an early prototype where you don't yet know what questions to ask) or complex, multi-step tasks where a participant getting stuck needs a human judgment call about whether to let them struggle or intervene. The cost is real: scheduling live sessions, paying a facilitator's time per participant, and a ceiling on how many sessions you can realistically run in a week.

Unmoderated testing — participants complete tasks independently, usually through a remote-testing tool that records their screen, clicks, and (often) a think-aloud audio track, with no one watching live. Nobody is present to ask a follow-up, redirect confusion, or catch a participant quietly giving up outside the recorded task window. What you get is exactly what the participant did and said unprompted — nothing more. This trade-off is precisely why unmoderated testing is cheaper, faster, and more scalable (you can run 30 sessions overnight across time zones for the cost of a few moderated ones), but structurally worse at explaining why something went wrong beyond what the participant happened to narrate on their own.

Moderated Unmoderated
Facilitator present live? Yes No
Can probe / ask follow-ups in real time? Yes No — only what's said/done unprompted
Cost per participant High (facilitator time, scheduling) Low (tool cost, no live time)
Speed to run a full study Slower — bound by scheduling Faster — parallel, async, any timezone
Best fit Exploratory, early-stage, or complex multi-step tasks Well-defined tasks, larger sample, tight timeline
What you lose Nothing structural — this is the richer method The ability to redirect, clarify, or dig into an ambiguous reaction

Decision factors, in the order that actually matters:

  1. How well-defined is the task? A rough concept where you're still discovering what confuses people needs a human who can adapt mid-session — moderated. A specific, scripted task ("find and change your billing address") doesn't need adaptive probing to be informative — unmoderated works.
  2. Timeline and budget. Five moderated sessions in a week is realistic; thirty is not, without a research team dedicated to it. Unmoderated tools remove the scheduling bottleneck entirely.
  3. Does "why" matter more than "what"? If the deliverable is a list of specific reasoning behind failures (needed to redesign with confidence), moderated wins even at higher cost. If the deliverable is a broader signal across more people (does this pattern replicate, not just why did one person struggle), unmoderated's larger reachable sample is the better trade.

A strong interview answer names the specific factor that decided it, not just "it depends" — e.g., "I'd run this unmoderated because the task is a well-defined checkout flow and I need signal from 15+ participants across devices in three days, which a moderated schedule can't fit."

A worked scenario. Say you're evaluating a redesigned account-settings page two weeks before launch, with a research budget for eight sessions. If the redesign consolidated three previously separate settings screens into one and you genuinely don't know how users will react to the new structure, that uncertainty argues for moderated sessions — you want a human present to ask "what did you expect to find under this section" the moment a participant hesitates. If instead the redesign is a straightforward visual refresh of an already-familiar settings page and the goal is simply confirming nothing regressed, unmoderated sessions across a larger, more representative sample answer that question just as well for a fraction of the cost per participant.


Heuristic Evaluation vs. Usability Testing

Heuristic evaluation is an expert review, not a study with real users: one or more trained evaluators inspect an interface and flag violations against a known checklist — almost always Nielsen's 10 usability heuristics, covered in interaction-design-principles. No participants, no recruiting, no scheduling — an evaluator can produce a findings list in a day. Its limitation is structural, not a matter of the evaluator's skill: it only surfaces problems that map to a known category the checklist already anticipates. A confusing interaction that doesn't cleanly violate any of the 10 heuristics — a mismatch between this specific product's users' mental model and the interface, for instance — can pass a heuristic review clean and still trip up real users constantly.

Usability testing puts an actual member of the target audience in front of the real interface attempting a real task. It's slower and costs more per finding, but it finds problems nobody predicted in advance, because the evaluator isn't checking against a fixed list — they're watching what an unscripted human actually does.

Nielsen's own guidance on heuristic evaluation is specific and worth citing by name in an interview: a panel of 3-5 evaluators, each reviewing independently before findings are pooled, finds the large majority of heuristic-violation problems — a single evaluator misses too much (different reviewers naturally spot different violations), and beyond 5 the marginal new findings drop off sharply, the same diminishing-returns curve user-research-methods describes for the 5-user usability-testing rule, just applied to expert reviewers instead of end users. Critically, Nielsen frames heuristic evaluation as a complement to user testing, not a replacement for it — the two catch different problem categories (known heuristic violations vs. everything else), and a mature evaluation process runs both: a cheap heuristic pass early to catch the obvious, known-category issues before spending user-testing budget on problems a checklist would have caught for free.

Heuristic Evaluation Usability Testing
Who's involved Expert evaluator(s), no real users Real target users
Speed / cost Fast, cheap Slower, costlier
Finds Known categories of problem (heuristic violations) Problems nobody predicted, including ones outside any checklist
Best used Early, to catch the obvious before user-testing budget is spent Whenever the question is "will real users actually succeed at this"
Nielsen's guidance 3-5 independent evaluators, pooled 5 users per round, iterate between rounds

Interview framing that signals depth: "I'd run a quick heuristic evaluation first to catch the known-category issues cheaply, then spend the usability-testing budget on what that pass can't catch" is a stronger answer than treating the two as interchangeable synonyms for "checking the design," or worse, treating heuristic evaluation as a substitute for ever talking to a real user.


Task Success Metrics

Once a usability test is running, three metrics do the actual diagnostic work, and the interview mistake is naming only one as if it were sufficient on its own.

  • Completion rate — did the participant finish the task? Usually binary (finished / not finished), sometimes scored as partial credit (finished with help, finished via a workaround). This is the bluntest signal: a low completion rate means the design has a hard blocker somewhere, full stop.
  • Time-on-task — how long the task took. Useful for comparing efficiency between a current design and a proposed one, or between two candidate flows, but misleading read in isolation: slower isn't automatically worse. A participant who takes longer because they're carefully reading a confirmation screen before a destructive action is not "struggling" — they're doing exactly what a well-designed warning should make them do. Time-on-task needs the other two metrics next to it to interpret correctly.
  • Error rate — how many mistakes (wrong clicks, backtracks, invalid form submissions) happened before success or abandonment. This is what separates a clean success from a lucky one.

Why you need all three together, not any one alone: a participant can complete a task quickly with a high error rate — several wrong turns, corrected fast — which looks like a success in the completion-rate and time-on-task numbers alone, but the error rate is telling you the design produced confusion the participant happened to recover from, not that the design worked. Conversely, a participant who completes a task slowly with zero errors might be reading carefully in a genuinely well-designed, information-dense flow, not struggling at all. Read one metric alone and you'll draw the wrong conclusion in both directions; read all three together and the pattern becomes diagnostic — for example, high completion + high time + high errors together describe a design that's survivable but not usable, which is a specific, actionable finding none of the three metrics states by itself.

Pattern Completion Time Errors Likely read
A High Low Low Genuinely working design
B High Low High Participant got lucky, not that the design is good
C High High Low Possibly fine — careful reading, not confusion (check qualitatively)
D High High High Survivable but not usable — real friction the user fought through
E Low Hard blocker; time/error numbers are almost secondary to this

A worked example. Two candidate checkout flows are each tested with 8 participants attempting the same purchase task. Flow A: 100% completion, average 45 seconds, an average of 0.2 errors per session. Flow B: 100% completion, average 38 seconds, an average of 1.6 errors per session. Read time-on-task alone and Flow B looks like the winner — faster. Read all three together and the story flips: Flow B's participants are completing faster by recovering from more mistakes, not because the flow is clearer, while Flow A's participants move a little slower with almost no wrong turns at all. Shipping Flow B on the strength of its time-on-task number alone would ship the design with more actual friction in it. This is exactly the kind of finding that later rolls up into product-level metrics — design-metrics-and-product-thinking covers how task-level usability numbers like these connect to business metrics like activation and drop-off.


System Usability Scale (SUS)

The System Usability Scale, developed by John Brooke in 1986, is the standard attitudinal instrument for usability — it measures perceived usability (what the participant believes about the system after using it), not task success. This is a direct application of the attitudinal/behavioral axis from user-research-methods: SUS lives squarely in the attitudinal cell, and it is not a substitute for the behavioral task-success metrics above — a design can produce a high completion rate with a low SUS score (it worked, but the participant found it effortful and unpleasant) or the reverse (a participant who liked how the interface felt but actually failed the task). Reporting SUS alone, without task metrics, tells you how the design felt, not whether it worked.

The instrument is 10 fixed statements, alternating positive and negative wording, each rated on a 5-point Likert scale from "Strongly disagree" (1) to "Strongly agree" (5):

# Statement Wording
1 I think that I would like to use this system frequently Positive
2 I found the system unnecessarily complex Negative
3 I thought the system was easy to use Positive
4 I think that I would need the support of a technical person to use this system Negative
5 I found the various functions in this system were well integrated Positive
6 I thought there was too much inconsistency in this system Negative
7 I would imagine that most people would learn to use this system very quickly Positive
8 I found the system very cumbersome to use Negative
9 I felt very confident using the system Positive
10 I would need to learn a lot of things before I could get going with this system Negative

Scoring, item by item: for the odd-numbered, positively worded items, subtract 1 from the raw score. For the even-numbered, negatively worded items, subtract the raw score from 5. This converts every item onto a common 0-4 scale where higher always means "more usable," regardless of whether the statement itself was worded positively or negatively — without this step, a "5" on a negative item (strong agreement that the system is cumbersome) would wrongly look as good as a "5" on a positive item. Sum all 10 converted scores (range 0-40) and multiply by 2.5 to produce a final score on a 0-100 scale.

Worked example — a participant answers: 4, 2, 5, 1, 4, 2, 5, 1, 4, 2 (items 1-10 in order):

Item Raw Wording Conversion Converted
1 4 Positive 4 − 1 3
2 2 Negative 5 − 2 3
3 5 Positive 5 − 1 4
4 1 Negative 5 − 1 4
5 4 Positive 4 − 1 3
6 2 Negative 5 − 2 3
7 5 Positive 5 − 1 4
8 1 Negative 5 − 1 4
9 4 Positive 4 − 1 3
10 2 Negative 5 − 2 3

Sum = 34. Final SUS score = 34 × 2.5 = 85.

Interpretation. A raw 0-100 SUS score isn't a percentage and shouldn't be read like a test grade — it's benchmarked against the distribution of SUS scores collected across hundreds of real studies. Bangor, Kortum, and Miller's widely cited benchmark research puts the average SUS score at roughly 68; a score above that is above-average perceived usability relative to other products that have been SUS-tested, and a score below roughly 51 falls into the bottom, "poor" range of that same distribution. The number 85 in the worked example above, in that context, reads as a strong result — well above the 68 average.

Interview-relevant point: naming that SUS is attitudinal, that a score needs the ~68 benchmark to mean anything (a bare "72" means nothing without knowing what average looks like), and that it should be reported alongside behavioral task metrics rather than instead of them, is a substantially stronger answer than knowing SUS exists as a "usability score you can get from a questionnaire."

Why alternating wording is deliberate, not arbitrary. Mixing positive and negative statements guards against acquiescence bias — the tendency of some respondents to agree with whatever is put in front of them regardless of content. A participant who reflexively agrees with every statement produces a mediocre, self-cancelling score once the negative items are converted, rather than an artificially inflated one; a participant answering thoughtfully produces a real signal either way. Losing that alternation (for instance, rewording every item to be positive "for clarity") quietly breaks the instrument's validity, which is why SUS is always administered with the original wording, not a paraphrased version.


Think-Aloud Protocol

The think-aloud protocol asks a participant to verbalize their thoughts continuously while attempting a task — what they're looking at, what they expect a control to do before they interact with it, what confuses them, what they're about to try next. It's the technique that turns a usability-test recording from "here's what the participant did" into "here's why they did it," which is exactly the gap between a symptom and a diagnosis: watching someone click the wrong button tells you a mistake happened; hearing them say "I assumed this was where I'd find my account settings" tells you the mental model that produced the mistake, which is what you actually need to fix the design.

Common facilitation pitfalls, all of which quietly replace the participant's real reasoning with the facilitator's assumptions:

  • Leading the participant. "Do you see the button there in the corner?" doesn't observe whether the participant would have found the button — it tells them where to look, contaminating the very thing being tested. The neutral version is silence, or at most a content-free prompt: "keep talking me through what you're doing."
  • Filling silence for them. Participants often go quiet while concentrating, and an uncomfortable facilitator jumps in to fill the gap — but that silence is frequently the most diagnostic moment in the session (confusion, hesitation, a decision being weighed), and talking over it destroys the data. The correct response to silence is usually a neutral re-prompt ("what are you thinking right now?"), not narrating for them or moving the task along.
  • Asking about intent instead of watching behavior. "Would you click this?" asks the participant to predict their own future behavior, which is exactly the kind of self-report user-research-methods warns is unreliable — people are poor predictors of their own actions, especially under artificial, low-stakes test conditions where the consequence of a wrong guess is nothing. The fix is procedural: don't ask, present the actual task and watch whether they actually click it. What people say they'd do and what they actually do diverge often enough that the question itself is close to worthless next to direct observation.
Pitfall Weak prompt Neutral alternative
Leading "Do you see the button there in the corner?" "Keep talking me through what you're doing."
Filling silence (facilitator jumps in to explain) "What are you thinking right now?"
Asking intent "Would you click this?" Present the task, watch what they actually click

The unifying pattern across all three pitfalls: each one substitutes something easier to get (a confirming nod, filled silence, a stated intention) for the harder, more valuable signal the method exists to capture — unprompted behavior and the reasoning behind it. A facilitator who catches themselves about to lead, fill silence, or ask about intent has a simple fallback in all three cases: stop talking, and let the next few seconds of silence do the work instead.


A/B Testing vs. Usability Testing

Both claim to "test" a design, and interviewers use the phrase "how would you validate this" partly to see whether a candidate reaches for the right one.

A/B testing randomly splits real traffic between two or more variants and measures the difference in a target metric (conversion, click-through, retention) with statistical confidence. It's quantitative and behavioral in the user-research-methods framework, and it answers which variant performs better on a defined metric. It requires two things usability testing doesn't: an already-live product with real traffic, and enough volume flowing through the tested flow to reach statistical significance in a reasonable window — on a low-traffic page, an A/B test can take months to produce a trustworthy result.

Usability testing is qualitative, runs on a handful of participants, and answers why something isn't working — the reasoning, confusion, and mental-model mismatches behind a number, not just the number itself.

Usability Testing A/B Testing
Data type Qualitative Quantitative
Sample size Small (5-8 per round) Large (enough for statistical power)
Answers Why isn't this working Which variant performs better
Requires live traffic? No Yes, and enough of it
Right for Pre-launch validation, diagnosing a bad metric, early concepts with no traffic yet Optimizing an already-live, high-traffic flow with two clear variants

When each is the right tool, stated plainly: usability testing is the right call before launch (there's no traffic to A/B test yet) or whenever the question is "why is this metric bad" (a number alone doesn't explain itself). A/B testing is the right call once a flow is live with real volume and you have two well-defined variants competing for a statistically confident answer to "which one wins." The two are often sequential, not competing: usability testing surfaces why checkout abandonment is high and generates a specific fix hypothesis; A/B testing then measures whether that specific fix actually moves the metric at scale. Proposing an A/B test before you understand why the current design is failing risks testing the wrong fix with real user traffic — usability testing is what narrows down what's worth A/B testing in the first place.

Worked example. A checkout flow has a 60% completion rate and stakeholders want to "A/B test our way to a better number." Before any variant exists, five moderated usability-testing sessions find that most participants abandon at a shipping-cost reveal that appears only after they've entered payment details — they feel misled, not merely slowed down. That finding produces a specific, testable hypothesis (surface shipping cost earlier in the flow), which is now exactly the kind of two-variant, high-traffic question A/B testing is built to answer with statistical confidence. Running the A/B test first, with no usability-testing input, would have left the team guessing at which of a dozen plausible checkout changes to test, on real traffic, with real revenue at stake per losing variant. case-study-redesign-a-checkout-flow, later in this track, works through a fuller version of exactly this scenario end to end.


Common Facilitation Pitfalls Across Every Method

Three failure patterns cut across every method in this subject, and interviewers listen for whether a candidate names them unprompted:

  • Recruiting bias. Testing with people who are already familiar with the product — internal employees, power users, colleagues — under-represents the confusion a genuine first-time user would hit, because the participant pool has already adapted to the product's quirks. A usability test that only recruits people who "get it" will systematically under-report real problems, regardless of how well the session itself is facilitated.
  • Sample-size confusion. Two opposite mistakes, both rooted in the same misunderstanding: treating a 5-user qualitative usability-testing finding as if it were statistically significant ("4 out of 5 users struggled" is a strong qualitative signal, not a percentage that generalizes to the full user base), and the reverse — over-investing in a large qualitative sample (30+ moderated sessions) that adds little proportional insight beyond what user-research-methods's 5-user rule already predicts you'll find in the first round. Qualitative usability testing and quantitative A/B testing answer different kinds of questions and need different-shaped evidence; borrowing one method's rigor claims for the other's sample size is the tell.
  • Confirmation bias in interpreting ambiguous behavior. A participant hesitating before a click is genuinely ambiguous — it could mean confusion, or it could mean careful reading. A facilitator who already expects the design to fail will read hesitation as confirmation of a problem; one who expects it to succeed will read the same hesitation as thoroughness. The mitigation, same as in user-research-methods's synthesis-bias discussion, is process rather than willpower: note the ambiguous behavior as ambiguous, and resolve it with the participant's own think-aloud narration rather than the facilitator's prior expectation of the outcome.

Why this section matters as a closer: an interviewer who has already heard a candidate correctly distinguish moderated from unmoderated, heuristic evaluation from usability testing, and usability testing from A/B testing is testing one more thing with a follow-up like "what could go wrong with your own study" — whether the candidate can turn the same critical eye on their own method, not just on the product being tested. Naming these three biases unprompted, before being asked, is usually the strongest signal in the whole conversation.


Quick Reference: Matching the Method to the Question

Every method in this subject answers a specific question and fails when used to answer a different one. This is the table to reconstruct out loud when an interviewer asks "how would you validate this":

Question being asked Method Data shape
"Does this violate known usability principles?" Heuristic evaluation Fast, expert judgment against a checklist
"Can real users actually complete this task?" Usability testing (moderated or unmoderated) Small qualitative sample, task-based
"How efficiently and cleanly did they do it?" Completion rate, time-on-task, error rate together Behavioral, per-task
"How did it feel to use, subjectively?" SUS Attitudinal, standardized 0-100 score
"Why did they get stuck?" Think-aloud protocol layered on any of the above Qualitative reasoning, not just outcome
"Which of two live variants performs better?" A/B testing Quantitative, large sample, statistical confidence

No single row in this table is a complete answer to "how would you validate this design" on its own — the interview-strong move is picking the row that matches the actual question, naming why the others don't fit as well, and being ready to explain what you'd do next once that method's findings come back.

Ready to test your knowledge?

Practice questions

We use cookies for product analytics to improve OmniAtlas. See our Privacy Policy.