Skip to content
Design & UXAll articles

Usability Testing: Types, Metrics, and How to Run a Test

Usability testing is the cheapest way to find out your product confuses people before launch does it for you. Here's how to run one: the types and methods, the metrics that matter, how many users you actually need, and what to do with what you learn.

Occasional field notes on building software, no spam

Idealogic: the usability testing process

Usability testing is the practice of watching real people attempt real tasks in your product so you can see exactly where they stumble. Not what they say they would do, not what they think of the colors. What they actually do when you hand them a goal and get out of the way. A single usability test session is the cheapest, bluntest way to find out your interface confuses people, and it is a lot less painful to learn that in a test than from a churn chart three months after launch.

This guide covers the parts that matter when you run one for real: the types of usability testing and when each fits, the specialized techniques worth knowing, how to run a test step by step, how many users you genuinely need (the honest version), the metrics that actually tell you something, and how to turn a pile of observations into fixes your team will build. It pairs well with a broader UX audit, where testing is one method among several.

The short version

  • Usability testing means watching real people attempt real tasks in your product so you can see where they stumble, not asking what they think of the design.
  • The main types split by moderation (moderated or unmoderated) and location (remote or in-person), with guerrilla testing as the fast, informal option. Most teams blend a few.
  • Beyond task-based sessions, techniques like tree testing, first-click testing, the five-second test, and card sorting each answer a sharper, narrower question.
  • The working rule of thumb is five users per round, which the Nielsen Norman Group found uncovers about 85% of the usability problems in a given flow.
  • Running a test is a six-step loop: set a goal, write neutral tasks, recruit five representative users, observe in silence, analyze recurring friction, then ship prioritized fixes.
  • Track a few real metrics (task success rate, time on task, error rate, plus a SUS or SEQ score) instead of chasing every number you can capture.
  • Rank findings by impact times frequency so the changes that block people from the core job get fixed first, then retest to confirm.

What usability testing is

Usability testing is observation, not opinion-gathering. You give someone a task ("find a plan that fits a team of ten and start a trial"), then you watch, mostly in silence, and note every place they pause, backtrack, or pick the wrong thing. The friction you see is the data. What people tell you they think afterward is interesting, but it is secondary, because users are famously bad at predicting their own behavior.

That is the trait that makes it powerful. A stakeholder can argue all day that a button is obvious; one person missing it twice in a row ends the argument. Testing replaces taste with evidence, which is why it sits at the center of any serious user-centered design practice rather than off to the side as a nice-to-have.

One caveat worth saying early: usability testing tells you whether something works and where it breaks. It does not tell you what people want built, and it will not replace the discovery research you do before there is anything to test. Different job, different tool.

Types of usability testing

There is no single right method. There is the one that fits your question, your timeline, and your budget. The common types of usability testing sort along two axes, moderation and location, plus a quick-and-dirty variant. Here is how they compare.

TypeWhat it is forEffort
ModeratedUnderstanding why, probing reasoning on rough or early designsHigh (a facilitator per session)
UnmoderatedConfirming a pattern at scale once you know what to checkLow (a tool does the running)
GuerrillaA fast gut-check, cheap and rough, no recruitingVery low (whoever is nearby)
RemoteReaching users wherever they are, over video or a testing toolMedium (setup, then it scales)
In-personWatching body language and subtle hesitation up closeHigh (scheduling and a room)

A few opinions on that table. Moderated and unmoderated are not rivals; they are a sequence. You run a handful of moderated sessions early to understand the reasons behind the friction, then switch to unmoderated tests to check whether the fix holds across more people. Guerrilla testing gets unfairly dismissed as sloppy, but for an early-stage product it is often the highest return on an afternoon you will find. And "remote vs in-person" matters less than it used to. Remote moderated testing gets you most of what a room would, minus the travel.

The split that confuses people most is moderated versus unmoderated, so it is worth being precise: moderated means a human runs the session live and can ask "what made you click there?"; unmoderated means the user is alone with a tool and you watch the recording later. One trades speed for depth, the other trades depth for scale.

Specialized usability testing techniques

Task-based sessions are the backbone, but some questions are narrow enough to deserve a sharper instrument. These techniques are quick, often unmoderated, and each isolates one thing.

TechniqueWhat it answersWhen to reach for it
Tree testingCan people find things in your navigation, stripped of visual designBefore you commit to a menu or IA structure
First-click testingDo people click the right thing first on a taskValidating a single screen or entry point
Five-second testWhat sticks after a brief glance: message, hierarchy, brandLanding pages and first impressions
Card sortingHow users expect content to be grouped and labeledDesigning or reworking information architecture

Tree testing and card sorting are the two you reach for when the problem is structure rather than surface. Card sorting asks users to group and label content the way they expect it, and tree testing then checks whether they can actually navigate the structure you built from that input. Run them in that order and you design an information architecture around how people think, not how your org chart is shaped.

First-click testing earns its keep because of one stubborn finding: the first click predicts almost everything. Jeff Sauro's first-click analysis at MeasuringU reports that when a user's first click is down the right path they eventually succeed about 87% of the time, and when it is down the wrong path only about 46% do. Get the first click right and the rest of the task tends to fall into place, which is why a five-minute first-click test on a key screen is worth so much more than its cost.

The five-second test (sometimes written the 5-second test) sits at the opposite end: it measures the instant impression. Show a screen for five seconds, take it away, and ask what someone remembers. If nobody can say what the page is for, no amount of careful copy further down will save it. It is the cheapest way to pressure-test a landing page or a dashboard design before you invest in the details.

How to run a usability test, step by step

The backbone of a good test is the same whether you are checking one screen or a whole onboarding flow. The order is what keeps it honest: skip the goal and you will collect observations you cannot rank.

The usability testing process in five steps, left to right. Plan: set a goal and write realistic tasks; Recruit: find about five people who match your real users; Test: run the sessions and stay quiet while you observe; Analyze: review recordings and group recurring friction; Fix: turn findings into prioritized design changes. An arrow runs across all five, and a dotted return arrow loops from Fix back to Plan, showing that testing repeats in rounds.
Five steps, run in rounds: each pass should leave you with fixes, not just notes
  1. Set a goal. Decide the one thing this round is meant to answer: "can a new user reach the first value moment without help?" A vague goal produces a vague report. This is also where you decide what not to test, which keeps the session short enough that people stay sharp.
  2. Write tasks. Turn the goal into realistic tasks phrased as outcomes, not instructions. "Buy a gift card for a friend" is a task; "click the Gift Cards menu item" is just narrating your own UI back at the user. Neutral wording is the whole game. The moment you hint at the path, you have stopped measuring anything.
  3. Recruit. Find people who resemble your actual users, not your colleagues. Five or so per group is plenty (more on that below). Screen for the behavior that matters. If you are testing a fintech flow, "has opened a bank account online" beats "is between 25 and 40."
  4. Run the sessions. Hand over the task and go quiet. The hardest discipline in this whole process is not rescuing someone the second they struggle. Their struggle is the finding. Ask them to think aloud (the think-aloud protocol), and when they ask "is this right?", gently bounce it back: "what would you do if I were not here?"
  5. Analyze. Watch the recordings and group what repeats. One person fumbling a button is noise; four people fumbling the same button is a fix. Tag each issue by how badly it hurt and how often it showed up, so the ranking writes itself.
  6. Fix. Translate the findings into specific design changes, in priority order. A test that ends in a tidy report nobody builds from was theater. The output you want is a short list an engineer or designer can pick up on Monday.

The tasks are where most tests are won or lost. A few examples of the neutral phrasing you are aiming for: "Find a phone plan for two people and add it to your cart." "You just got paid; move 200 dollars into savings." "Cancel your subscription." Each names a goal a real person would have and says nothing about which button to press. Most of the calendar, by the way, lands on planning and analysis, not the sessions themselves, which usually surprises first-timers.

How many users do you actually need

For most qualitative tests, around five users per round is the working answer. That is not a shortcut, it is well-supported. The Nielsen Norman Group's long-standing guideline, from Jakob Nielsen's research on testing with five users, is that roughly five users uncover about 85% of the usability problems in a given flow. The model behind it is simple: if a single user finds a typical 31% of the problems, five independent users cover most of what is there, and each extra user past that finds less and less you did not already know. Run a sixth, a seventh, and you mostly watch the same issues repeat.

The honest nuance people skip: that 85% holds for one type of user on one set of tasks. If your product serves two genuinely different audiences (say, patients and clinicians in a healthtech tool), you need five of each, because they will trip over different things. And if you have crossed from "what is broken?" into "what percentage convert?", you have left usability testing for quantitative research, where the numbers get much bigger.

Five users will not tell you everything. They will tell you the things that are wrong enough to matter, and that is almost always the list worth acting on first.

The bigger lever is not sample size, it is frequency. Five users this month, fix what you find, then five more next month beats twenty users once a year. Small, repeated rounds catch regressions early and keep the product honest as it grows.

Usability testing metrics that matter

Watching people is qualitative, but you still want numbers to track progress across rounds and to settle debates with people who were not in the room. The trick is picking a few metrics that map to something real. The international standard for usability, ISO 9241-11, frames it as three dimensions, effectiveness, efficiency, and satisfaction, and the practical metrics fall neatly under them.

MetricWhat it measuresHow to read it
Task success rateEffectiveness: the share of users who finish the taskTrack per task; the trend across rounds matters more than any single number
Time on taskEfficiency: how long finishing takesCompare against your own baseline, not an absolute
Error rateSlips and wrong turns per taskLower is better; clustered errors point straight at the fix
System Usability Scale (SUS)Perceived usability, ten questions scored 0 to 100Above the benchmark is good news, below it is a flag
Single Ease Question (SEQ)Perceived difficulty of one task, rated 1 to 7Ask it right after each task, while it is fresh

A few notes on reading these. Task success rate, time on task, and error rate are the behavioral core, and Nielsen's usability metrics treats success, time, errors, and satisfaction as the base set. They are most useful as a trend: a checkout that climbs from 60% to 90% success across three rounds is the story, not the raw figure.

For perceived usability, two questionnaires do most of the work. The System Usability Scale, created by John Brooke in 1986, is ten questions that resolve to a single 0 to 100 score; it is not a percentage, so you read it against the benchmark, and the average across 500 studies is about 68. The Single Ease Question is even lighter: one question after each task on a seven-point scale, where the typical average lands around 5.5. Neither replaces watching people, but both give you a number you can defend and compare over time.

Common usability testing mistakes

Most tests that fail do not fail in the session; they fail in how they were set up or read. The recurring mistakes are worth naming so you can catch them before they cost you a round.

  • Leading tasks. The fastest way to ruin a test is to write the answer into the task. "Click Settings, then Billing" measures nothing. Phrase the goal, never the path.
  • Rescuing people. The instinct to help the second someone struggles is strong, and it deletes your most valuable finding. Sit on your hands and let the silence do its work.
  • Testing only colleagues. Your teammates know the product too well to get lost the way a real user will. Recruit for behavior, not proximity.
  • One giant study. A single twenty-person test before launch feels rigorous and arrives too late to change anything. Trade it for four rounds of five, spread across the build.
  • A report nobody builds from. Fifty color-coded observations with no ranking is a document, not a plan. If a finding does not point to a change, it is not done.

Turning findings into fixes

A test is only worth running if it ends in changes. The trap is the beautiful findings deck: fifty observations, color-coded, that nobody can act on because nothing is ranked. Skip it. What a team can actually use is a short, prioritized list where each item carries the problem, the evidence, and a rough sense of effort.

Rank by impact times frequency. A confusing label on a settings page four users glanced at is not the same as a checkout button two of five could not find. Lead with the things that block people from the core job, tag each with rough effort, and the "major, low-effort" items get fixed this sprint instead of debated next quarter. The same severity thinking we use in a full UX audit applies here. Testing just feeds it sharper evidence.

Then close the loop. The fix is a hypothesis until the next round of testing confirms it, which is exactly why the process bends back on itself. Usability testing is not a gate you pass once before launch; it is a habit that runs through the whole UX design process, most usefully right after you can tell UI from UX and start treating them as separate problems to test.

Want to know where your product loses people? Test it with real users
Talk to our design team

Frequently asked questions

  • Usability testing is watching real people try to complete real tasks in your product, so you can see where they hesitate, get stuck, or give up. It is not asking people what they think of a design. It is observing what they actually do with it. The output is a ranked list of friction points and the evidence behind each one, which is far more reliable than opinions in a meeting.

  • The main types of usability testing split along two axes: who runs the session and where it happens. Moderated tests have a facilitator running things live; unmoderated tests let users work alone through a tool. Remote tests reach people wherever they are; in-person tests put everyone in one room. Guerrilla testing is the quick, informal version. Most teams mix a few: moderated sessions to understand the why, then unmoderated tests to confirm the pattern at scale.

  • Around five users per round is the working number for most qualitative tests. The Nielsen Norman Group's well-known finding is that roughly five users surface about 85% of the usability problems in a given flow, and you get more value from running several small rounds than one large one. If you are testing very different user groups or chasing statistical numbers rather than problems, you will want more, but for finding what is broken, five is enough to start fixing.

  • In moderated testing a facilitator runs the session live, asks follow-up questions, and probes the reasons behind what they see. In unmoderated testing users work through tasks alone using a tool, and you review the recordings afterward. Moderated tells you why something happened and is better for early, rough designs; unmoderated is faster, cheaper, and easier to run at scale once you know which tasks to check.

  • Run a usability test in six steps: set a clear goal, write realistic tasks tied to that goal, recruit five or so people who match your real users, run the sessions while staying quiet and observing, analyze the recordings to spot recurring friction, then turn the findings into prioritized design fixes. The discipline that matters most is writing neutral tasks and resisting the urge to help people while they struggle.

  • Track a small mix of behavioral and attitudinal metrics. Task success rate tells you whether people can finish at all, time on task tells you how hard the finishing was, and error rate shows where they slip. For perceived usability, the System Usability Scale (SUS) gives a 0 to 100 score, and the Single Ease Question rates one task from 1 to 7. A handful of these, tracked round over round, beats a dashboard full of numbers nobody reads.

  • Usability testing is qualitative and small: you watch a handful of people use one design and learn why it confuses them. A/B testing is quantitative and large: you show two live variants to thousands of real users and measure which performs better on a metric like conversion. Usability testing explains the why and works before launch; A/B testing measures the what and needs live traffic. They answer different questions, and mature teams use both.

  • Run it early and often, not once before launch. Test rough prototypes to catch structural problems while they are cheap to fix, test again after major design changes, and keep testing after release to catch regressions as the product grows. Small, repeated rounds tied to whatever you are shipping next beat one big study a year, which usually lands too late to change anything.