Measuring Quality

Testing finds problems. Measuring tells you whether the version you just shipped is better than the one before it — a different question, answered by different methods.

Two methods are worth your time. Benchmarking turns the experience into numbers you can compare across releases and against rivals. Expert review puts trained eyes over the interface against known principles, with nobody to recruit and no sessions to schedule.

Neither replaces watching a real person struggle. Both are cheaper, faster and repeatable, which means you can run them often enough to see a trend. A number on its own means very little. It means something beside another number — the last version, a rival, a target. Choose the comparison before you measure, not after you have seen the result.

UX Benchmarking

Pick measures that survive a redesign. Task success rate, time on task, error count, and a standardised questionnaire — SUS or UMUX-Lite — give you a number per task and one per product.

Then hold everything else still. Same tasks, same wording, same participant profile, same sample size — and make that sample far bigger than a qualitative round. Five people will surface problems; they will not give you a stable average. Change the tasks between rounds and you have two unrelated numbers and no benchmark.

Compare against three things: the previous version, your nearest competitor, and a target set in advance. Quarterly is usually enough. Monthly becomes a reporting job that nobody reads.

Publish the number when it drops, too. A benchmark reported only when it flatters you is a marketing asset, not a measurement.

Expert Review

Three to five experienced people walk the interface separately against a fixed list of heuristics, then pool their findings — one reviewer working alone misses too much. The method's other name is heuristic evaluation. Nielsen's ten are the usual list; visible system status, a match to the real world and recognition rather than recall are three of them. Every issue gets a location, the heuristic it breaks, and a severity from cosmetic to catastrophic.

It is cheap. A day of someone's time clears out the obvious rubbish — the unlabelled icon, the error message that names a database table — so your user sessions spend their expensive minutes on what only real users reveal.

It finds problems. The severity ratings are a trained guess at which ones hurt — a queue to check, not a verdict. An expert predicts; a user demonstrates.

No questions on this lesson yet. Highlight a passage to ask about it, or use Ask a question.