Methods

How I work.

The same few methods run through every case. They are stated once here so the case studies can stay short.

Brock Dubbels placing a sticky note on a glass wall while colleagues cluster notes around him
D2P2

Describe → Diagnose → Predict → Prescribe

A/B tests tell you what happened, after customers have paid for it. This sequence is how you learn why, before they do.

  1. 1

    Describe

    What is actually happening, and to whom?

    Go where the work happens and watch it before asking about it. Behavior in the aisle is more reliable than memory in a survey.

    • Contextual inquiry & shop-alongs
    • Think-aloud / RITE protocol analysis
    • Longitudinal & diary studies
    • Faultline (counterfactual) interviews
    • Mental-models elicitation (Show, Not Tell)
  2. 2

    Diagnose

    What is working, what is not, and why?

    Turn observations into something a team can argue with: a count, a map, a funnel, a benchmark.

    • Thematic analysis
    • Funnel modeling (SQL)
    • Screen & container audits
    • Usability benchmarking (SUS)
    • Customer-jobs matrices & story maps
  3. 3

    Predict

    What happens if we change it — before we build it?

    Estimate the effect with the cheapest evidence that can answer the question, then spend more only where it matters.

    • Wizard-of-Oz & paper prototypes
    • A/B, A/B/C & multivariate tests
    • Monte Carlo simulation
    • Critical path & Markov-chain analysis
    • Calibrated multi-rater evaluation
    • Mixed-effects modeling
  4. 4

    Prescribe

    What is the smallest change that moves the outcome?

    Hand the team a decision, not a deck: what to build first, how to know it worked, and when to stop.

    • PRFAQs & product requirements
    • Phased implementation plans
    • KPI review structures
    • Design direction & interaction models
Why this order

Learn before customers pay for the lesson.

The same change costs ten times more at each step down the line. So answer the question at the cheapest step that can answer it.

  1. $1

    Pre-development · Show-Not-Tell prototype

    Find out what people expect before anything is built. A handful of sessions; no customer meets an unfinished product.

  2. $10

    During development · MVP test

    Build enough to ship to a sample. Weeks of engineering, and real customers meet the unfinished version.

  3. $100

    After hand-off · production change

    Ship, measure, roll back. Every team touches it, and every customer lives with it until the lagging numbers come in.

Common practice across the industry, as I presented it to Marqeta’s product leadership in 2022. A rule of thumb for orders of magnitude, not a measured ratio.

Leading metrics, not lagging ones

Exercise minutes lead weight loss; résumés sent lead interview requests. In a product, customer effort leads, satisfaction and recommendation (CSAT, NPS) follow, and revenue lags. Perceived ease of use alone explains about 37–39% of the variation in likelihood to recommend (Sauro, MeasuringU). When effort improves, the survey scores catch up later, and the surveys themselves arrive after an escalation, mostly from unhappy customers. Revenue reports on decisions made months ago, after the customers who left are gone. Effort can be observed at the project level, in a session, before the build. D2P2 measures the leading end and ties it to the lagging end, so a team can show value before the quarter closes. Not every useful research measure is a revenue measure, but every useful product measure should have a defensible path to revenue.

Two-way doors still cost customer intent

A two-way door is a decision you can reverse. The code can be rolled back; the customer’s visit cannot. Every experiment in production spends real intent: the shopper who met the weaker variant and left does not come back for the stronger one. Production tests are for confirming what is already understood, not for finding out.

Signature method · since 2005

Show, Not Tell

Never explain the screen. What people do before anyone hints is the mental model they walked in with, not the one we are trying to install.

Expectation → encounter → prediction → synthesis → consequential action

Step 0 · before anything is built
  1. What decision?What are we trying to decide?
  2. What evidence?What would let us make that decision?
  3. What consequence?What outcome should the decision eventually affect?

A prototype is an instrument for producing evidence about a decision. Without a named decision, a session only collects reactions.

  1. 1

    Describe

    Build the model

    Goals, expectations, prior knowledge and context, not just the task sequence or a frozen persona. Scenario archetypes cross disposition with circumstance to show how the same person meets a different problem.

  2. 2

    Diagnose

    Find the discrepancy

    The gap between the experience people expect and the one they encounter.

  3. 3

    Predict

    Make hypotheses encounterable

    Each design idea becomes a visual hypothesis: characteristic → experience hypothesis → prototype → evidence.

  4. 4

    Prescribe

    Commit with evidence

    Build from what held across variants, and name the behavior it should change.

SNT began in a training game I designed for Benedictine Health Services. A trainee is shown a scenario and asked one question, what do you do here?, with nothing behind it to prompt or correct.

The method has one governing constraint: the environment poses the hypothesis; the facilitator never does. If a screen tests whether someone distinguishes urgency from severity, that distinction lives in what the screen shows, not in a question that names it. The moment a facilitator says “notice how this involves both urgency and severity,” the session stops being mental-models elicitation and becomes usability testing: a different instrument, measuring a different thing.

The checkCould a stranger have run this session from a script, moving the participant from screen to screen, without ever describing what was on them? If not, I have already told.
Core questions, every session
  • What did you expect?
  • What is this for?
  • What would you do first?
  • What do you expect to happen next?

The full probe sequence, from the Benedictine game

  1. What do you do here?The first, silent question. What they do before any hint arrives is the mental model they walked in with.
  2. Tell me moreRefuses a vague answer without supplying the missing content.
  3. What will happen?Captures a prediction before any outcome is visible, so there is nothing yet to reason backward from.
  4. Is this what you want?Checks the stated action against the stated goal, where a gap between intention and belief usually shows first.
  5. What will happen if it does not work?Forces a contingency model into view, or reveals that none exists.
  6. Time, quality, preferenceTrade-off questions that locate where people are actually willing to spend cost, which is rarely where they say they would.
  7. SynthesisCommit to one answer across all of it. This is where a borrowed or half-formed model tends to break.

It guards against performing for the researcher

A participant told what is being tested performs toward the hypothesis instead of revealing one: demand characteristics, in Orne’s original sense (1962).

It guards against the plausible story told afterward

Asked to explain a choice after the fact, people often construct a plausible account rather than an accurate one (Nisbett & Wilson, 1977). Capturing the prediction before the reveal, and the response as it happens, leaves nothing to confabulate about later.

Style studies find a territory, not a winner

Deliberately different treatments, A through D, are probes into a design space, not candidates for the final product. Synthesis locates a region (too corporate ← viable → too expressive; too sparse ← viable → too dense) and names what held across variants: one variant’s density, another’s typography, a third’s restraint. Read the variants for convergence, divergence, invariants and reorganization, then build n+1: combine the characteristics people repeatedly recognize and use. That region becomes the containers, components, type and spacing rules of the design system.

Content studies test whether the proposition produces action

Once the container is stable, the question changes: does the content create enough understanding and value for the person to continue? Three observable states: I don’t get it (no model has formed), I get it (intelligible, but understanding alone does not create behavior), I want this (enough relevance to take the next consequential step). Comprehension, relevance and motivation are explanatory variables; action is the result. Follow the consequence: experience → behavior → conversion → revenue.

The same discipline has carried across very different instruments: customer-experience work at Amazon Style, platform research at Nielsen, and a Wizard-of-Oz controller for Maetri, where the sequence compresses into two fields, what the participant predicts and what they actually say or do, logged separately and never narrated between. The interface changes. The discipline does not.

Signature approaches

Design for Curiosity

An approach to activation and interaction developed in the customer-experience work on Amazon Style: make exploration rewarding, so people learn the product by wanting to see what happens next rather than by being instructed.

Faultline interviewing

A counterfactual interview method — “methods of imagination” (Dubbels, 2020). Asking how things might have gone differently (“if only…”) surfaces the constraints and alternatives people actually weigh, which direct questions miss.

Multi-rater evaluation

Calibrated observers, agreement statistics (ICC, Krippendorff’s α) and mixed-effects models with random intercepts and slopes by rater, separating the thing being evaluated from the person evaluating it. The same discipline applies to rating AI output.

Evidence qualification

Every claim carries its status — measured, client data, projected, modelled, target or demonstrated. Building on Itamar Gilad’s Confidence Meter and multi-trait, multi-method measurement, it is the rule behind this site and behind Maetri.

Tolerance for complexity

After Herbert Simon: solving a problem means representing it so the solution becomes transparent. Much of product work is finding the representation a team can act on.

Toolkit
Qualitative
  • Cognitive ethnography
  • Contextual inquiry
  • Think-aloud / RITE
  • Diary & longitudinal studies
  • Thematic analysis
Quantitative
  • Mixed-effects models
  • Inter-rater reliability
  • Factorial designs
  • Monte Carlo simulation
  • Critical path analysis
  • Markov chains
  • A/B & multivariate testing
  • SUS benchmarking
AI evaluation
  • Human-in-the-loop evaluation
  • Agentic AI evaluation
  • Semantic evaluation
  • Multimodal signal interpretation
Tools
  • R
  • Python
  • SPSS
  • SQL
  • Figma
  • UserTesting
  • Contentsquare

See the methods at work in the featured cases, or in the Equinix SmartView mixed-methods study.