Playbook ┬╖ 6 minute read
How to Run User Research for AI Products
AI user research works when it studies the current task before proposing a system, tests with realistic failures rather than curated successes, measures whether users' trust matches actual reliability, and runs long enough for the novelty effect to fade before conclusions are drawn.
AI user research has traps conventional research does not: novelty distorts early enthusiasm, trust is poorly calibrated, and demonstrations hide the failures that determine whether a product is safe. This playbook covers avoiding them, drawing on FISTA Solutions' web and mobile work.
When is this worth doing?
Before designing an AI feature, and again after a prototype exists but before it is built properly.
The first round is about the task; the second is about the interaction. Skipping the first produces products that automate something nobody found difficult.
What does the sequence look like?
| Step | Purpose |
|---|---|
| 1. Study the current task | Before proposing anything |
| 2. Find the real difficulty | Which part is actually hard |
| 3. Test with realistic failures | Not curated successes |
| 4. Measure trust calibration | Not satisfaction |
| 5. Run long enough | Past the novelty period |
| 6. Include the sceptics | They generalise better |
Step 1 тАФ Study the current task first
Watch people do the work before proposing a system. Where the time goes, which parts are difficult, what gets skipped under pressure, and what the workarounds are.
This consistently surprises. The part a product team assumes is hard is frequently routine, and the real difficulty is somewhere nobody looked тАФ usually in gathering information rather than in producing output.
Observation beats interviews here. People describe their process as they believe it to be, and what they actually do differs in ways that matter.
Step 2 тАФ Find where the difficulty actually is
Separate the mechanical parts from the judgement parts.
Mechanical work тАФ finding, summarising, formatting, checking тАФ is where AI helps most and where automation is safest. Judgement work is where it helps least and where automating it causes the most harm.
Products that target the judgement because it looks more impressive produce resistance from the people doing the work, who correctly identify that their expertise is being replaced badly.
Step 3 тАФ Test with realistic failures
Include wrong outputs in the sessions, at a realistic rate, and watch what users do.
Do they notice? Do they check? Do they accept it? How do they recover? Those behaviours determine whether the product is safe, and a session using only good outputs answers none of them.
This is the step most often skipped because it feels like sabotaging your own product. It is the most informative part of the research.
Step 4 тАФ Measure trust calibration
Ask users how confident they are in each output and compare against whether it was correct.
Over-trust тАФ high confidence in wrong answers тАФ is a design problem. It usually means the interface presents everything with equal certainty, giving users no signal about when to check.
Under-trust means users check everything, which removes the benefit. That is also a design problem, usually caused by earlier visible failures with no explanation. Both are fixable through how outputs are presented.
Step 5 тАФ Run long enough for novelty to fade
Revisit participants after two to four weeks of real use.
First-session enthusiasm is not evidence. People respond positively to new things, and the response decays as the novelty does. The second conversation, after the tool has become ordinary, is where the honest assessment arrives.
Watch usage data alongside. Participants who describe a tool as useful and stopped opening it three weeks ago are telling you something the interview did not.
Step 6 тАФ Include the people who are sceptical
Recruit across the population that will use the system, including those who did not volunteer.
Research conducted with enthusiasts produces findings that do not generalise. The sceptics frequently have specific reasons тАФ they have seen the system be wrong in their domain тАФ and those reasons are the most useful input available.
Include the people whose work changes most. They know the edge cases, and they have the strongest incentive to tell you where the design fails.
How do you research something that does not exist yet?
With realistic prototypes rather than descriptions. Wizard-of-Oz studies, where a person produces the outputs, work well and are cheap.
Descriptions produce reactions to the idea, which are uninformative. A participant working with realistic outputs тАФ including the wrong ones тАФ produces behaviour you can observe.
The effort is worth it. Research on descriptions reliably over-predicts adoption.
What about research with customers rather than staff?
The same principles apply with an added constraint: customers rarely tolerate being studied for long, and their tolerance for failure is lower.
Keep sessions short, use realistic scenarios drawn from their actual interactions, and be explicit that they are seeing something in development. Customers who feel experimented on without being told respond badly and tell others.
Who needs to be involved?
A researcher, someone from the product team who will act on the findings, and access to the people who do the work.
Research conducted without product involvement produces a report. Product people who watched the sessions change their designs.
How long does it take?
Two to four weeks for a round, plus a follow-up after a few weeks of use. Compressing it removes the follow-up, which is where the honest findings are.
What are the common failure modes?
Proposing before observing. Targeting judgement rather than mechanical work. Curated demonstrations. Measuring satisfaction. Single sessions. And recruiting only enthusiasts.
How do you know it worked?
A clear account of where the difficulty actually is, observed behaviour when the system fails, trust that matches reliability, and findings that hold up at the follow-up.
What does it cost?
Mostly people's time rather than tooling. The expensive version is the one that stalls halfway and leaves the organisation with neither the old state nor the new one, which is why a narrow first pass beats a comprehensive plan nobody finishes.
Budget the work as an operated change rather than a project with an end date, because most of these need a maintenance tail. See AI total cost of ownership.
What should you do first?
Watch three people do the task you are considering automating, without proposing anything. The most useful finding usually arrives in the first hour.
How FISTA Solutions helps
FISTA Solutions runs this work alongside client teams rather than around them: research that studies the current task before proposing a system, sessions that include realistic failures so trust calibration can be observed, evidence produced as the work proceeds, and handover that leaves your people able to continue without us. Delivery runs through AI agents, AI enablement, and forward deployed engineers. The record is 150+ projects for 50+ companies across 12+ countries, with 47% average efficiency gains where measured.
To run this with support, message FISTA on WhatsApp, or read how to design an AI feedback loop.
Share-ready article cover
Download the generated social format.
Clear answers
Questions raised by this field note.
Straightforward guidance for evaluating scope, fit, and the next step.
01What is different about researching AI products?
Novelty effects and trust calibration. Users respond enthusiastically to something new regardless of usefulness, and they either over-trust or dismiss outputs rather than calibrating, and both distort early research findings.
02Why test with failures?
Because how users handle a wrong answer determines whether the product is safe. A session using only successful outputs tells you the interface is pleasant and nothing about what happens when it is wrong.
03What is trust calibration?
Whether users' confidence in the system matches its actual reliability. Over-trust means accepting wrong answers; under-trust means checking everything and gaining nothing. Both are design problems rather than user problems.
04How do you handle novelty effects?
By running longer studies and revisiting participants after a few weeks. Enthusiasm in a first session is not evidence of value, and the second conversation is usually considerably more informative.
05Who should be included?
The people whose work changes most, including those who are sceptical. Research conducted only with volunteers and enthusiasts produces findings that do not generalise to the population that will have to use it.
Continue exploring
Related capabilities
Start with the hard problem
Need the outcome owned, not merely analyzed?
Tell us where delivery is constrained. WeтАЩll map the fastest credible path from intent to verified production.