Product teams are making more decisions across more customer segments in less time. Yet the evidence behind those decisions still depends on research methods that are difficult to scale. Synthetic users introduce a new possibility: teams can explore journeys, reactions and edge cases before committing scarce research and development resources. The risk is confusing plausible AI output with reliable customer evidence. The value of synthetic user testing will depend on how closely simulation is connected to real data, real behaviour and real validation.
Product Research Is Reaching Its Scalability Limit
Modern product development creates a structural mismatch. Release cycles are shorter, product surfaces are larger and user segments are more varied. A single feature can generate questions about onboarding, pricing, adoption and retention. Each may deserve research, but interviews, moderated usability sessions and beta programmes cannot expand at the same rate.
In Maze’s Future of User Research Report 2025, 55% of 800 product professionals said demand for user research had increased, while 63% named time and bandwidth as their leading challenge. The same report found that organisations embedding research into strategy and operations reported 2.7 times better outcomes than teams that rarely incorporated user insight. Research is becoming more valuable as capacity becomes harder to stretch.
Modern products now generate more hypotheses than teams can realistically validate through interviews, usability sessions and limited beta programmes. Product organisations therefore research the largest decisions and rely on intuition for many smaller ones. Traditional research has not lost value; human research time has become too valuable to distribute equally across every variation.
Synthetic Users Are Changing How Teams Test Product Decisions
Synthetic users move product research earlier in the development cycle. Instead of waiting until a prototype is complete and participants are recruited, teams can simulate how different customer profiles might move through onboarding, respond to pricing, adopt a feature or encounter friction in a navigation path. This does not make a final decision. It makes weak assumptions visible sooner.
Consider a team redesigning onboarding for a B2B platform. Testing every combination with administrators, managers and daily users could take weeks. Synthetic testing can run the scenarios first, reveal recurring friction and narrow the field to the strongest options. Human research can then focus on understanding why real users behave as they do.
A study presented at IEEE BigData 2024 generated three sets of 1,000 synthetic product reviews for product-desirability testing. The datasets showed sentiment correlations of 0.93 to 0.97 with the source patterns, although researchers found a modest bias toward positive sentiment. The result is encouraging for AI product testing, but the bias matters: simulation may reproduce broad patterns while smoothing over the dissatisfaction and contradiction that often contain the most valuable insight. The study therefore positions synthetic data as useful where real test data are limited, rather than as a substitute for them.
Synthetic users expand the number of decisions teams can examine before those decisions become expensive. Their purpose is not to certify that customers will accept an idea. It is to improve which ideas reach real validation.
Simulation Is Valuable Only When It Reflects Reality
A synthetic user is only as useful as the evidence used to construct it. A fluent response can sound like customer understanding while being little more than an average of familiar online opinions.
Trustworthy simulation should be grounded in signals a company already holds: CRM records, product analytics, support tickets, behavioural events, prior interviews and historical usage. Analytics show what users did. Support conversations reveal where they struggled. CRM data establishes commercial and organisational context. Interviews explain motivations that event logs cannot capture.
A major Stanford study illustrates the difference grounding makes. Researchers created agents representing 1,052 people using two-hour interviews, then tested them across surveys, personality measures and behavioural games. The interview-grounded agents matched participants’ survey answers at 85% of the participants’ own two-week consistency level. They also achieved an 80% correlation on personality tests and 66% on economic games. Agents built from richer interviews outperformed versions given only demographic data.
For product leaders, demographics and generic personas are not enough. A synthetic “operations manager” built from a job title and company size may talk convincingly about efficiency. A version grounded in support history, system permissions, integration constraints and actual task sequences is more likely to expose why a simple workflow change would fail inside the customer’s organisation.
Synthetic users do not replace customer understanding. They amplify whatever understanding already exists.
Poor data is therefore not neutral. Missing segments, outdated behavioural records and overrepresented positive feedback become embedded in the simulation. Nielsen Norman Group’s evaluation found that synthetic-user responses were often favourable, vague and insufficient for prioritisation. Its recommendation is to treat the output as hypotheses for further research, not as evidence for final decisions.
The governing question should shift from “How realistic does the AI sound?” to “Which customer evidence produced this result, and how representative is it?”

The Future Is a Validation Loop, Not Synthetic vs. Real Users
Synthetic and real users operate at different levels of certainty.
Synthetic users are suited to breadth. They can explore scenarios, challenge assumptions and identify patterns worth investigating. Real-user research provides depth: lived context, unexpected behaviour and the ability to explain why a decision works in one environment but fails in another. Product analytics shows whether those findings survive contact with actual usage.
The more mature workflow is:
Idea → Synthetic testing → Analytics comparison → Real-user validation → Release → Behavioural feedback

Synthetic testing narrows the option space. Existing analytics checks whether predicted behaviour resembles known patterns. Interviews and usability sessions examine the highest-impact assumptions. Post-release data then feeds back into future simulations.
This changes where human research creates the most value. Researchers spend less time screening minor variations and more time investigating consequential questions: why customers abandon a workflow, how buying groups make trade-offs or what users do when the product no longer matches operating reality.
A simulation that cannot be compared with later real-world outcomes remains a persuasive story. One whose predictions are tracked and recalibrated becomes part of an organisational learning system.
Continuous Product Simulation Requires an Operational Foundation
Building a dependable synthetic-user capability across the product lifecycle is an operational challenge.
The evidence required for simulation is usually scattered across analytics, CRM platforms, support tools, research repositories and operational systems. Definitions differ, identities do not always match and historical data may reflect conditions that no longer exist. Without integration, teams end up prompting a model with static personas and calling the result customer intelligence.
A continuous capability needs a governed data flow that can update customer representations, preserve segment differences, record which sources informed a result and compare predicted behaviour with actual outcomes. It also needs clear boundaries around consent, privacy and customer-data use. The Stanford team maintained audit logs and withdrawal controls for the agents representing study participants – an important reminder that more realistic simulation creates stronger governance obligations.
This is where Twendee’s role as an AI deployment partner becomes relevant. The work is not simply generating synthetic personas. It is connecting product and customer data, integrating simulation into development workflows and building the validation layer that compares AI-generated insight with analytics and real-user evidence. That foundation allows synthetic testing to support feature prioritisation, pre-release evaluation and post-release learning without becoming an isolated experiment.
Competitive advantage will not come from producing the largest population of synthetic users. It will come from building the strongest operational link between simulation, product decisions and real customer behaviour.
Conclusion
Synthetic users are unlikely to replace traditional product research. They will change where human research delivers the greatest return. As product teams use AI to explore more scenarios before shipping, stronger organisations will reserve direct customer access for decisions that require genuine context, judgement and empathy and use simulation to make those decisions sharper.
Twendee helps enterprises build AI-enabled product testing workflows grounded in real data, connected systems and continuous validation. Visit the Twendee website, follow Twendee on LinkedIn, or book a conversation through Twendee’s Calendly to explore how synthetic testing can become a reliable part of your product-development lifecycle.



