ENTRY ANGLES
Vertical-specific synthetic research for mid-market · API integration into existing product tools · Behavioral foundation model pre-training
VERTICALS
CAPABILITIES
Behavioral data synthesis, Statistical modeling, Enterprise sales, Data partnership development
CVS Health is making product stocking decisions based on consumers who don't exist.
That sentence sounds like a liability. It's actually competitive infrastructure.
Simile has built what it calls synthetic research: AI-generated populations that behave like your real customers, respond to your hypothetical products, and deliver statistically reliable signal before you've spent a dollar on manufacturing or a day on focus groups.
The way it works: Simile trains a customer model using your existing behavioral data – purchase history, return patterns, support interactions, demographic signals. That training process takes roughly seven months and builds a population of synthetic agents that mirror the preferences, sensitivities, and contradictions of your real customer base. Then you can ask that population anything. Which SKU wins in aisle five? What price point triggers abandonment? Which new feature would pull power users from a competitor?
CVS Health ran the experiment. The synthetic population's predictions matched what real consumer panels actually said 85% of the time – a benchmark that makes most traditional research methods look expensive by comparison. The average focus group costs between $50,000 and $200,000 to run. It takes weeks to recruit, schedule, and analyze. Simile collapses that into days, at a fraction of the cost, with a customer model that only gets more accurate over time.
The $76 billion market research industry has been telling companies they need more data, better surveys, longer feedback cycles. Simile's thesis is the opposite: you already have the data, and the bottleneck is synthesis. What CVS actually needed wasn't more consumer panels – it needed a faster way to test what the panels would say.
This matters beyond retail. Any organization sitting on behavioral data and still running expensive validation cycles is a potential customer. Healthcare, financial services, subscription software, consumer packaged goods – the pattern is the same. You have years of how your customers act. You're still spending months figuring out how they'll react.
Simile has raised $300 million. That's not a research tools company. That's infrastructure for how large organizations will test product decisions going forward.
The model here isn't "replace market research." It's "move the feedback loop from post-launch to pre-design."
Traditional research operates on a question-then-wait cadence. You ship, you survey, you analyze, you iterate. That cycle is measured in quarters. Simile compresses it to days by asking synthetic populations trained on your own behavioral data instead.
The 85% accuracy rate is the credible signal that breaks the trust barrier. You're not betting on AI to invent what customers think – you're betting that the patterns in your historical data are predictive enough to simulate what customers will say next. For organizations with years of transaction data, that's a reasonable bet.
For builders, the structural shift is here: validated research is becoming a software problem. The bottleneck isn't access to consumers. It's the tooling to synthesize what you already know about them. Whoever controls that synthesis layer controls a new kind of competitive advantage – pre-validated product decisions that arrive faster than competitors can even schedule their focus groups.
The $300M raised means enterprise sales cycles are already closing. The question isn't whether this works. It's which industries use it to get three product cycles ahead while everyone else is still recruiting panelists.
Simile's model proves something: behavioral data is already predictive. The unlock is tooling that treats it as such.
The first angle is vertical-specific synthetic research products. Simile serves large enterprises with years of transaction data. The gap is mid-market organizations – regional retailers, direct-to-consumer brands, B2B SaaS companies – that have enough behavioral signal to train a useful model but not the budget or data science teams to build one. A productized synthetic research platform targeting these segments, with lighter-touch onboarding and pre-built industry models, is the natural wedge.
The second angle is research-adjacent software that integrates synthetic populations into existing product workflows. Product managers don't want a new research tool – they want research to live inside the tools they already use. An API layer that plugs synthetic consumer response into Figma prototypes, Notion specs, or A/B testing platforms creates stickiness without requiring workflow change.
The third angle is the underlying model infrastructure. Simile's training process takes seven months on proprietary behavioral data. Any organization that can compress that training cycle – or pre-train foundation models on industry-specific behavioral datasets – has something to sell. Think "behavioral foundation model, fine-tuned to your customer base in days, not months."
The practical entry point: partner with a DTC brand sitting on 500K+ customer transaction records. Offer to train a lightweight synthetic population model, then run three product decisions through it against their usual research process. If the synthetic model matches the research outcomes within 80%, you have the case study. If it gets there faster, you have the pitch.