Andrew Dai spent fourteen years at Google DeepMind before leaving to raise $55M for a visual AI company that hasn't shipped a product. The investors include Nvidia and Jeff Dean. They're betting on a gap.
ENTRY ANGLES
Build a visual understanding API for AI agent developers before Elorian's full product ships · Create GUI automation enhancement tools using Elorian's visual reasoning capability
VERTICALS
CAPABILITIES
AI/ML engineering, Computer vision, Agent development, Developer tooling
Fifty-five million dollars. No product. No customers. Three hundred million dollar valuation.
That's Elorian, founded in 2025 by Andrew Dai — a fourteen-year veteran of Google DeepMind who contributed to the early large language model research that helped make ChatGPT possible — and Yinfei Yang, who built multimodal AI systems at Apple before leaving to co-found something more ambitious.
Their investors include Striker Ventures, Menlo Ventures, Altimeter Capital, Nvidia, and Jeff Dean — who ran Google Brain for a decade. These are not people who fund hobby projects.
The thesis is simple to state and hard to execute: current AI is visually illiterate. Ask GPT-4 to count fingers in a hand-drawn sketch, trace a wire through a circuit diagram, or read a floor plan, and it underperforms a careful ten-year-old. The reason is architectural: most frontier models translate images into text before reasoning about them. Visual information gets compressed. Spatial relationships get approximate representations. The conversion step is where understanding breaks down.
Elorian is building models that reason natively in visual space — understanding spatial relationships and physical constraints without the translation step. Instead of describing what they see, the models think in images directly.
The practical target is AI agents: systems that need to look at a computer screen the way a human does, understand what's on it, and interact with it. GUI navigation, engineering schematics, legal document review, robotic systems that need to understand what they're looking at. Every agentic workflow that interacts with real software has a visual problem. Elorian's bet is that the next generation of agents requires native visual intelligence — not vision as a feature bolted onto a language model, but as a first-class reasoning mode.
Pre-product. Pre-revenue. Three hundred million dollars says the window is now and the team is the right one.
The visual reasoning problem isn't a minor gap in current AI capabilities. It's a structural ceiling.
Every enterprise workflow built for AI agents runs into it eventually. The sales tool that needs to read a competitor's product screenshot. The legal system that needs to compare two contract redlines. The infrastructure agent that needs to interpret a network diagram. Each requires visual understanding that current models patch around with workarounds — OCR, image captioning, multi-step translation pipelines — rather than solve.
The market timing argument for Elorian is that native visual reasoning is where language reasoning was in 2020: the capability exists in early form, the architecture to scale it is still being built, and the company that gets there first will own the infrastructure layer the way early language model companies owned text embedding.
Andrew Dai and Yinfei Yang are not generalists making a product bet. They built the underlying research at Google DeepMind and Apple that made today's frontier models possible. This is a founder-timing match: the people who built the prior generation of AI stepping out to build the specific capability the next generation needs.
The window is defined by when the incumbents catch up. OpenAI has GPT-4V. Google has Gemini's vision stack. Both are retrofitting better visual understanding onto text-first architectures. A company building visual reasoning natively from the start has a structural advantage that compounds — and the window that matters is the two to three years before frontier labs ship native visual reasoning at scale.
Elorian's core technology has an immediate B2B application that doesn't require waiting for the full model: a visual understanding API for AI agent developers. Every team building computer-use agents needs better visual parsing than today's frontier models provide. Sell the capability as infrastructure before the flagship product ships.
Elorian's technology is enabling infrastructure, not an end product. Companies building specialized agents for healthcare imaging, architectural review, legal document comparison, or manufacturing quality control need exactly this capability. Build relationships and investment positions in vertical agent companies that will bet their stack on Elorian's foundation — they're identifiable now, before Elorian has a public product.
The current generation of GUI automation tools — RPA vendors, Zapier-style platforms, computer-use wrappers — are all limited by the quality of their visual understanding. A drop-in replacement or enhancement layer that dramatically improves visual parsing for existing automation workflows is a near-term revenue opportunity that doesn't require waiting for Elorian's model to ship.