AI

Rhoda AI Exits Stealth With $450M for Robot Intelligence

Rhoda AI left 18 months of stealth on March 10, 2026 with a $450M Series A at a $1.7B valuation and FutureVision, a robot foundation model pre-trained on hundreds of millions of internet videos.

Rhoda AI Exits Stealth With $450M for Robot Intelligence — article cover
On this page6 SECTIONS
  1. Why VCs Put $450M Into a Series A
  2. FutureVision: Predicting Actions From Video
  3. From Internet Video to the Factory Floor
  4. The Data Flywheel Is the Real Bet
  5. What It Means for Robotics Developers
  6. Sources

On March 10, 2026, Palo Alto-based Rhoda AI emerged from eighteen months of stealth with two announcements at once: a $450 million Series A (valuing the company at $1.7 billion, per Reuters and Bloomberg) and FutureVision, a foundation model for robot intelligence. The round drew Capricorn Investment Group, Khosla Ventures, Leitmotif, Matter Venture Partners, Mayfield, Premji Invest, Prelude Ventures, Temasek, and Xora, plus individual investor John Doerr.

The timing itself says something. The money landed one day after Nscale set the record for Europe’s largest-ever round (see Nscale’s $2B Series C) — but the direction is the opposite. One bet is compute infrastructure; the other is the intelligence layer that goes inside robots. The capital markets apparently think both roads deserve heavy commitment.

Why VCs Put $450M Into a Series A

Part of the pitch is the roster. Co-founder and CEO Jagdeep Singh is a repeat deep-tech founder; Chief Science Officer Eric Ryan Chan is a Stanford researcher in computer vision and generative modeling who was previously a generative model architect at WorldLabs; co-founder Gordon Wetzstein heads Stanford’s Computational Imaging Lab.

What stretches a Series A to $450 million, though, is the scale of the task: the team is building a cross-embodiment robot foundation model, not control software for one machine. Premji Invest’s Patnam stated the investment logic plainly: such a model can “kick-start a powerful data flywheel, creating a compounding advantage in capturing the long tail of real-world edge cases.”

FutureVision: Predicting Actions From Video

FutureVision is built on a proprietary architecture Rhoda calls Direct Video Action (DVA). Training runs in two stages: pre-training on internet-scale video — hundreds of millions of public clips — to learn motion and physics priors, then post-training on smaller robot datasets to map video predictions into actions and tune embodiment-specific behaviors.

At runtime it operates closed-loop: observe, predict future states as video, act, re-observe, on cycles of a few hundred milliseconds. The design inherits the world-model playbook directly — the model first understands how the scene will evolve, then decides how to move. Bringing a new task online requires roughly ten hours of teleoperation data, which against per-task custom engineering in traditional robotics is an order-of-magnitude improvement.

From Internet Video to the Factory Floor

Rhoda’s go-to-market choice is pragmatic: manufacturing and logistics. The press release cites a high-volume manufacturing evaluation in which a component-processing task ran under two minutes per cycle without human help and beat customer KPIs. Leitmotif’s Wiese framed the customer perspective: “The real challenge isn’t solving it once, it’s delivering consistent, reliable output under real-world production conditions.”

On the business side, Rhoda currently deploys with industrial partners but intends to license FutureVision to other hardware and software platforms over time — selling the robot brain, not the robot. The structural resemblance to the LLM industry’s API-ification path is hard to miss.

The Data Flywheel Is the Real Bet

Line up three facts and the actual object of the investment comes into focus: the pre-training corpus is public internet video (near-zero acquisition cost), post-training needs only ten-hours-scale teleoperation data, and every factory deployment generates proprietary robot data nobody else can get. Each turn of the flywheel lowers the entry cost of the next task while widening competitors’ data deficit.

This is the shared script of the 2026 robotics race: video pre-training supplies the “common sense,” and data returned from deployments supplies the “expertise.” Whoever scales deployment first turns the physical world’s long tail into a moat.

What It Means for Robotics Developers

Three practical takeaways. First, the robot foundation model is becoming a procurable, licensable component — if you build for a specific embodiment or vertical, the competitive question is shifting from “build or not build” to “whose intelligence layer.” Second, the video-pre-training-plus-light-teleoperation recipe has now been market-validated more than once; new teams should differentiate on data and domain, not on re-creating pre-training from scratch. Third, closed-loop control is latency-sensitive (few-hundred-millisecond cycles), so edge deployment and inference cost are hard engineering constraints — factor a model vendor’s edge strategy into selection before committing.

Sources

AI-assisted summary compiled from the sources above, reviewed by a human before publishing.

FOUND_THIS_USEFUL?

Support more practical AI articles, tutorials, and build notes.

BUY_ME_A_COFFEE
SHAREXEMAIL