Last week I attended Actuate, a developer conference for the robotics and autonomy industry hosted by Foxglove, a data platform for physical AI.
I entered the conference with five questions:
Has general purpose autonomy reached it’s GPT-3 moment?
Are VLAs or world models more important to autonomy?
What data sources are most important to solving and scaling physical AI?
Are full-stack (hardware + software) robots more likely to succeed than hardware agnostic physical AI models running across different form factors?
How has the software stack underpinning this new age of robotics taken shape?
The industry hasn’t reached consensus on points 2-4, but I saw compelling arguments from every side of each debate, and I left with a clearer map and a few real answers.
Here were my key takeaways:
Specialized robots have had their GPT-3 moment; general-purpose autonomy has not. Narrow tasks now run reliably at scale. Generalist models are making large improvements but are stuck at 5–20 second horizons or need task-specific post-training to perform tasks reliably.
The industry has not converged on models, data or full-stack vs software-only:
World models vs. VLAs is a spectrum. Dyna, Veeda, Wayve, and 1X treat world modeling as necessary for cross-embodiment generalization; Physical Intelligence uses world models lightly to generate visual subgoals; and Sunday Robotics reached 99.1% zero-shot with none in their stack.
Data: Every data source is useful but companies prioritize them differently. Teleoperation buys dexterity but cannot scale, egocentric video scales cheaply but misses contact forces, and simulation offers the largest potential volume.
Full-stack vs hardware agnostic software: Full-stack is winning and all players operating at notable scale have hardware and software tightly integrated. Still, generalizable hardware-agnostic software players are chasing the largest markets.
The autonomy software stack is consolidating around a few defaults, but the core stays in-house. Scaled players build their own simulation. NVIDIA owns compute, Applied Intuition is consolidating simulation, Foxglove is becoming the observability standard, and labeling incumbents face pressure from synthetic data and new collection networks.

Has robotics hit its GPT-3 moment?
Specialized robots have reached the GPT-3 moment
Scale among vertical robots is proven. Waymo is at 500,000 paid rides per week, Zoox just opened up public access in San Francisco and Las Vegas. Zipline completes a delivery every 30 seconds. Bonsai Robotics has deployed hundreds of autonomous agriculture robots.
General-purpose robots can hillclimb individual tasks with enough post-training. Many teams can hill-climb one task to production reliability with enough post-training; Sunday Robotics’ ACT-2 hit 99.1% zero-shot across 785 attempts in 31 unseen homes and 9 garment types, on one fixed checkpoint with no per-home adaptation. Dyna-1 ran continuous laundry fulfillment at a commercial laundromat in Sacramento at over 99%.
General-purpose autonomy has not
Time horizons are seconds, not minutes. Generalist AI’s one-shot in-context learning works on 5–20 second primitives, averaging roughly a 6-second horizon on benchmarks; longer behavior must be chained from multiple prompts.
Absolute success rates are low. Dyna-2 moved on-robot success from 20% to 53% across 14 tasks and three hardware platforms, with just 1–2 hours of task-specific post-training per task.
The field agrees on the size of the gap. 1X says we need “thousands and thousands times more data”; newly launched Veeda says “pure imitation learning doesn’t bring you all the way.”
World Models vs. VLAs
A VLA turns pixels and language directly into motor commands while a world model predicts how the environment will evolve, projecting future states. Founders and researchers are in disagreement over which of the two models are more important:
Companies on both sides of the argument have seen success. Sunday “solved” laundry folding without world models, Generalist can one-shot ~5 second tasks without world models, and PI is making progress with a lightweight world model. Waymo is at scale while relying on world models, Wayve is leveraging world models and approaching a public launch, and Dyna’s robots used world models to “solve” laundry folding.
What everyone agrees on. Almost nobody runs a world model as a runtime controller: Dyna keeps its action stream independent so the robot stays reactive, and Wayve uses GAIA for synthetic data and offline evaluation while the driving policy stays end-to-end. Veeda’s taxonomy makes the distinction - passive synthesis, trajectory prediction, and interactive simulation, where only the third closes the learning loop.
Importance of data sources
No source is sufficient alone and none is strictly dominant. The Data Wars panel rejected the idea of a war outright, but companies have prioritized different data sources.
Companies agree that multiple data sources are important but they don’t agree on which is most important
Simulation is the fastest-moving and most scalable layer. Simulation could be the most important due to its unbounded scalability. Veeda’s Sim 1.0 → 3.0 progression ends in generative world models. Wayve is running on world simulators, which have also been a key pillar of Waymo’s success.
Full-stack vs. Hardware-Agnostic Software
Full-stack is winning. Scaled players have hardware and software tightly integrated. Still, generalizable hardware-agnostic software players are chasing the largest TAM.
Wayve’s public launch will be the first win for hardware-agnostic software. Wayve, which builds AI driver models built to work across a variety of vehicles, is close to a public launch in London.
1X is building a development ecosystem. While the company is building the robot and world model, they will provide data collection rigs (cameras and liDAR headset, tactile pressure-sensing gloves, wrist, ankle and waist sensors) where people can train individual skills and tasks. 1X will provide APIs, cloud storage, and simulation tools. More details to come out in September.
Physical Intelligence may go full-stack. While Physical Intelligence has positioned themselves as hardware agnostic software, they have begun hiring hardware engineers and have job postings that indicate they will be designing their own hardware stack (Manufacturing engineer job posting) (LinkedIn post).
The Autonomy Software Stack
Companies write their own ROS: Open-source ROS is the starting point but most players build their own. Software vendors exhibiting at the conference integrate with highly customized, proprietary robot operating systems and bespoke internal tools still outweigh third-party dev tools. Companies typically prototype on ROS 2 and go in-house after hitting performance ceilings at scale.
Internally built tools are still the norm: Between a third and a half of engineering time is spent building custom, internal tools.
NVIDIA is still dominant in compute and their moat is software and CUDA. Even Unitree humanoids ship with Jetson onboard. A few teams complained about pricing and thin production deployment tooling but stay for CUDA and the Isaac ecosystem.
Simulation: Scaled players built in-house while others are buying from vendors. At-scale AV players and OEMs keep simulation in-house for IP and safety reasons. Below that line, Applied Intuition is gaining adoption in AV and defense sectors. Manipulation-grade simulation and model evaluation remain open frontiers, where newer entrants like Lightwheel (SimReady assets, RoboFinals eval platform) are growing.
Foxglove is becoming the industry default and expanding up the stack. Already the standard for robotics visualization and debugging, its new Remote Access Gateway moves it into live operations: centralized robot operating centers, incident triage, and remote intervention demoed at roughly 125ms cross-country latency.
Labeling is growing but squeezed from two sides. Scale, Encord, and Labelbox remain the enterprise annotation standards. Pressure is building from training mixes shifting toward majority synthetic data, and from an industrializing collection layer upstream of annotation (Instawork’s 30,000-workers collecting egocentric data, Lightwheel).
Thank You
Thank you to Foxglove and all of the speakers & exhibitors for such a great conference! Can’t wait for Actuate 2027.
If you’re an operator or investor in the space, I’d love to hear from you. Please reach out at kgetsiv@indeed.com
Thanks for reading!
Footnotes & Disclosures:
The information presented in this publication is for informational purposes only and should not be construed as investment advice or a recommendation to buy or sell any security. The views expressed are solely those of the author and do not represent the views of any company, employer, or affiliated party.
All data and analysis are derived from public sources, including company websites. No confidential or proprietary information has been used or referenced. While the information is believed to be accurate and from reliable sources, no warranty is made as to its completeness or accuracy, and no liability is accepted for errors or omissions. Past performance is not indicative of future results.




