Four generations of models, a research partnership, and a clearer answer than we expected to the question of where a learned model actually belongs in an industrial robot.
Key Takeaways
- Vision foundation models proved most useful as open-world perception, feeding estimators and controllers that stay classical.
- Across four generations of navigation research, the variable that mattered was the representation handed to the controller, not the network that produced it.
- Ati Robotics is applying that lesson with OccupancyReact, a learned local policy conditioned on the occupancy and cost maps its robots already produce.
Recently in our test facility, amidst much cheer, we watched a robot drive itself to a specific object across the room. Our robots do harder things in customer factories every day. What made this run worth watching was how it navigated: no LiDAR occupancy grid, no scan-matched pose, no metric map at all. Just a camera, a sequence of images from an earlier walk through the space, and a 3D foundation model turning pairs of those images into geometry.
Not long ago this was a research demo in a university corridor. Now it runs on our own hardware, on a stack we can modify and fine-tune. Getting there took a research partnership and multiple published papers, and taught us something more useful than whether the demo works.
THE BET
Why we went looking for a vision model
Our production stack is geometric and classical at its core, and it has worked well: 3D maps built from LiDAR, inertial sensing and wheel encoders for motion prediction, localization by fast scan matching, a model predictive controller keeping cross-track and heading errors in check, and task-specific fine-tuned models for visual perception. It has run in customer factories for years, and the important parts are deterministic by design.
At Ati Robotics, safety is paramount. Every behavior that moves a heavy industrial machine around people must remain inspectable, bounded and testable. That is a great feature when you are certifying a machine that moves a tonne of steel past a person, and a limiting constraint once the list of things you want the fleet to do outgrows what any team can code. So we went looking for approaches that bring together the best of rule-based development and general human intuition, with the ultimate hope of solving the planning problems that demand tight, precise maneuvers. To that effect, we began a sponsored research engagement on mapless visual navigation with the Robotics Research Centre at IIIT-Hyderabad, under Prof. K. Madhava Krishna. The collaboration has produced work at IROS, NeurIPS and ICRA, and several quite different answers to the same question: how much of navigation should be learned?
GENERATION ONE
Foundation features are the durable win
SparseLoc attacks a problem every warehouse robotics company knows: dense LiDAR maps are precise, but enormous. SparseLoc builds a sparse semantic-topometric map instead, using vision-language foundation models to identify open-vocabulary landmarks without task-specific training, and localizes against this map with Monte Carlo localization. On KITTI it maintains average global localization error below 5 m and 2° while retaining roughly one five-hundredth of the points used by dense mapping methods.
The shape of the answer
The foundation model here supplies open-world perception: it turns visual observations into semantically meaningful landmarks (fire extinguisher, doorway, pillar) without requiring a warehouse-specific recognition model. The estimator remains classical: a particle filter. A pre-trained model provides general perception; classical estimation provides geometric and probabilistic rigor. That division of labor is an important pattern in practical foundation-model robotics.
The same group's SegMASt3R adds an interesting qualification: off-the-shelf 2D foundation models can initially outperform a geometry-oriented model such as MASt3R at segment matching, but after task-specific fine-tuning the ordering reverses and the geometry-grounded model wins. General-purpose visual representations are valuable; explicit geometric inductive bias still matters when correspondence must survive extreme viewpoint changes.
GENERATION TWO
General navigation models, and the price of generality
The next wave came from General Navigation Models (GNM, ViNT and NoMaD). The premise: one policy, trained across many robot bodies, navigating from camera images to an image goal with no metric map at all. We evaluated ViNT in Isaac Sim, then on a robot in one of our halls. Out of the box it reached the goal in most runs. That is a model trained on other people's robots, in other people's buildings, driving ours.
The failures clustered in specific scenarios: tight maneuvers, in-place turns, a persistent gentle sway. They share a root cause: embodiment-agnosticism is the design. A policy trained to be indifferent to the body it drives has little reason to care about our wheelbase, steering geometry, payload, turning radius or trailer jackknife envelope. With no metric anchor and no notion of safe speed for this machine, the practical answer is to hold speed low and let the model steer.
For a research agenda aimed at generality, a reasonable trade. For us, it is backwards. We know our kinematics exactly (measured, modeled and validated), and discarding that knowledge to buy portability to robots we do not build is not a trade we would choose. Keep the open-world perception; give the geometry back.
GENERATION THREE
Geometry without a global map
The most interesting idea of the three is the one now running on our robots: MASt3R-Nav. Metric maps demand global consistency: every point and pose registered into one frame and kept that way. Topological maps avoid that burden by discarding almost all the geometry. MASt3R-Nav threads between them.
MASt3R is a 3D foundation model: give it images of the same scene and it returns dense pixel correspondences and per-pixel 3D coordinates, known as pointmaps. MASt3R-Nav builds a graph whose nodes are individual pixels: matched pixels across frames joined by zero-cost edges, since they are the same point observed more than once; pixels within a frame joined by edges weighted by the 3D distance MASt3R gives between them. A shortest-path search from a goal pixel assigns every pixel in view a cost-to-go. The result is a WayPixel costmap, and that image of costs conditions a learned controller the authors call PixelReact.
The controller reads the costmap and predicts a handful of local waypoints that carry the robot to the goal. The navigation is geometrically precise without ever being globally consistent. There is no metric model of the building, only relative geometry between pairs of images.
It leads every prior approach on the imitation task and on the four-task average:
| Method | Representation | Imitate SPL | Avg SPL (4 tasks) |
|---|---|---|---|
| GNM | Image-relative | 78.79 | 26.55 |
| PixNav | Object-relative | 42.42 | 23.09 |
| RoboHop | Object-relative | 57.56 | 32.19 |
| ObjectReact | Object-relative | 60.60 | 33.36 |
| MASt3R-Nav | Pixel-relative | 93.94 | 52.79 |
HM3D-IIN validation, 36 scenes, Success weighted by Path Length. Source: MASt3R-Nav, ICRA 2026.
WHAT WE LEARNED
The model was never the interesting variable
Line up these three generations and a pattern appears that has little to do with which network anyone chose. What changed each time, and what determined how well the robot drove, was the representation handed to the controller.
For a company that builds one family of robots, knows their dimensions to the millimeter and already carries a 3D LiDAR, the far end of that ladder is home ground.
GENERATION FOUR
Bring back the geometry we already know
MASt3R-Nav led us to a question we had not expected to be asking: if the useful abstraction is a geometrically meaningful cost representation conditioning a learned controller, what should that representation be on a robot that already carries a 3D LiDAR?
We do not need a neural network to rediscover which parts of the floor are occupied. Our production robots already estimate that reliably. We also know the robot footprint, kinematics, steering limits and safety margins. Throwing those facts away simply to make the policy more general would solve a problem we do not have. So our current experiments move one step further along the ladder. We call the approach OccupancyReact.

OccupancyReact. The policy combines a robot-centric view of free space, navigation cost and goal location with the robot's current motion state, then predicts a short five-waypoint trajectory rather than an instantaneous velocity command.
Instead of conditioning the policy on raw images, detected objects or WayPixel costs, we hand it a local robot-centric occupancy map, a goal-conditioned cost-to-go field and a goal mask. The occupancy map answers where the robot can physically move. The cost field answers where it should move. The learned policy is left with the part that is hard to hand-engineer: how this particular machine should maneuver through that geometry.
OccupancyReact is not itself a foundation-model system. That is precisely why it belongs in this story. Working with foundation-model navigation changed our view of what should be learned and what should stay explicit. The transferable idea was the structured representation and the learned local policy. The model used to build the representation mattered less.
Like MASt3R-Nav, the controller is trained in simulation. The test that matters is whether its closed-loop behavior survives the move to the floor unchanged. Hardware trials are underway.
HOW WE ARE ENGINEERING IT
Start simple, expose the failure modes
We started in deliberately controlled conditions. Random occupancy grids supplied the geometry, Dijkstra supplied a globally valid cost-to-go field, and an existing classical local controller acted as the expert. At every step the policy saw only a robot-centered crop of the occupancy and cost fields, and learned the local behavior needed to make progress toward the goal.
One early design choice mattered more than it first appeared: predicting a short sequence of local waypoints behaved better than predicting instantaneous velocity commands. A short trajectory gives the network a small amount of temporal structure. It expresses intended motion rather than a single reaction, and in our simulation experiments produced visibly cleaner local behavior.
Global planning, localization and collision geometry already have reliable engineering solutions. What we want to learn is the local behavior that sits between a geometrically valid plan and a precise physical maneuver.
WHERE WE ARE HEADING
Learn the last two meters
The clearest articulation of the question came from Prof. Madhava Krishna:
"Indeed we expect things to be easier with a LIDAR based cost-map. However my own feeling though is if we have near perfect LIDAR based cost-maps we can go with A* or a classical planner as well. But perhaps there is still some meat there for policy learning."
Prof. K. Madhava Krishna, Robotics Research Centre, IIIT-Hyderabad
He is right. Where we have a good costmap, A* and a well-tuned controller are excellent. The harder part is often the end of the journey: docking, pallet approach, threading into a narrow bay, converging on a final pose within a centimeter, or controlling a trailer while it settles behind the tractor. These are the maneuvers where a skilled operator is visibly better than our controllers, and where we have written many rules approximating what that operator simply does.
Whether the best teacher is an existing planner, reinforcement learning, carefully collected human demonstrations, or some combination of them remains an engineering question rather than an article of faith. That is the space in which we expect learning to earn its place: capturing the local maneuvering strategy that becomes increasingly awkward to express as a growing set of handcrafted rules, while free space and safety envelopes stay explicit.
"Grid cost-conditioned policy sounds great! It immediately gets rid of the imperfect-costmaps issue due to the RGB-only perception-planning backend. Although BEV-based learnt control policies exist, they might not be using dense grid costmaps in particular."
Prof. Sourav Garg, Robotics Research Centre, IIIT-Hyderabad
THE GENERAL CLAIM
Learn what benefits from learning
The broader architectural lesson is to expand the learned component only where it consistently earns the right to replace an engineered primitive. Localization, planning, control and safety can remain explicit, inspectable modules. Foundation models can supply context, semantics and increasingly useful representations around them.
Over time, some of those boundaries will move. That is fine. A learned component should enter the production stack when it is more useful and more reliable than the primitive it replaces, whatever its size or generality.
We now have vision-driven navigation running on our own hardware, a research partnership producing work at top robotics venues, and, most valuable, a clearer view of the boundary.
The answer is narrower than "foundation models everywhere," but considerably more useful: learn the pieces that benefit from learning, preserve the geometry we already know, and make the boundary between the two an engineering decision. That boundary is what we mean by Physical AI.
The next post takes one concrete instance: teaching a robot a three-point turn, from a one-line brief to a primitive the planner can call.
What This Means on Your Floor
What this means on your floor
The navigation that moves a pallet across your plant is already solved and already deterministic. The work here is aimed at the last two meters, where a robot docks, approaches a pallet or threads a narrow bay, and where every second saved multiplies across a shift. Getting that right shortens cycle times without loosening a single safety constraint.
References
- P. Paul et al., "SparseLoc: Sparse Open-Set Landmark-based Global Localization for Autonomous Navigation," IEEE/RSJ IROS, 2025, pp. 11057-11064, doi: 10.1109/IROS60139.2025.11245963.
- S. Garg et al., "ObjectReact: Learning Object-Relative Control for Visual Navigation," CoRL, 2025, arXiv:2509.09594.
- V. Garg et al., "MASt3R-Nav: WayPixel Navigation in Relative 3D Maps," IEEE ICRA, 2026, arXiv:2605.24111.

