Site-specific perception and policy models, post-trained on customer data during rollout.
Key Takeaways
- Most of what makes a robot work in a specific factory is a set of small, purpose-trained models, each one closing a gap that appeared during deployment.
- Ati Robotics turns each site quirk, such as a cable across an aisle or a new payload, into a dataset collected by the fleet and a compact model post-trained during the rollout itself.
- These models run on the robot or on a camera's edge box, stay bounded and inspectable, sit beside the certified safety layer rather than inside it, and are cheap to retrain when the site changes.
A cable lies across an aisle. To a lidar it is nearly invisible; to a generic obstacle detector it is either nothing or an emergency stop, and both answers are wrong. The correct behavior is to see the wire for what it is and roll over it smoothly, at speed, without waking up the whole line. On our robots that behavior is a small segmentation model, trained on that class of cable, running on the robot's own compute.
That model is not an exception. It is the pattern. Most of what separates a robot that works from a robot that works in this factory is a collection of models like it: small, purpose-trained, each one closing a gap that appeared during deployment.
The gap
Robots do not fail in factories for lack of intelligence. They fail on specifics: a wire, metal burr near a machining cell, a payload a few centimeters taller than last week's. No amount of general capability anticipates these, because every plant accumulates its own physical dialect. Deployment is the process of learning that dialect, and the question is only whether you learn it with heuristics and exclusion zones, which slow the robot down and pile up as debt, or with models.
We learn it with models. Each quirk becomes a dataset, collected by the fleet already driving the site, and then a compact model post-trained during the rollout itself. The cable segmentation model above is one. Sub-5-centimeter objects detected and avoided in live aisles is another. None of these is glamorous, and that is the point: this is what deploying actually consists of.
Reading the floor
Tugging adds a harder perception problem than driving: the robot is responsible for a train it cannot fully see. One model watches for humans in and around the trolleys: a person stepping between tug and load is detected even when the trolley occludes them; the payload envelope is tracked through blind corners. A second watches the train itself: how the trolleys follow and swing, and whether the follower is where the tow geometry says. Perception here is not a convenience; it is what makes towing at speed defensible in a shared aisle.
Acting on what it sees
Payload handling works the same way: one robot picks payloads of varying sizes and types because identifying and engaging the payload is a vision model, not a mechanical preset. The alternative is a robot variant per payload, which is how fleets stop scaling.
The factory grows eyes
Not all of these models ride on robots. Ati Eye puts them on fixed cameras: traffic management in zones shared by humans, our robots and other vendors' robots, including recognizing third-party bots and yielding to them. The same pipelines watch staging areas and supermarkets (occupancy, slot state, what is ready to pick), and the orchestration layer (Part 8) consumes what the cameras see.
Why small
These models execute on the robot's embedded compute or a camera's edge box, inside the control loop's time budget. A site quirk arrives with dozens to hundreds of examples, not millions, and a compact model post-trains on that in days, inside the deployment window. Its behavior is bounded and inspectable: a cable segmentation model can be validated against its envelope and its failure modes enumerated, which lets it sit adjacent to the certified safety layer without ever being inside it. When the quirk changes (a new trolley, a new payload), retraining is cheap and local.
That loop depends on owning the pipeline: our perception stack is ours end to end. Nothing to license, nobody to wait for. When a deployment surfaces a new quirk, the loop is short: the fleet collects the data, a model is post-trained and validated, the site gets the update. Run that loop enough times and deployment compounds: every quirk becomes a small model, and every model makes the next deployment faster.
None of this is an argument against large models; Part 2 covered our research there. It is an argument about fit: on the floor the work is specific, the data is site-shaped, and the compute is on the robot. Small, purpose-trained models are what that combination selects for.
Next: A Robot That Knows It's Lost, measuring localization confidence with a committee of experts.

