The kidnapped-robot problem is not finding the pose. It is knowing when not to trust one.
Key Takeaways
- Confidence cannot come from the same sensor and map that produced the error. It has to come from agreement between independent systems that fail under different conditions.
- Ati Robotics lets 3D scan matching propose a pose while camera place recognition, a WiFi radio map and GPS check it, and any layer can answer "unknown".
- When any expert disagrees, the robot reports itself lost and waits instead of guessing, a quiet prerequisite for lights-out operation.
A vehicle in fleet mode was powered off, moved, and powered on elsewhere in the factory. It resumed its trip: the localizer had returned a pose a meter from the last known position (wrong, but on the planned path). Nothing flagged the localization as bad, so the robot drove, stopping only on cross-track error. Near a ramp or a dock door, the same failure is a safety incident.
The pose estimate was not the failure; the failure was a localizer that always returns a pose and never reports that it does not know.
Measuring confidence
A confidence score attached by the localizer fails for a structural reason: it is computed from the same sensor and map that produced the error. Two aisles built to the same drawing return nearly identical LiDAR scans; match against the wrong one and the score is high, because the geometry matches. An internal stability check (re-seed from the returned pose, test whether the maximum stays put) helps, and is still the same sensor against the same map.
There is no single confidence score that holds across sites and conditions. So the design corroborates the pose instead, with systems whose sensors fail under different conditions.
The experts
The pose proposal comes from 3D branch-and-bound scan matching, built on the open-source 3D-BBS. The search is exhaustive: every candidate translation and heading is considered, and the result is the global optimum of the match score, not a local search's best result. What makes this affordable is the bound structure: the map is a hierarchy of voxel grids at coarser and coarser resolutions, and a coarse voxel upper-bounds the score of every pose inside it. Branches that cannot beat the current best are dropped whole; the rest are refined, with thousands of candidates scored concurrently on the GPU. On fleet data it recovers poses to 0.16 m average error. Its failure mode is the aliasing case above (a genuinely identical aisle produces a genuinely high score), which is why it is not trusted alone.
The camera checks the geometry. Visual place recognition runs open-vocabulary detection, segmentation and embedding models over the scene, keeps the static objects (racks, machines, fixed structure) and stores each object's descriptor and centroid in the map frame. A query image is matched at two levels: descriptor similarity, then geometric consistency of the centroid arrangement. On fresh runs of an already-mapped site its standalone success fell to 30% and 8%, on feature-sparse straights and under dynamic occlusion, which rules it out as a localizer, not as a check: its errors are uncorrelated with LiDAR's, because identical aisles rarely hold identical objects. It is, honestly, the weakest seat on the committee, and the fix is to bound the search: mapping already produces the site's own object list (labels, descriptors, locations), so localization matches against that curated list instead of asking an open-vocabulary model to describe the scene and pick the objects out. A bounded search over known objects, in place of an open question.
The radio map covers the case where geometry and appearance fail together. Access points have fixed identities and position-dependent signal strength, so a one-time survey predicts the WiFi fingerprint at any pose. The verifier compares prediction against what the robot hears and returns confirmed, rejected, or, with fewer than four common access points, abstain. A wrong-aisle recovery puts the robot in range of the wrong access points regardless of what scan and camera suggest. Thresholds self-calibrate against the site's own fingerprint distribution; nothing is tuned per factory. Verdicts are currently logged, not acted on, until they justify gating a recovery, and fingerprints are logged even at sites with no radio map yet, so deployed vehicles accumulate the survey on their own.
The fourth expert is GPS. Under a factory roof it has nothing to say, and abstaining is a first-class answer in this design, so it simply sits out. Outdoors it speaks exactly where the other three thin out: long fences and open tarmac give scan matching little to grip, yards hold few stable objects for the camera, and WiFi fades past the last access point. Outdoor recovery (yards, dock aprons, the routes between buildings) stands to gain the most.
Fusion and performance
The combination is asymmetric: scan matching proposes; the others check. When the experts agree, the pose seeds the localizer and the robot resumes work. When any disagrees, the recovery stops: the robot reports itself lost and waits rather than act on a disputed pose. Confidence here is not a number any component reports about itself; it is consistency across components that fail differently. A committee of experts in the plain sense: independent systems, fixed roles, combined by rule.
The experts run in parallel (the three sensor-heavy ones CUDA-accelerated on the vehicle's GPU) and the full panel returns a pose and a verdict in about five seconds. At five seconds, recovery stops being an operator-assisted procedure and becomes something the robot does whenever it wakes unsure of its position.
Recovery is anchored to stations. A lost robot does not need a pose anywhere on the map; it needs to know which station it is at: that is where trolleys are picked, dropped and parked, and where a recovered pose is immediately useful. Framing the question as which of N stations makes the search sharp and the failure modes measurable in advance: pairwise distances between station signatures, computed offline, show which stations a site could confuse and where thresholds must be tight.
Status
The panel is in shadow mode at two customer sites, logging verdicts against real recoveries. What carries over from this work: no pose is accepted on one system's word, every layer can return "unknown", and disagreement stops the robot instead of letting it guess. It is also a quiet prerequisite for lights-out operation: a fleet cannot run without people until it can get lost, and get found, without one.
Next: Reversing a Trolley, Classically: bidirectional MPC, jackknife constraints, and a problem where the right amount of machine learning was none.

