A one-line brief, a week of AI agents, and a lot of cutting.
Andrej Karpathy's autoresearch inspired the shape of this experiment: let an AI agent change one thing, give it a cheap score, and run the loop. Ours proposed a motion maneuver, simulated it against vehicle geometry, and scored whether the robot fits the aisle.
The motivation: "make it turn around"
The Ati Robotics Sherpa XT Lite tugs up to 1,500 kg and pivots in place: a 0 m turn radius inside roughly 1.5 m of swept envelope.
The Sherpa 10K pulls 10,000 lb (≈4,500 kg). Getting there meant a drivetrain rebuilt for the duty cycle and a steered axle in place of independent left-and-right traction. Far more capable hauler, different kinematics, no in-place rotation.
Customers liked the platform, then hit the same wall in every real aisle: "it can't turn around here."
The brief was a single sentence: build a reliable, planner-callable three-point turn (forward, reverse, forward) in the smallest aisle the geometry allows. Over a week, agents wrote most of the spec, math, simulation and tests. Nearly all the work that mattered was taking things out.
Act 1: from a one-liner to a real spec
The first prompt was the brief itself. What came back was a placeholder: the maneuver in words, two constraints, and no contact with how our planner composes routes or what the vehicle can do.
It got serious when we anchored the agent in real data: logs from a manual three-point turn (motor CAN, controller inputs, SLAM poses), with instructions to find the segment structure. Now the spec had ground truth: a known-good human run to fit against and to beat.

Least-squares fits where the operator held steering lock came back almost embarrassingly clean: RMS residuals around 1 cm, R_min ≈ 1.0 m at full steering, implied wheelbase 0.85 m, within a hair of CAD. The bicycle model described this vehicle exactly. Every geometry decision after that had a closed-form check beside it.

Act 2: the simulator as a thinking tool
We wrote a 200-line simulator that read routes in our production format and forward-integrated the kinematics, so anything it accepted, the planner would accept too. That throwaway became the workbench: design a route → expand → simulate → plot.

The design that died first
Our first variant wasn't a three-point turn at all: pre-shift sideways with a Bézier lane change, then a tighter forward/reverse. Clean on paper, dead in minutes. A cubic Bézier with matched end tangents, lateral offset D and longitudinal run Lx has endpoint curvature κ = 8D/(3·Lx²). The run, not the offset, is what kills it. Curvature scales as 1/Lx², and a tight aisle is exactly where Lx is short. It stays in the design folder marked rejected so nobody rediscovers why.
Three things the plots surfaced
- The blend penalty. Our planner's turn primitive is a cubic Hermite blend between tangent vectors. At the arc angles this maneuver uses, peak curvature inside the blend is
1.5/Rrather than1/R. Ask for radius R and the chassis absorbs 50% more curvature. That factor propagates into aisle widths, feasibility checks and the envelope math. - The reverse-leg collapse. In a symmetric turn the reverse leg sweeps the tangent from −ψ to +ψ, driving
Σsin(θ) ≈ 0in our path expansion and collapsing the y-component under the blend rescaling. The result was a flat line where an arc belonged. Fix: split the reverse turn at the arc's midpoint, wheresin(θ_mid) = 0makes tangent continuity exact. - Body envelope vs. rear-axle path. In a plot of a 3.5 × 3.5 m aisle, the rear-axle path stayed cleanly inside the box while the body clipped a corner. Customers care about the swept envelope. Padding by the body diagonal
√(ℓ² + w²)separates the two. A textbook check. What the simulator supplied was the reason to look.
Act 3: shrinking the surface area
The first feasibility analysis had eight knobs: ψ, R, drift, wheelbase, R_min, body length, body width, blend penalty. Far too many to hand a planner caller.
The constraint chain collapsed them. Aisle aspect ratio sets a feasible band for ψ, 60° to 70.5° (arccos(1/3)). Drift goes to zero at the bottom of that band, so ψ pins to 60°, and R follows from aisle width and ψ. The rest are platform constants. What a planner sees is two numbers, aisle width W and length L, and the shorthand ["3pt_turn", W, L] fell out mechanically.
Then we stopped generating and started auditing. Three questions, all answered by grep:
- Where does the solver compose the dispatchable route? Exactly one site, so a two-line hook suffices.
- Are there existing three-point-shaped waypoints? Zero, so there was no migration burden.
- Does anything already own the name? One function, parking-specific, different geometry, no conflict.
That cut the integration plan from "redesign the route schema" to "add two lines."
Act 4: handing it off to another AI
The spec went to a coding agent, packaged as documentation, tests and constraints. That was enough for the agent to build the feature without inventing anything, and without being able to disturb existing route planning.
The piece that made it work was a regression net: every existing route had to come through the new code completely unchanged. With that guarantee in place, handing the implementation to an agent stopped being a risk decision.
The handoff itself was a self-contained brief: the spec, the tests, and the handful of constraints not derivable from either. Those were written down explicitly rather than left for the agent to infer.
What we actually learned
The headline isn't "AI wrote the PRD." AI is very good at volume: long specs, alternative designs, exhaustive test grids. What made this work was repeated cutting: killing the lane-change variant, collapsing eight knobs to two, deleting scratch directories, carving the newer format out of the first PR. That loop of proposing, simulating, scoring and discarding is what Physical AI looks like at the geometry level.
Real data is the cheapest thing you can give an agent.
The spec was guesswork until we handed the agent logs from a manual run.
A throwaway simulator pays for itself in a week.
Two hundred lines, deleted the day the PR landed. It surfaced three bugs and settled every design tradeoff.
Agent-ready handoffs demand more discipline than human ones.
Every ambiguity is a place the agent guesses, and guesses propagate. Consistent naming, sourced-vs-derived constants and an explicit list of non-obvious constraints separate a clean landing from a week of rework.
The 10K now turns around in an aisle barely bigger than its own footprint, behind a planner call two numbers wide.
Customers asked us to teach it to drive like a car. The harder part was teaching ourselves to write a spec like an engineer.


