Reproducing part of ROSA, and what did not reproduce
In July, NVIDIA Research and Stanford published ROSA (arXiv 2607.01088): a serving system for robot fleets where many robots share a pool of GPUs instead of each carrying its own. I have spent the summer on the operator's side of exactly that question - what happens on a floor when the models a fleet needs stop fitting on the robot - so I tried to reproduce one row of their results on hardware I could actually get. These are my notes on what held, what did not, and what I think a fleet operator should take from it. Everything below is from my own runs. Where a number from the paper appears, it appears with the paper's hardware and model, and never in the same sentence as one of mine.
What I set out to reproduce
Table 1, P4, the paper's "assemble kit" pipeline: a System-1 action model at a 200 ms p99 deadline, a System-2 planner called once per ten System-1 calls at 2,000 ms, a safety judge at 2 Hz / 500 ms, a task monitor at 0.5 Hz / 2,000 ms, with 200 ms of physical action per step. Two arms, in the shape of the paper's Figure 10: a capped arm, where each robot's send rate is held at the rate the scheduler predicts the pool can qualify, and an uncapped arm with no rate control. A "qualified action" means both the latency deadline and the invocation rate were met, the paper's own definition in §4.2.
My rig, which is not theirs
- Two single-GPU H200 virtual machines from a cloud provider, one Ray Serve replica, CUDA graphs on. A separate 8-vCPU host for ingress and the load generator. The paper's experiments run on one 8x H200 server.
- Models: the planner and monitor on Qwen2.5-VL-7B-Instruct and the safety judge on Qwen2.5-VL-3B-Instruct, which are the models the paper's Figure 3 declares for those roles. System-1 is my one stand-in: Qwen2.5-VL-7B-Instruct in place of GR00T N1.6. That substitution is the biggest gap between my runs and the paper, and it is why none of my numbers belong beside theirs.
- One fixed prompt and a capped 8-token reply, not real camera observations.
- Latency tables profiled once on this rig type and frozen for both arms: batch sizes 1 through 16, ten request-rate fractions per batch size, four roles.
- The scheduler is my reimplementation of what the paper describes - the binary search over a shared action rate with a packing step (§3.4.2) - not the authors' code, which had not been released when I ran this.
What reproduced
At 16 robots on P4, both arms, across two separate runs and three scored points per arm: the capped arm qualified 62.08 actions per second per System-1 GPU and the uncapped arm 80.00, identical to two decimals across all three points. My rig has one System-1 GPU, so that is also the whole-rig figure, on two single-GPU H200 VMs with a Qwen2.5-VL stand-in for System-1.
My reading, not a measurement: uncapped ahead of capped at 16 robots is what you would expect before the pool saturates. The paper reports the same direction at its lowest robot count, 8, where the uncapped baselines edge the scheduled system on 8x H200 with GR00T N1.6 (§4.2).
What did not reproduce, and why
At 32 and 64 robots my runs excluded themselves. A guard I had written before the run flagged every 32- and 64-robot point as driver-bound: the busiest load-generator process sat at 94 to 100 percent of a CPU core. I believe the backlog was inside the generator rather than on the GPUs, but that is inferred, not measured. So I have no valid statement about 32 or 64 robots, and I am not going to make one.
I think the root cause is a fidelity error in my client, not the hardware. Both of my arms fired System-1 requests on a fixed 200 ms schedule whether or not the previous action had returned. The paper's §4.3 reports 145.5 raw actions per second at 64 robots for its uncapped baseline on 8x H200 with GR00T N1.6. A client firing on a fixed schedule regardless of returns would produce far more than that, so I now read the paper's uncapped client as one observation in flight per robot with no rate cap on top. Mine had no cap and no in-flight limit, which is a stricter load than the paper applied. The fix is designed and not yet run: one in-flight request per robot in both arms, an in-flight gauge written into the run record, and a CPU-only validation before the next GPU window.
What I take from it, as someone who talks to operators
Three things, held loosely.
The models were not the hard part. Standing up four roles behind a gateway and getting qualified actions out at 16 robots took days. The scheduling did the work, and the paper is candid that the scheduling is where the difficulty lives: ILP-based packing per configuration (§3.4.2), a combinatorial search over mixed task configurations that they prune rather than enumerate (§3.4.3), and a runtime re-solve they describe as time-consuming and disruptive to robots that were otherwise unaffected (§3.4.4).
Rate control lives in the robot, not the server. The capped arm only exists if the robot client agrees to pace itself. On a real floor that is a conversation with whoever owns the robot stack, and it should happen before anyone promises a p99.
The 16-robot row is a server-side number. It says nothing about the network between the robot and the pool, and the paper is explicit that its assumptions hold for stable industrial networks (§3.2). A fleet SLO is written at the robot, and the hop from the GPU to the robot is outside everything above.
What I would ask an operator
If you run a fleet that serves models off the robot today, I would like to know two things: how the serving rate is set at the robot end, and what the request log looks like for one shift. The second is the input that would turn these notes from a reproduction into a sizing. I read every reply myself.
The longer argument for why this layer belongs on the floor is here: nectaredge.co/why-fleet-compute/
Run artifacts: pp20230 (two ingress shapes) and pp38996, 2026-09. Checked against the paper's Figure 3 and Figure 10.