Rigyd assets at training scale: 8,192 environments on one GPU
Throughput, memory and stability of Rigyd exports in a Franka scene at 1,024 to 8,192 parallel environments on one GPU, measured on MuJoCo Warp and Isaac Lab.
Rigyd Engineering · Published September 29, 2026
When you can measure what you are speaking about, and express it in numbers, you know something about it.
A question robotics teams ask before trusting a third-party asset: how does it behave when a training run copies its scene thousands of times onto one GPU? This post answers with numbers instead of a claim.
What it buys: a rollout of 10 million environment-steps (a typical PPO training budget for a manipulation task) costs 10M / 270k = 37 s of physics on this card. Policy inference and gradient updates come on top; for a small policy those are usually of the same order, so physics is not the bottleneck at this rate.
Hardware and software: one NVIDIA A100 80 GB PCIe (Azure NC24ads_A100_v4) and one Tesla T4 16 GB (NC4as_T4_v3). MuJoCo Warp 3.13 · Isaac Sim 5.1 + Isaac Lab 2.3.2 (PhysX GPU pipeline) · MuJoCo 3.13. Raw data and the benchmark protocol are available on request.
What the benchmark measures
The setup. A Rigyd object on a table next to a robot arm that moves at random and knocks it about. The scene is copied 1,024 to 8,192 times on one graphics card; each copy is one “environment”, the unit robot-learning teams count in.
Speed. How long one step of physics takes for all copies at once, and how many environment-steps per second that gives. Training time is set by this number.
Memory. How much of the card each copy uses, and the largest number of copies that fits.
Stability. Whether anything breaks over 10,000 steps: a simulation that produces “not a number”, an object sinking into the table, a hinge forced past its stop, an object flung off into the distance.
Fairness. Every number is the difference against the same scene with no object, so the robot is not charged to the asset. Three objects, from a plain box to a Rubik’s cube with 310 collision pieces, on the two simulators most robot-learning teams train with: MuJoCo Warp and NVIDIA Isaac Lab.
What the results say
The assets work at training scale. The Rubik’s cube, our hardest export, runs 8,192 copies on one A100 card in both simulators, and none of 4,096 copies broke in 10,000 steps.
Speed is competitive where it counts. On MuJoCo Warp the cube adds 26 ms per step to an empty scene; 270,000 environment-steps per second is training-grade throughput. Isaac Lab is about three times slower on the same asset, a property of that simulator’s collision engine, not of the export.
The cost driver is known. The number of collision pieces sets both speed and memory; a 310-piece cube costs 3 to 6 times a 37-piece mug. That is a dial we control at export time.
Absolute environments-per-GPU depend on the robot, the policy and the observation stack, which belong to the training team. We measure what the asset adds to a fixed reference scene and report the delta.
Method: the asset’s marginal cost in a fixed scene
Reference scene
Ground plane, 0.6 × 0.6 m table with its top at 0.4 m, Franka Panda on the table edge (MuJoCo Menagerie in Warp, Isaac Lab’s own Franka in Isaac).
Asset 2 cm above the table centre; every 50 steps new uniform random targets for the 7 arm joints, gripper open; every 500 steps a 0–2 N impulse on the asset root held 10 steps.
Per-environment randomization at reset: mass × U(0.8, 1.2), sliding friction U(0.3, 0.9). Timestep 0.002 s as exported, one physics step per environment step, headless, no rendering.
Assets, as exported
Baseline (a): the scene with no asset. Baseline (b): a 6 cm, 100 g box. Mug: rigid, 37 CoACD hulls. Rubik’s cube: 7 bodies, 6 hinges, 310 hulls, 16.5k hull vertices.
Preparation only: visual geoms stripped, the export’s own ground plane removed, a free joint on the root. Contact excludes, solref/solimp, friction, damping and armature stay as shipped.
Measurements per stack, tier and N
- Throughput: 100 warm-up steps, then the median of three 1,000-step runs, CUDA-graph captured in Warp.
- GPU memory: nvidia-smi at steady state minus idle, divided by N.
- Max N: doubling in a fresh process until allocation fails or a step exceeds 100 ms.
- Stability: 10,000 steps at N = 4,096; an environment is flagged on any NaN, an asset contact deeper than 5 mm (Warp only, PhysX exposes forces not depths), a hinge more than 5° past its limit, or the asset more than 2 m from the table.
- Inventory: bodies, joints, hulls, vertices, and the contact and constraint capacity the stack allocated (4× the CPU-observed maximum, the spec’s rule).
Two collision variants in Isaac
The Rigyd USD exports carry two collision variants, and Isaac numbers are given for both:
hulls uses the same CoACD hulls as the MJCF export and is the apples-to-apples
comparison with Warp; sdf uses signed-distance fields on the visual meshes at
resolution 256.
Result 1 of 4: the Rubik’s cube, 8,192 environments on one A100 in both stacks
Warp steps 8,192 cubes in 30 ms (270k environment-steps per second); Isaac Lab needs 95 ms with the same hulls and 204 ms with the SDF variant. Neither produced a NaN.
| Environments | Warp A100 | Isaac Lab A100, hulls | Isaac Lab A100, SDF | Warp T4 |
|---|---|---|---|---|
| 1,024 | 6.0 ms · 171k/s · 9.4 MB | 36.8 ms · 28k/s · 6.8 MB | 92.3 ms · 11k/s · 10.4 MB | 10.0 ms · 102k/s · 9.0 MB |
| 2,048 | 8.5 ms · 241k/s · 9.1 MB | 46.7 ms · 44k/s · 4.0 MB | 119.4 ms · 17k/s · 7.5 MB | out of memory |
| 4,096 | 14.5 ms · 282k/s · 8.9 MB | 61.4 ms · 67k/s · 2.6 MB | 150.1 ms · 27k/s · 4.0 MB | out of memory |
| 8,192 | 30.4 ms · 270k/s · 8.8 MB | 94.9 ms · 86k/s · 2.3 MB | 204.3 ms · 40k/s · 3.7 MB | out of memory |
Each cell is milliseconds per step, thousand environment-steps per second, and GPU memory per environment.
Warp’s 8.8 MB per environment is the contact-capacity rule, not the hulls: 4 × the 238 contacts observed on CPU = 952 contact slots per environment, and Warp’s collision scratch is sized by that. With a 2× rule the cube costs 4.5 MB per environment (28 ms per step) and reaches 16,384 environments. The SDF step time grows during an episode as cubes get knocked around, so its 200-step max-N trial (97 ms) understates the 3,000-step number (204 ms).
Result 2 of 4: asset cost, what each export adds to the empty scene
At 8,192 environments the cube adds 26 ms per step in Warp and 89 ms in PhysX; the mug adds 6 ms in Warp and 14 ms in PhysX. Hull count drives the cost in both stacks.
| Asset, minus baseline (a) | Warp A100 | Isaac Lab A100, hulls | Isaac Lab A100, SDF | Warp T4 |
|---|---|---|---|---|
| 6 cm box (baseline b) | +0.9 ms · +0.01 MB | +4.3 ms · +0.00 MB | +4.3 ms · +0.00 MB | +7.6 ms · +0.01 MB |
| Mug, 37 hulls | +6.2 ms · +0.75 MB | +13.9 ms · +0.31 MB | +22.3 ms · +2.00 MB | +28.8 ms · +0.75 MB |
| Rubik’s cube, 310 hulls | +26.4 ms · +8.22 MB | +88.7 ms · +0.94 MB | +198.0 ms · +2.36 MB | out of memory at 2,048; +7.1 ms at 1,024 |
A100 baselines: Warp 4.0 ms per step and 0.60 MB per environment, Isaac Lab 6.2 ms and 1.32 MB. Each cell is the tier minus its own stack’s baseline, at 8,192 environments.
The lighter tiers stay fast. The empty scene reaches 2.1 M environment-steps per second on Warp and 1.3 M on Isaac Lab at 8,192; the box and the mug cost little in Warp.
| Thousand environment-steps per second at 8,192 | Warp A100 | Isaac Lab A100, hulls | Isaac Lab A100, SDF | Warp T4 |
|---|---|---|---|---|
| No asset (baseline a) | 2,054k | 1,317k | 1,317k | 921k |
| 6 cm box (baseline b) | 1,669k | 779k | 779k | 498k |
| Mug, 37 hulls | 800k | 408k | 287k | 217k |
Result 3 of 4: max environments and memory
Both stacks hold 8,192 cubes on an 80 GB A100; Warp is memory-bound (72 GB under the 4× rule), PhysX is time-bound (16,384 runs but at 159 ms per step). The T4 holds 1,024 cubes.
| Max N passing | Warp A100 | Isaac Lab A100, hulls | Isaac Lab A100, SDF | Warp T4 |
|---|---|---|---|---|
| No asset (baseline a) | 131,072 (24 ms, 69 GB; next: memory) | 32,768 (12 ms, 41 GB; next: cap) | 32,768 (12 ms, 41 GB; next: cap) | 16,384 (10 ms, 9 GB; next: memory) |
| 6 cm box (baseline b) | 131,072 (34 ms, 70 GB; next: memory) | 32,768 (18 ms, 41 GB; next: memory) | 32,768 (18 ms, 41 GB; next: memory) | 16,384 (17 ms, 9 GB; next: memory) |
| Mug, 37 hulls | 32,768 (46 ms, 42 GB; next: memory) | 32,768 (38 ms, 51 GB; next: cap) | 32,768 (42 ms, 58 GB; next: cap) | 8,192 (30 ms, 11 GB; next: memory) |
| Rubik’s cube, 310 hulls | 8,192 (27 ms, 72 GB; next: memory) | 8,192 (84 ms, 18 GB; next: time) | 8,192 (97 ms, 29 GB; next: time) | 1,024 (9 ms, 9 GB; next: memory) |
“Next” says why the following doubling was not accepted: memory = allocation failed, time = slower than 100 ms per step, cap = not attempted because Isaac’s USD stage creation for 65,536 environments ran past an hour. Isaac’s per-environment memory is smaller than Warp’s because PhysX shares cooked hulls and sizes its buffers globally; its fixed ~5 GB Kit process is removed by the baseline subtraction.
Result 4 of 4: stability over 10,000 steps
No tier produced a NaN in either stack. The cube is the best-behaved asset in both: fewer flags than the primitive box in Warp, one flung environment in Isaac.
| Flagged environments, N = 4,096 | Warp A100 | Isaac Lab A100, hulls | Isaac Lab A100, SDF | Warp T4 |
|---|---|---|---|---|
| 6 cm box (baseline b) | 11.2 % · 431 deep contacts, 33 flung | 0.0 % · clean | 0.0 % · clean | 11.5 % · 440 deep contacts, 33 flung |
| Mug, 37 hulls | 16.6 % · 633 deep contacts, 52 flung | 0.2 % · 7 flung | 0.1 % · 6 flung | 16.6 % · 621 deep contacts, 64 flung |
| Rubik’s cube, 310 hulls | 8.1 % · 296 deep contacts, 38 flung | 0.0 % · 1 flung | 0.1 % · 3 flung | 7.3 % · 66 deep contacts, 9 flung at N = 1,024 |
The Warp and Isaac columns are not comparable with each other: the Warp scene drives the Menagerie Franka at kp = 4,500, which pins objects against the table and launches them; Isaac Lab’s Franka config uses stiffness 80, so its arm is gentle. Within a stack, compare each tier with the box row: the box carries the scene’s own violence. Warp’s “deep contact” flag is mostly transient impact penetration (a 0.4 m fall off the table lands at 2.8 m/s, 5.6 mm per 2 ms step); the cube’s deep contacts are with the arm and the floor, only 210 of 9,800 with the table it rests on.
Finding: hull count is the cost
310 hulls on a 57 mm cube cost 3× the mug per step in Warp and 6× in PhysX, and set the contact capacity that dominates Warp’s memory. In PhysX the SDF variant is a further 1.4 to 2.2× on top of hulls. Both are dials set at export time, which is what makes the cost predictable per asset.
Inventory and capacity (Warp, as allocated)
| Asset | Bodies | Joints (free + hinges) | Hulls | Hull vertices | Mass | Contacts: CPU max / slots | Constraints: CPU max / slots |
|---|---|---|---|---|---|---|---|
| No asset (baseline a) | 0 | 0 | 0 | 0 | 0 g | 0 / 64 | 3 / 256 |
| 6 cm box (baseline b) | 1 | 1 | 1 | 0 | 100 g | 4 / 64 | 19 / 256 |
| Mug, 37 hulls | 1 | 1 | 37 | 2,321 | 100 g | 35 / 140 | 141 / 564 |
| Rubik’s cube, 310 hulls | 7 | 7 | 310 | 16,459 | 90 g | 238 / 952 | 959 / 3836 |
All exports ship timestep 0.002 s, Euler, Newton solver with 100 iterations, pyramidal cone; the scene keeps these. PhysX ran TGS at the same timestep with Isaac Lab’s defaults; buffers were raised so no overflow message appeared in any reported run.
Generate an asset, drop it into your MuJoCo or Isaac Lab training scene, and tell us what you measure. Raw data and the protocol are available on request.