NanoDrive
Fast multi-agent traffic simulation
An autoregressive model for joint traffic motion, with cached temporal attention and spatial relations recomputed at each step.
SMART and NanoDrive
SMART on the left; NanoDrive on the right. Ego is yellow.
Saved PyTorch rollouts from the 100-minute NanoDrive and SMART checkpoints. Eight display scenes; these videos do not run inference.
Generate a rollout
Generate eight seconds of joint motion on your device with the actual NanoDrive model.
SMART is a saved comparison sample. The NanoDrive view starts at the observed scene; Generate runs new ONNX inference in your browser. Model files download on the first run.
Architecture
NanoDrive uses optimized causal attention, cached temporal keys and values, and compact spatial neighborhoods to reduce repeated work compared with SMART. Spatial relations are recomputed from current poses at each step.
Training and inference
Matched CPU and GPU measurements, plus training runs up to 200 minutes.
CPU inference · PyTorch
Xeon 8562Y+ · 8 threads · FP32 · one future · median and range of 3 runs.
GPU inference
Generation benchmark details
² Generating 32 joint futures at eight concurrent futures: 3.71× on H100 and 5.04× on L40S, FP32, 16 scenarios, same device within each comparison. SMART uses the NVlabs/CAT-K behavior-cloning implementation with token reuse. Timings use the older 100-minute checkpoints; the corrected-input quality pilots below use separate 25-minute checkpoints. These results do not establish equal-quality speedup at convergence.
Longer training runs
Legacy prepared input · H100 · 512 development scenes. Quality: independent budget runs, mean and range of two seeds. Exposure: seed 1817.
Training loss
Seed 1817 · 1-minute means. The models use different loss targets and recipes.
Corrected-input 25-minute pilots
| Seed | NanoDrive scene visits | SMART scene visits | Exposure ratio | Realism · NanoDrive / SMART |
|---|---|---|---|---|
| 817 | 536,248 | 29,854 | 18.0× | 0.7338 / 0.7114 |
| 1817 | 606,746 | 28,364 | 21.4× | 0.7392 / 0.6880 |
Training comparison details
¹ Repeated presentations of 15,624 distinct training scenes. Systems differ in physical batches, targets, precision, and optimization. Offline preparation and evaluation are excluded. Seed 817 resumed; retained time omits some discarded work. Realism is WOSAC-2025 on 128 repeatedly consulted development scenarios.