NanoDrive
Fast multi-agent traffic simulation

An autoregressive model for joint traffic motion, with cached temporal attention and spatial relations recomputed at each step.

Loading rollout…8 seconds · joint simulation
7.6Mlearned parameters
18–21×scene exposure in 25 minutes¹
3.71×H100 generation speed²
~40 MBFP32 model + motion codebook

SMART and NanoDrive

SMART on the left; NanoDrive on the right. Ego is yellow.

Saved PyTorch rollouts from the 100-minute NanoDrive and SMART checkpoints. Eight display scenes; these videos do not run inference.

Generate a rollout

Generate eight seconds of joint motion on your device with the actual NanoDrive model.

SMART is a saved comparison sample. The NanoDrive view starts at the observed scene; Generate runs new ONNX inference in your browser. Model files download on the first run.

0.0 s

Architecture

NanoDrive uses optimized causal attention, cached temporal keys and values, and compact spatial neighborhoods to reduce repeated work compared with SMART. Spatial relations are recomputed from current poses at each step.

NanoDrive architecture across two time steps: reusable map encoding, temporal keys and values, and observed spatial interactions at 0.5 and 1.0 seconds

Training and inference

Matched CPU and GPU measurements, plus training runs up to 200 minutes.

CPU inference · PyTorch

Matched PyTorch CPU latency for SMART and NanoDrive on four scenes

Xeon 8562Y+ · 8 threads · FP32 · one future · median and range of 3 runs.

GPU inference

Paired H100 and L40S generation latency across future batch sizes
Generation benchmark details

² Generating 32 joint futures at eight concurrent futures: 3.71× on H100 and 5.04× on L40S, FP32, 16 scenarios, same device within each comparison. SMART uses the NVlabs/CAT-K behavior-cloning implementation with token reuse. Timings use the older 100-minute checkpoints; the corrected-input quality pilots below use separate 25-minute checkpoints. These results do not establish equal-quality speedup at convergence.

Longer training runs

NanoDrive and SMART WOSAC scores at 25, 100 and 200 minutes, and scene visits during training

Legacy prepared input · H100 · 512 development scenes. Quality: independent budget runs, mean and range of two seeds. Exposure: seed 1817.

Training lossSeparate NanoDrive and SMART training-loss curves across 200 minutes

Seed 1817 · 1-minute means. The models use different loss targets and recipes.

Corrected-input 25-minute pilots
SeedNanoDrive scene visitsSMART scene visitsExposure ratioRealism · NanoDrive / SMART
817536,24829,85418.0×0.7338 / 0.7114
1817606,74628,36421.4×0.7392 / 0.6880
Training comparison details

¹ Repeated presentations of 15,624 distinct training scenes. Systems differ in physical batches, targets, precision, and optimization. Offline preparation and evaluation are excluded. Seed 817 resumed; retained time omits some discarded work. Realism is WOSAC-2025 on 128 repeatedly consulted development scenarios.

Measurements and run provenance