Request: shareable RL post-training throughput benchmarks + sizing guidance across Hopper/Blackwell/Vera
- Dominant language
- Python
- Stars
- 2k
- Forks
- 561
- Avg merge
- 4d 5h
- Merged PRs (30d)
- 145
Description
### Summary
A customer is doing AI-factory sizing for post-training and is asking whether NVIDIA has
measured, externally-shareable **RL throughput** benchmarks with NeMo-RL across recent GPU
generations (Hopper, Blackwell, Vera), plus general sizing / performance-scaling guidance on
newer GPU portfolios.
### Questions
1. Are there externally shareable RL throughput benchmarks across Hopper, Blackwell, and Vera?
2. Is there sizing guidance for RL post-training to estimate cluster/GPU count for a target
throughput?
3. What performance/efficiency gains should be expected moving Hopper -> Blackwell -> Vera
for RL workloads?
### Context
- Use case: customer AI-factory sizing for post-training.
- Raised in #swdl-nemofw-support on 2026-07-13 by Flora Huang (florah@nvidia.com).
- Slack thread: https://nvidia.slack.com/archives/C083Y1XED43/p1783956056969079
- Companion request for **SFT/LoRA** throughput filed on NVIDIA-NeMo/Megatron-Bridge.
### Type
Question / documentation request (not a bug).
Contributor guide
Assessment
This issue has not been assessed yet.