NVIDIA-NeMo / NVIDIA-NeMo/RL

Request: shareable RL post-training throughput benchmarks + sizing guidance across Hopper/Blackwell/Vera

Open
#3,180 0 comments 0 reactions 1 assignee Claimed by @terrykong View on GitHub
Documentation Speed
Dominant language
Python
Stars
2k
Forks
561
Avg merge
4d 5h
Merged PRs (30d)
145

Description

### Summary
A customer is doing AI-factory sizing for post-training and is asking whether NVIDIA has
measured, externally-shareable **RL throughput** benchmarks with NeMo-RL across recent GPU
generations (Hopper, Blackwell, Vera), plus general sizing / performance-scaling guidance on
newer GPU portfolios.

### Questions
1. Are there externally shareable RL throughput benchmarks across Hopper, Blackwell, and Vera?
2. Is there sizing guidance for RL post-training to estimate cluster/GPU count for a target
throughput?
3. What performance/efficiency gains should be expected moving Hopper -> Blackwell -> Vera
for RL workloads?

### Context
- Use case: customer AI-factory sizing for post-training.
- Raised in #swdl-nemofw-support on 2026-07-13 by Flora Huang (florah@nvidia.com).
- Slack thread: https://nvidia.slack.com/archives/C083Y1XED43/p1783956056969079
- Companion request for **SFT/LoRA** throughput filed on NVIDIA-NeMo/Megatron-Bridge.

### Type
Question / documentation request (not a bug).

Contributor guide

Open the contributing guide

Assessment

This issue has not been assessed yet.

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.