THUDM / THUDM/slime

[Feature Request] Add minimal 1-GPU script for debugging and learning

Open Beginner friendly
#1,628 0 comments 1 reaction 0 assignees View on GitHub

Nobody has claimed this yet.

Dominant language
Python
Stars
8.5k
Forks
1.3k
Avg merge
5h 36m
Merged PRs (30d)
22

Description

Motivation

All existing run scripts under scripts/ require 4-8 GPUs (e.g., run-qwen2.5-0.5B-reproducibility.sh uses 8 GPUs, test scripts use 4 GPUs). This makes it difficult for new contributors to:

  • Understand the slime training flow on a single development GPU
  • Debug issues without reserving a multi-GPU node
  • Quickly iterate when developing new features (e.g., adding model support)
Proposal

Add a minimal 1-GPU run script scripts/run-qwen3-0.6B-minimal.sh that:

  • Uses Qwen3-0.6B (smallest supported Qwen3 model)
  • Runs on a single GPU with --tensor-model-parallel-size 1, --pipeline-model-parallel-size 1
  • Uses small batch sizes (global batch size 16, rollout batch size 4) and short response lengths (512 tokens)
  • Covers the full GRPO training loop (rollout + training + eval) on GSM8K
Existing alternatives
Script GPUs Limitation
scripts/run-qwen2.5-0.5B-reproducibility.sh 8 Designed for reproducibility, not minimal debugging
examples/true_on_policy/run_simple.py 1 (default) Only covers true on-policy mode, not the standard training flow
tests/test_qwen2.5_0.5B_debug_rollout_then_train.py 8 Test-oriented, not a standalone learning script

None of these provide a simple 1-GPU script for the standard training flow.

Additional Context

I'm happy to submit a PR for this if the team thinks it would be useful. The script is a single file with no changes to existing code.

Contributor guide

Open the contributing guide

First steps

  1. Read the whole issue, then the project's contributing guide.
  2. Comment on the issue to say you are picking it up — it saves two people doing the same work.
  3. Fork the repository and make your change on a branch.
  4. Open a pull request that references the issue number.

Research direction

Compare scripts/run-qwen2.5-0.5B-reproducibility.sh with examples/true_on_policy/run_simple.py and tests/test_qwen2.5_0.5B_debug_rollout_then_train.py to identify the standard training flow and required invocation. Add scripts/run-qwen3-0.6B-minimal.sh with the listed one-GPU, batch-size, response-length, and GSM8K settings. Done means rollout, training, and evaluation complete on one GPU.

Written by the indexing model from the issue text.

Assessment

Tech stack
bash, python
Domain
machine-learning
Issue type
Feature
Difficulty
2/5
Estimated time
1-3 hours
Activity status
Stale
Clarity
Clearly specified
Newbie friendliness
68/100

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.