Determine whether a large dynamics RL batch refusal is necessary or avoidable by execution layout
Nobody has claimed this yet.
- Dominant language
- Python
- Stars
- 10.8k
- Forks
- 989
- Avg merge
- 6h 29m
- Merged PRs (30d)
- 85
Description
Status reconciliation — September 11, 2026 (Schulman)
Open feasibility investigation; no incorrect refusal has been established. The exact historical replay remains blocked by missing fixture/profile state. Existing phased-execution comparisons do not prove a memory improvement or matching gradient execution.
Planned lane: Schulman and subagents, lower priority than reproducible #848/#870 boundaries. Recover an exact witness from retained evidence if possible; otherwise state that limit and use a separately identified representative case, without claiming reconstruction. Do not spend new GPU work trying to reproduce an unbound historical state.
Historical report (preserved):
Type: feasibility investigation, not a demonstrated incorrect refusal.
Owner: Schulman. Exact-witness replay is blocked by missing historical state; general planner work continues under #848.
Shannon's retail49 R0e run stopped on a TrainerRankMemoryError refusal. Preserved run-state v300 proves successful updates through batch 53; batches 54–56 contain rollout-only records and next_number is 57. Thus 57 describes driver progress, not a proven failed optimizer step. The guard raised before this attempted execution; this is distinct from a hard CUDA OOM.
The log reports 208288 packed / 568147 logical tokens, predicted 114.938 GiB against usable 63.101 GiB on one H200. Under the inspected source, token geometry describes the minimum wave's unsplit full-sharing plan, while the budget belongs to the final rejected split rung. At DP1, the matching 048 code makes one pair a top-level item, so this would be one indivisible pair, not the full eight-pair update. Historical ART/runtime-byte identity remains unconfirmed.
Preserve the exact failing batch/checkpoint and source/ART identities. Reconstruct top-level pair boundaries, per-history logical/packed lengths, request mix, sharing depth, cold/profiled state and original admitted/refused budget. Determine whether an indivisible input/request group forces the refusal, whether retained state is necessary, and whether a semantically identical execution schedule fits.
A reduced-history cap, changed pair weight or substitution of prepass values changes the experimental workload/objective; do not call such a change an equivalent execution fix. An auxiliary-first/generator-second backward schedule with one optimizer step is a candidate, not a qualified solution. If refusal is necessary, document the measured boundary and supported alternatives.
Evidence: /home/brad/shannon-logs/report-2026-09-09.md (R0e), associated run logs, and /home/brad/.local/share/schulman/retail49-program-20260909/analysis/separate-lora-phases-scope/REPORT.md. Cross-reference #848's broader cold/warm admission work. No GPU execution was performed to file this issue.
Current evidence gap: exact generated histories/tensors, selected profile/budget state and the pre-update G/aux state were not recovered. Queue indices and cal.log_trajectories metrics do not establish preserved raw inputs. No equivalent replay or necessary-refusal conclusion is claimed. Existing reconstruction: /home/brad/.local/share/schulman/memory-planning-audit-20260909/REPORT.md (SHA256 4c65c3c1e861eaf90286a19ecbdf43a53cdc138621184f2c3e8407f352b5fa95); concise evidence in adjacent comment-869-progress.md.
Contributor guide
First steps
- Read the whole issue, then the project's contributing guide.
- Comment on the issue to say you are picking it up — it saves two people doing the same work.
- Fork the repository and make your change on a branch.
- Open a pull request that references the issue number.
Research direction
Start with /home/brad/.local/share/schulman/memory-planning-audit-20260909/REPORT.md and adjacent comment-869-progress.md, then review the R0e report and separate-lora-phases-scope/REPORT.md. Recover the exact fixture, profile, budget, and state if possible; done means clearly documenting whether the refusal is necessary or an equivalent execution layout fits, without changing the workload or claiming an unverified replay.
Written by the indexing model from the issue text.
Assessment
- Tech stack
- python
- Domain
- machine-learning, performance
- Issue type
- Bug
- Difficulty
- 5/5
- Estimated time
- Over a week
- Activity status
- Active
- Clarity
- Mostly clear
- Newbie friendliness
- 25/100