ROCm / ROCm/aiter

[Bug] gfx942 CK SiLU MoE drops positive swiglu_limit in BF16/block-FP8 paths

Open
#5,502 1 comment 0 reactions 0 assignees View on GitHub

Nobody has claimed this yet.

Dominant language
Python
Stars
565
Forks
585
Avg merge
3d 4h
Merged PRs (30d)
366

Description

Summary

The selected two-stage CK SiLU path accepts a positive swiglu_limit at
fused_moe, but silently computes the unclamped activation. This blocks
qualification of SGLang39247's gfx942 shared-expert opt-in. It is not a claim
that all AITER FP8 paths fail, or that the eligibility patch introduced the bug.

Expected GLM activation: silu(min(gate, limit)) * clamp(up, -limit, limit).
This is not GPT-OSS's alpha/bias SwiGLU; changing the activation enum would not
preserve the model contract.

Public reproducer

SGLang commit 31946c5bb7d3824745cb02b3376f8135f345314b includes
test/manual/diagnose_glm_shared_clamp_gpu.py.
Run PYTHONPATH=python python test/manual/diagnose_glm_shared_clamp_gpu.py.
It invokes the production runner with8 routed +1 shared expert,H512,I256,
BF16 or native FP8 weights, M1/8/16 and clamp0/10. The reference uses FP32
dequantized-weight expert math; positive-clamp inputs are scaled to exercise
the bound. It logs exact selected metadata, module SHA and Torch/HIP versions.
The script is diagnostic: inspect each JSON passed flag, not exit0.

Reproduced again on gfx942 using official
lmsysorg/sglang-rocm@sha256:91ca32f56c662a8520c0ab8f24dffeb6a62ca6037dd91d880f1583a7381b03e9
(v0.5.19-rocm720-mi30x-20260911,Torch2.9.1/ROCm7.2), with the pinned
SGLang source above. No full model is needed.

  • BF16 unclamped M1/8/16 passes, NRMSE about0.0033.
  • FP8 unclamped M8/16 passes, NRMSE about0.0452/0.0433; M8 graphs pass.
  • Clamped BF16 all three M and FP8 M8/16 have NRMSE2.12-2.50 against the
    clamped reference, but match the unclamped reference (BF16 about0.0033,
    FP8 about0.040-0.044).
  • Separate FP8 M1 unclamped failure: NRMSE0.77485 with split-K2. Its cause
    is unresolved and should not be conflated with the clamp omission.

Source boundary and request

The selected ck_moe_stage1 signature has no clamp parameter. The host
dispatcher forwards swiglu_limit to selected FlyDSL/Opus wrappers, not this
CK path. A read-only check of current AITER main
08fd34e82 on September14 still finds that CK signature. We did not run that
new main as a GPU stack; newer MXFP4 metadata changes do not establish BF16/FP8
CK clamp support.

Is there a supported clamp-aware BF16/block-FP8 CK or FlyDSL dispatch for this
contract on gfx942? Otherwise the dispatcher should not silently select a
path that drops a positive bound. We have historical fixed-limit CK prototype
patches, but are not proposing a globally hard-coded limit10 or bundling a
backend rewrite into a shared-expert eligibility PR.

Prior investigation/results are public on
SGLang39247.
This diagnosis, reproducer work and report used Codex assistance.

Contributor guide

No contributing guide indexed for this repository

First steps

  1. Read the whole issue, then the project's contributing guide.
  2. Comment on the issue to say you are picking it up — it saves two people doing the same work.
  3. Fork the repository and make your change on a branch.
  4. Open a pull request that references the issue number.

Research direction

Run test/manual/diagnose_glm_shared_clamp_gpu.py first and inspect each JSON passed flag. Then trace the host dispatcher, selected ck_moe_stage1 signature, and the FlyDSL/Opus wrappers to determine whether a clamp-aware BF16/block-FP8 path exists for gfx942. Done means the positive swiglu_limit contract is supported or the dispatcher no longer silently selects a path that drops it.

Written by the indexing model from the issue text.

Assessment

Tech stack
python, pytorch
Domain
backend, machine-learning
Issue type
Bug
Difficulty
5/5
Estimated time
Over a week
Activity status
Active
Clarity
Mostly clear
Newbie friendliness
35/100

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.