[Question] Clarification on FP8 Micro-block Scaling and FP4 Support Timeline
Nobody has claimed this yet.
- Dominant language
- Python
- Stars
- 2.2k
- Forks
- 155
- PR merge metrics
- No merged PRs in 30d
Description
Hi cuTile team,
I have two specific questions regarding the support for Blackwell-specific hardware features:
- Automatic Micro-block Scaling for FP8
When using fp8 with ct.matmul, how is the Micro-block Scaling (1x16) handled?
Automation: Does the tileiras compiler automatically handle the scaling logic and hardware invocation (5th-gen Tensor Core) under the hood?
Explicit Scaling: If it is not fully automatic, how should we provide the scale-factor tiles to the ct.matmul operator? Currently, the ct.matmul(A, B) signature seems to only accept data tiles. Is there a plan for a signature like ct.matmul(A, B, A_scale, B_scale)?
- NVFP4 (FP4) Support Roadmap
The current documentation and samples focus on fp8 and bf16. Since Blackwell's throughput peak is tied to NVFP4:
When can we expect the support for 4-bit narrow-precision tiles in cuTile Python?
Thanks for this great library!
Contributor guide
First steps
- Read the whole issue, then the project's contributing guide.
- Comment on the issue to say you are picking it up — it saves two people doing the same work.
- Fork the repository and make your change on a branch.
- Open a pull request that references the issue number.
Research direction
Review the current cuTile Python documentation and samples for fp8, bf16, and ct.matmul, then compare them with the issue's questions about Micro-block Scaling and NVFP4. Done means documenting whether scaling is automatic, how scale-factor tiles are supplied if not, and the expected timeline for FP4 tile support.
Written by the indexing model from the issue text.
Assessment
- Tech stack
- python
- Domain
- compilers
- Issue type
- Documentation
- Difficulty
- 5/5
- Estimated time
- Over a week
- Activity status
- Quiet
- Clarity
- Needs clarification
- Newbie friendliness
- 25/100