llnl / llnl/professor

Study MI300A partition-mode effects on training

Open
#6 1 comment 0 reactions 0 assignees View on GitHub
study
Dominant language
Python
Stars
2
Forks
1
Avg merge
4d 21h
Merged PRs (30d)
1

Description

The MI300A exposes memory/compute partitioning modes that change how memory bandwidth and capacity are presented to processes. These may interact with unified-memory behavior, DataLoader host-memory pressure, and per-rank throughput in ways discrete GPUs don't exhibit.

Possible study: enumerate available partition modes on the target platform, run a fixed training configuration under each, and report throughput, memory behavior, and any failure modes.

Contributor guide

No contributing guide indexed for this repository

Research direction

Start by identifying the target MI300A platform and the fixed training configuration to use for the study. Enumerate its available partition modes, run the same training workload under each, and report throughput, memory behavior, DataLoader host-memory pressure, per-rank effects, and failure modes.

Written by the indexing model from the issue text.

Assessment

Tech stack
python
Domain
machine-learning, performance
Issue type
Feature
Difficulty
4/5
Estimated time
3-5 days
Activity status
Quiet
Clarity
Mostly clear
Newbie friendliness
42/100

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.