`rocFFT`: plan handle cache (`IDLE_HANDLES`) never evicts distinct-shape plans, leaking GPU memory
Nobody has claimed this yet.
- Dominant language
- Julia
- Stars
- 344
- Forks
- 79
- Avg merge
- 2d 23h
- Merged PRs (30d)
- 27
Description
Questionnaire
-
Does ROCm work for you outside of Julia, e.g. C/C++/Python?
yes -
Post output of
rocminfo. (Agent 5 is about the GPU)
ROCk module version 6.10.5 is loaded
=====================
HSA System Attributes
=====================
Runtime Version: 1.14
Runtime Ext Version: 1.6
System Timestamp Freq.: 1000.000000MHz
Sig. Max Wait Duration: 18446744073709551615 (0xFFFFFFFFFFFFFFFF) (timestamp count)
Machine Model: LARGE
System Endianness: LITTLE
Mwaitx: DISABLED
DMAbuf Support: YES
==========
HSA Agents
==========
*******
Agent 1
*******
Name: AMD EPYC 7A53 64-Core Processor
Uuid: CPU-XX
Marketing Name: AMD EPYC 7A53 64-Core Processor
Vendor Name: CPU
Feature: None specified
Profile: FULL_PROFILE
Float Round Mode: NEAR
Max Queue Number: 0(0x0)
Queue Min Size: 0(0x0)
Queue Max Size: 0(0x0)
Queue Type: MULTI
Node: 0
Device Type: CPU
Cache Info:
L1: 32768(0x8000) KB
Chip ID: 0(0x0)
ASIC Revision: 0(0x0)
Cacheline Size: 64(0x40)
Max Clock Freq. (MHz): 2000
BDFID: 0
Internal Node ID: 0
Compute Unit: 32
SIMDs per CU: 0
Shader Engines: 0
Shader Arrs. per Eng.: 0
WatchPts on Addr. Ranges:1
Memory Properties:
Features: None
Pool Info:
Pool 1
Segment: GLOBAL; FLAGS: FINE GRAINED
Size: 131302256(0x7d38370) KB
Allocatable: TRUE
Alloc Granule: 4KB
Alloc Recommended Granule:4KB
Alloc Alignment: 4KB
Accessible by all: TRUE
Pool 2
Segment: GLOBAL; FLAGS: EXTENDED FINE GRAINED
Size: 131302256(0x7d38370) KB
Allocatable: TRUE
Alloc Granule: 4KB
Alloc Recommended Granule:4KB
Alloc Alignment: 4KB
Accessible by all: TRUE
Pool 3
Segment: GLOBAL; FLAGS: KERNARG, FINE GRAINED
Size: 131302256(0x7d38370) KB
Allocatable: TRUE
Alloc Granule: 4KB
Alloc Recommended Granule:4KB
Alloc Alignment: 4KB
Accessible by all: TRUE
Pool 4
Segment: GLOBAL; FLAGS: COARSE GRAINED
Size: 131302256(0x7d38370) KB
Allocatable: TRUE
Alloc Granule: 4KB
Alloc Recommended Granule:4KB
Alloc Alignment: 4KB
Accessible by all: TRUE
ISA Info:
*******
Agent 2
*******
Name: AMD EPYC 7A53 64-Core Processor
Uuid: CPU-XX
Marketing Name: AMD EPYC 7A53 64-Core Processor
Vendor Name: CPU
Feature: None specified
Profile: FULL_PROFILE
Float Round Mode: NEAR
Max Queue Number: 0(0x0)
Queue Min Size: 0(0x0)
Queue Max Size: 0(0x0)
Queue Type: MULTI
Node: 1
Device Type: CPU
Cache Info:
L1: 32768(0x8000) KB
Chip ID: 0(0x0)
ASIC Revision: 0(0x0)
Cacheline Size: 64(0x40)
Max Clock Freq. (MHz): 2000
BDFID: 0
Internal Node ID: 1
Compute Unit: 32
SIMDs per CU: 0
Shader Engines: 0
Shader Arrs. per Eng.: 0
WatchPts on Addr. Ranges:1
Memory Properties:
Features: None
Pool Info:
Pool 1
Segment: GLOBAL; FLAGS: FINE GRAINED
Size: 132110904(0x7dfda38) KB
Allocatable: TRUE
Alloc Granule: 4KB
Alloc Recommended Granule:4KB
Alloc Alignment: 4KB
Accessible by all: TRUE
Pool 2
Segment: GLOBAL; FLAGS: EXTENDED FINE GRAINED
Size: 132110904(0x7dfda38) KB
Allocatable: TRUE
Alloc Granule: 4KB
Alloc Recommended Granule:4KB
Alloc Alignment: 4KB
Accessible by all: TRUE
Pool 3
Segment: GLOBAL; FLAGS: KERNARG, FINE GRAINED
Size: 132110904(0x7dfda38) KB
Allocatable: TRUE
Alloc Granule: 4KB
Alloc Recommended Granule:4KB
Alloc Alignment: 4KB
Accessible by all: TRUE
Pool 4
Segment: GLOBAL; FLAGS: COARSE GRAINED
Size: 132110904(0x7dfda38) KB
Allocatable: TRUE
Alloc Granule: 4KB
Alloc Recommended Granule:4KB
Alloc Alignment: 4KB
Accessible by all: TRUE
ISA Info:
*******
Agent 3
*******
Name: AMD EPYC 7A53 64-Core Processor
Uuid: CPU-XX
Marketing Name: AMD EPYC 7A53 64-Core Processor
Vendor Name: CPU
Feature: None specified
Profile: FULL_PROFILE
Float Round Mode: NEAR
Max Queue Number: 0(0x0)
Queue Min Size: 0(0x0)
Queue Max Size: 0(0x0)
Queue Type: MULTI
Node: 2
Device Type: CPU
Cache Info:
L1: 32768(0x8000) KB
Chip ID: 0(0x0)
ASIC Revision: 0(0x0)
Cacheline Size: 64(0x40)
Max Clock Freq. (MHz): 2000
BDFID: 0
Internal Node ID: 2
Compute Unit: 32
SIMDs per CU: 0
Shader Engines: 0
Shader Arrs. per Eng.: 0
WatchPts on Addr. Ranges:1
Memory Properties:
Features: None
Pool Info:
Pool 1
Segment: GLOBAL; FLAGS: FINE GRAINED
Size: 132110908(0x7dfda3c) KB
Allocatable: TRUE
Alloc Granule: 4KB
Alloc Recommended Granule:4KB
Alloc Alignment: 4KB
Accessible by all: TRUE
Pool 2
Segment: GLOBAL; FLAGS: EXTENDED FINE GRAINED
Size: 132110908(0x7dfda3c) KB
Allocatable: TRUE
Alloc Granule: 4KB
Alloc Recommended Granule:4KB
Alloc Alignment: 4KB
Accessible by all: TRUE
Pool 3
Segment: GLOBAL; FLAGS: KERNARG, FINE GRAINED
Size: 132110908(0x7dfda3c) KB
Allocatable: TRUE
Alloc Granule: 4KB
Alloc Recommended Granule:4KB
Alloc Alignment: 4KB
Accessible by all: TRUE
Pool 4
Segment: GLOBAL; FLAGS: COARSE GRAINED
Size: 132110908(0x7dfda3c) KB
Allocatable: TRUE
Alloc Granule: 4KB
Alloc Recommended Granule:4KB
Alloc Alignment: 4KB
Accessible by all: TRUE
ISA Info:
*******
Agent 4
*******
Name: AMD EPYC 7A53 64-Core Processor
Uuid: CPU-XX
Marketing Name: AMD EPYC 7A53 64-Core Processor
Vendor Name: CPU
Feature: None specified
Profile: FULL_PROFILE
Float Round Mode: NEAR
Max Queue Number: 0(0x0)
Queue Min Size: 0(0x0)
Queue Max Size: 0(0x0)
Queue Type: MULTI
Node: 3
Device Type: CPU
Cache Info:
L1: 32768(0x8000) KB
Chip ID: 0(0x0)
ASIC Revision: 0(0x0)
Cacheline Size: 64(0x40)
Max Clock Freq. (MHz): 2000
BDFID: 0
Internal Node ID: 3
Compute Unit: 32
SIMDs per CU: 0
Shader Engines: 0
Shader Arrs. per Eng.: 0
WatchPts on Addr. Ranges:1
Memory Properties:
Features: None
Pool Info:
Pool 1
Segment: GLOBAL; FLAGS: FINE GRAINED
Size: 132055844(0x7df0324) KB
Allocatable: TRUE
Alloc Granule: 4KB
Alloc Recommended Granule:4KB
Alloc Alignment: 4KB
Accessible by all: TRUE
Pool 2
Segment: GLOBAL; FLAGS: EXTENDED FINE GRAINED
Size: 132055844(0x7df0324) KB
Allocatable: TRUE
Alloc Granule: 4KB
Alloc Recommended Granule:4KB
Alloc Alignment: 4KB
Accessible by all: TRUE
Pool 3
Segment: GLOBAL; FLAGS: KERNARG, FINE GRAINED
Size: 132055844(0x7df0324) KB
Allocatable: TRUE
Alloc Granule: 4KB
Alloc Recommended Granule:4KB
Alloc Alignment: 4KB
Accessible by all: TRUE
Pool 4
Segment: GLOBAL; FLAGS: COARSE GRAINED
Size: 132055844(0x7df0324) KB
Allocatable: TRUE
Alloc Granule: 4KB
Alloc Recommended Granule:4KB
Alloc Alignment: 4KB
Accessible by all: TRUE
ISA Info:
*******
Agent 5
*******
Name: gfx90a
Uuid: GPU-9d4ab50ff31ca655
Marketing Name: AMD Instinct MI250X
Vendor Name: AMD
Feature: KERNEL_DISPATCH
Profile: BASE_PROFILE
Float Round Mode: NEAR
Max Queue Number: 128(0x80)
Queue Min Size: 64(0x40)
Queue Max Size: 131072(0x20000)
Queue Type: MULTI
Node: 4
Device Type: GPU
Cache Info:
L1: 16(0x10) KB
L2: 8192(0x2000) KB
Chip ID: 29704(0x7408)
ASIC Revision: 1(0x1)
Cacheline Size: 128(0x80)
Max Clock Freq. (MHz): 1700
BDFID: 50688
Internal Node ID: 4
Compute Unit: 110
SIMDs per CU: 4
Shader Engines: 8
Shader Arrs. per Eng.: 1
WatchPts on Addr. Ranges:4
Coherent Host Access: TRUE
Memory Properties:
Features: KERNEL_DISPATCH
Fast F16 Operation: TRUE
Wavefront Size: 64(0x40)
Workgroup Max Size: 1024(0x400)
Workgroup Max Size per Dimension:
x 1024(0x400)
y 1024(0x400)
z 1024(0x400)
Max Waves Per CU: 32(0x20)
Max Work-item Per CU: 2048(0x800)
Grid Max Size: 4294967295(0xffffffff)
Grid Max Size per Dimension:
x 4294967295(0xffffffff)
y 4294967295(0xffffffff)
z 4294967295(0xffffffff)
Max fbarriers/Workgrp: 32
Packet Processor uCode:: 92
SDMA engine uCode:: 9
IOMMU Support:: None
Pool Info:
Pool 1
Segment: GLOBAL; FLAGS: COARSE GRAINED
Size: 67092480(0x3ffc000) KB
Allocatable: TRUE
Alloc Granule: 4KB
Alloc Recommended Granule:2048KB
Alloc Alignment: 4KB
Accessible by all: FALSE
Pool 2
Segment: GLOBAL; FLAGS: EXTENDED FINE GRAINED
Size: 67092480(0x3ffc000) KB
Allocatable: TRUE
Alloc Granule: 4KB
Alloc Recommended Granule:2048KB
Alloc Alignment: 4KB
Accessible by all: FALSE
Pool 3
Segment: GLOBAL; FLAGS: FINE GRAINED
Size: 67092480(0x3ffc000) KB
Allocatable: TRUE
Alloc Granule: 4KB
Alloc Recommended Granule:2048KB
Alloc Alignment: 4KB
Accessible by all: FALSE
Pool 4
Segment: GROUP
Size: 64(0x40) KB
Allocatable: FALSE
Alloc Granule: 0KB
Alloc Recommended Granule:0KB
Alloc Alignment: 0KB
Accessible by all: FALSE
ISA Info:
ISA 1
Name: amdgcn-amd-amdhsa--gfx90a:sramecc+:xnack-
Machine Models: HSA_MACHINE_MODEL_LARGE
Profiles: HSA_PROFILE_BASE
Default Rounding Mode: NEAR
Default Rounding Mode: NEAR
Fast f16: TRUE
Workgroup Max Size: 1024(0x400)
Workgroup Max Size per Dimension:
x 1024(0x400)
y 1024(0x400)
z 1024(0x400)
Grid Max Size: 4294967295(0xffffffff)
Grid Max Size per Dimension:
x 4294967295(0xffffffff)
y 4294967295(0xffffffff)
z 4294967295(0xffffffff)
FBarrier Max Size: 32
*** Done ***
- Post output of
AMDGPU.versioninfo()if possible.
┌───────────┬──────────────────┬───────────┬─────────────────────────────────────────────────────────────────────────────────────────────────────────
│ Available │ Name │ Version │ Path ⋯
├───────────┼──────────────────┼───────────┼─────────────────────────────────────────────────────────────────────────────────────────────────────────
│ + │ LLD │ - │ /opt/rocm-6.3.4/lib/llvm/bin/ld.lld ⋯
│ + │ Device Libraries │ - │ /projappl/project_462000008/decristoforo/.julia/artifacts/b46ab46ef568406312e5f500efb677511199c2f9/amd ⋯
│ + │ HIP │ 6.3.42134 │ /opt/rocm-6.3.4/lib/libamdhip64.so ⋯
│ + │ rocBLAS │ 4.3.0 │ /opt/rocm-6.3.4/lib/librocblas.so ⋯
│ + │ rocSOLVER │ 3.27.0 │ /opt/rocm-6.3.4/lib/librocsolver.so ⋯
│ + │ rocSPARSE │ 3.3.0 │ /opt/rocm-6.3.4/lib/librocsparse.so ⋯
│ + │ rocRAND │ 2.10.5 │ /opt/rocm-6.3.4/lib/librocrand.so ⋯
│ + │ rocFFT │ 1.0.31 │ /opt/rocm-6.3.4/lib/librocfft.so ⋯
│ + │ MIOpen │ 3.3.0 │ /opt/rocm-6.3.4/lib/libMIOpen.so ⋯
└───────────┴──────────────────┴───────────┴─────────────────────────────────────────────────────────────────────────────────────────────────────────
1 column omitted
[ Info: AMDGPU devices
┌────┬─────────────────────┬────────────────────────┬───────────┬────────────┬───────────────┐
│ Id │ Name │ GCN arch │ Wavefront │ Memory │ Shared Memory │
├────┼─────────────────────┼────────────────────────┼───────────┼────────────┼───────────────┤
│ 1 │ AMD Instinct MI250X │ gfx90a:sramecc+:xnack- │ 64 │ 63.984 GiB │ 64.000 KiB │
└────┴─────────────────────┴────────────────────────┴───────────┴────────────┴───────────────┘
Reproducing the bug
- Describe what's not working.
AMDGPU.rocFFT's process-global plan cache (AMDGPU.rocFFT.IDLE_HANDLES, a HandleCache in src/cache.jl/src/fft/wrappers.jl) never releases a plan handle unless more than max_entries (32) idle handles share the exact same cache key ((context, fft_type, dims, eltype, inplace, region)). A workload that plans many FFTs of distinct shapes (rebuilding FFT plans for a sweep of different problem sizes in one process) accumulates one leaked rocfft_plan handle per shape forever, since no shape ever repeats 32+ times. This eventually exhausts GPU memory and a later, unrelated rocfft_plan_create call fails with ROCFFTError(rocfft_status_failure).
Note: cache.jl has a standing # TODO: https://github.com/JuliaGPU/AMDGPU.jl/blob/16f5974fe0ee0b1b63f249b6592caf5b335b6485/src/cache.jl#L6
- Provide MWE to reproduce it (if possible).
⚠️ DISCLAIMER: AI was involved when writing this scrit.
using AMDGPU
const N_SIZES = length(ARGS) >= 1 ? parse(Int, ARGS[1]) : 6000
const CHECKPOINT_EVERY = 50
# One real->complex forward plan + one complex->real inverse plan for a given length,
# exercised once and then dropped. This is exactly the access pattern
# SpeedyWeather.jl's SpectralTransform uses: one FFT plan per latitude ring (ring lengths
# differ across a Gaussian/octahedral grid), rebuilt from scratch for every distinct
# resolution in a benchmark sweep.
function run_one!(len::Int)
x = ROCArray(rand(Float32, len))
p_fwd = AMDGPU.rocFFT.plan_rfft(x, (1,))
y = p_fwd * x
p_inv = AMDGPU.rocFFT.plan_brfft(y, len, (1,))
z = p_inv * y
AMDGPU.synchronize()
return nothing
end
function cache_stats()
idle = AMDGPU.rocFFT.IDLE_HANDLES.idle_handles
n_keys = length(idle)
n_idle = isempty(idle) ? 0 : sum(length(v) for v in values(idle))
return n_keys, n_idle
end
function main(n_sizes::Int)
println("AMDGPU version: ", AMDGPU.pkgversion(AMDGPU))
try
AMDGPU.versioninfo()
catch err
println("AMDGPU.versioninfo() failed: ", err)
end
free0, total0 = AMDGPU.info()
println("Initial GPU memory: ", Base.format_bytes(free0), " free / ", Base.format_bytes(total0), " total")
println("Planning $n_sizes distinct FFT lengths (16, 18, 20, ... in steps of 2)...")
n_done = 0
for i in 1:n_sizes
len = 16 + 2 * (i - 1)
try
run_one!(len)
catch err
bt = catch_backtrace()
println()
println("=== REPRODUCED after $i distinct plan sizes (last size = $len) ===")
n_keys, n_idle = cache_stats()
free, total = AMDGPU.info()
println("cached_plan_shapes=$n_keys idle_handles=$n_idle free=$(Base.format_bytes(free))")
showerror(stdout, err, bt)
println()
println("RESULT: FAIL (reproduced) after $i distinct FFT plan sizes")
exit(1)
end
n_done = i
if i % CHECKPOINT_EVERY == 0
GC.gc() # force finalizers (release_plan!) to run
AMDGPU.synchronize()
free, total = AMDGPU.info()
n_keys, n_idle = cache_stats()
println(
"[$i/$n_sizes] free=$(Base.format_bytes(free)) " *
"cached_plan_shapes=$n_keys idle_handles=$n_idle"
)
end
end
free_end, _ = AMDGPU.info()
println()
println(
"RESULT: did not reproduce with n_sizes=$n_sizes (free memory dropped from ",
Base.format_bytes(free0), " to ", Base.format_bytes(free_end),
"). Try a larger n_sizes argument, or run on a GPU with less free memory."
)
return
end
main(N_SIZES)
On LUMI this results in the following error:
julia> main(N_SIZES)
AMDGPU version: 2.8.0
AMDGPU versioninfo
AMDGPU.versioninfo() failed: ErrorException("could not load symbol \"hiptensorGetVersion\":\n/opt/rocm-6.3.4/lib/libhiptensor.so: undefined symbol: hiptensorGetVersion")
Initial GPU memory: 63.896 GiB free / 63.984 GiB total
Planning 6000 distinct FFT lengths (16, 18, 20, ... in steps of 2)...
[50/6000] free=63.168 GiB cached_plan_shapes=100 idle_handles=100
[100/6000] free=62.385 GiB cached_plan_shapes=200 idle_handles=200
[150/6000] free=61.555 GiB cached_plan_shapes=300 idle_handles=300
[200/6000] free=60.664 GiB cached_plan_shapes=400 idle_handles=400
[250/6000] free=59.768 GiB cached_plan_shapes=500 idle_handles=500
[300/6000] free=58.758 GiB cached_plan_shapes=600 idle_handles=600
[350/6000] free=57.760 GiB cached_plan_shapes=700 idle_handles=700
[400/6000] free=56.738 GiB cached_plan_shapes=800 idle_handles=800
[450/6000] free=55.705 GiB cached_plan_shapes=900 idle_handles=900
[500/6000] free=54.672 GiB cached_plan_shapes=1000 idle_handles=1000
[550/6000] free=53.580 GiB cached_plan_shapes=1100 idle_handles=1100
[600/6000] free=52.500 GiB cached_plan_shapes=1200 idle_handles=1200
[650/6000] free=51.396 GiB cached_plan_shapes=1300 idle_handles=1300
[700/6000] free=50.309 GiB cached_plan_shapes=1400 idle_handles=1400
[750/6000] free=49.213 GiB cached_plan_shapes=1500 idle_handles=1500
[800/6000] free=48.094 GiB cached_plan_shapes=1600 idle_handles=1600
[850/6000] free=47.000 GiB cached_plan_shapes=1700 idle_handles=1700
[900/6000] free=45.904 GiB cached_plan_shapes=1800 idle_handles=1800
[950/6000] free=44.785 GiB cached_plan_shapes=1900 idle_handles=1900
[1000/6000] free=43.674 GiB cached_plan_shapes=2000 idle_handles=2000
[1050/6000] free=42.561 GiB cached_plan_shapes=2100 idle_handles=2100
[1100/6000] free=41.449 GiB cached_plan_shapes=2200 idle_handles=2200
[1150/6000] free=40.338 GiB cached_plan_shapes=2300 idle_handles=2300
[1200/6000] free=39.211 GiB cached_plan_shapes=2400 idle_handles=2400
[1250/6000] free=38.092 GiB cached_plan_shapes=2500 idle_handles=2500
[1300/6000] free=36.973 GiB cached_plan_shapes=2600 idle_handles=2600
[1350/6000] free=35.861 GiB cached_plan_shapes=2700 idle_handles=2700
[1400/6000] free=34.750 GiB cached_plan_shapes=2800 idle_handles=2800
[1450/6000] free=33.631 GiB cached_plan_shapes=2900 idle_handles=2900
[1500/6000] free=32.496 GiB cached_plan_shapes=3000 idle_handles=3000
[1550/6000] free=31.367 GiB cached_plan_shapes=3100 idle_handles=3100
[1600/6000] free=30.240 GiB cached_plan_shapes=3200 idle_handles=3200
[1650/6000] free=29.113 GiB cached_plan_shapes=3300 idle_handles=3300
[1700/6000] free=28.002 GiB cached_plan_shapes=3400 idle_handles=3400
[1750/6000] free=26.875 GiB cached_plan_shapes=3500 idle_handles=3500
[1800/6000] free=25.756 GiB cached_plan_shapes=3600 idle_handles=3600
[1850/6000] free=24.613 GiB cached_plan_shapes=3700 idle_handles=3700
[1900/6000] free=23.486 GiB cached_plan_shapes=3800 idle_handles=3800
[1950/6000] free=22.359 GiB cached_plan_shapes=3900 idle_handles=3900
[2000/6000] free=21.225 GiB cached_plan_shapes=4000 idle_handles=4000
[2050/6000] free=19.939 GiB cached_plan_shapes=4100 idle_handles=4100
[2100/6000] free=17.883 GiB cached_plan_shapes=4200 idle_handles=4200
[2150/6000] free=15.824 GiB cached_plan_shapes=4300 idle_handles=4300
[2200/6000] free=13.848 GiB cached_plan_shapes=4400 idle_handles=4400
[2250/6000] free=11.781 GiB cached_plan_shapes=4500 idle_handles=4500
[2300/6000] free=9.734 GiB cached_plan_shapes=4600 idle_handles=4600
[2350/6000] free=7.656 GiB cached_plan_shapes=4700 idle_handles=4700
[2400/6000] free=5.615 GiB cached_plan_shapes=4800 idle_handles=4800
[2450/6000] free=3.584 GiB cached_plan_shapes=4900 idle_handles=4900
[2500/6000] free=1.533 GiB cached_plan_shapes=5000 idle_handles=5000
=== REPRODUCED after 2524 distinct plan sizes (last size = 5062) ===
cached_plan_shapes=5033 idle_handles=5033 free=576.000 MiB
ROCFFTError(code rocfft_status_failure, an operation failed)
Stacktrace:
[1] check
@ /projappl/project_462000008/decristoforo/.julia/packages/AMDGPU/HUKS4/src/fft/error.jl:38 [inlined]
[2] macro expansion
@ /projappl/project_462000008/decristoforo/.julia/packages/AMDGPU/HUKS4/src/utils.jl:239 [inlined]
[3] rocfft_plan_create(plan::Base.RefValue{Ptr{AMDGPU.rocFFT.rocfft_plan_t}}, placement::AMDGPU.rocFFT.rocfft_result_placement_e, transform_type::AMDGPU.rocFFT.rocfft_transform_type_e, precision::AMDGPU.rocFFT.rocfft_precision_e, dimensions::Int64, lengths::Vector{Int64}, number_of_transforms::Int64, description::Ptr{Nothing})
@ AMDGPU.rocFFT /projappl/project_462000008/decristoforo/.julia/packages/AMDGPU/HUKS4/src/fft/librocfft.jl:92
[4] create_plan(xtype::AMDGPU.rocFFT.rocfft_transform_type_e, xdims::Tuple{Int64}, T::Type, inplace::Bool, region::Tuple{Int64})
@ AMDGPU.rocFFT /projappl/project_462000008/decristoforo/.julia/packages/AMDGPU/HUKS4/src/fft/wrappers.jl:40
[5] #get_plan##0
@ /projappl/project_462000008/decristoforo/.julia/packages/AMDGPU/HUKS4/src/fft/wrappers.jl:11 [inlined]
[6] pop!(f::AMDGPU.rocFFT.var"#get_plan##0#get_plan##1"{AMDGPU.rocFFT.rocfft_transform_type_e, Tuple{Int64}, Type{Float32}, Bool, Tuple{Int64}}, cache::HandleCache{Tuple{HIPContext, AMDGPU.rocFFT.rocfft_transform_type_e, NTuple{N, Int64} where N, Type, Bool, Any}, Tuple{Ptr{AMDGPU.rocFFT.rocfft_plan_t}, Int64}}, key::Tuple{HIPContext, AMDGPU.rocFFT.rocfft_transform_type_e, Tuple{Int64}, DataType, Bool, Tuple{Int64}})
@ AMDGPU /projappl/project_462000008/decristoforo/.julia/packages/AMDGPU/HUKS4/src/cache.jl:49
[7] get_plan
@ /projappl/project_462000008/decristoforo/.julia/packages/AMDGPU/HUKS4/src/fft/wrappers.jl:11 [inlined]
[8] plan_rfft(X::ROCArray{Float32, 1, AMDGPU.Runtime.Mem.HIPBuffer}, region::Tuple{Int64})
@ AMDGPU.rocFFT /projappl/project_462000008/decristoforo/.julia/packages/AMDGPU/HUKS4/src/fft/fft.jl:149
[9] run_one!(len::Int64)
@ Main ./REPL[4]:8
[10] main(n_sizes::Int64)
@ Main ./REPL[6]:17
[11] top-level scope
@ REPL[7]:1
[12] __repl_entry_eval_expanded_with_loc(mod::Module, ast::Any, toplevel_file::Ref{Ptr{UInt8}}, toplevel_line::Ref{Int32})
@ REPL /pfs/lustrep4/users/decristoforo/.julia/juliaup/julia-1.12.7+0.x64.linux.gnu/share/julia/stdlib/v1.12/REPL/src/REPL.jl:301
[13] toplevel_eval_with_hooks(mod::Module, ast::Any, toplevel_file::Any, toplevel_line::Any)
@ REPL /pfs/lustrep4/users/decristoforo/.julia/juliaup/julia-1.12.7+0.x64.linux.gnu/share/julia/stdlib/v1.12/REPL/src/REPL.jl:308
[14] toplevel_eval_with_hooks(mod::Module, ast::Any, toplevel_file::Any, toplevel_line::Any)
@ REPL /pfs/lustrep4/users/decristoforo/.julia/juliaup/julia-1.12.7+0.x64.linux.gnu/share/julia/stdlib/v1.12/REPL/src/REPL.jl:312
[15] toplevel_eval_with_hooks
@ /pfs/lustrep4/users/decristoforo/.julia/juliaup/julia-1.12.7+0.x64.linux.gnu/share/julia/stdlib/v1.12/REPL/src/REPL.jl:305 [inlined]
[16] eval_user_input(ast::Any, backend::REPL.REPLBackend, mod::Module)
@ REPL /pfs/lustrep4/users/decristoforo/.julia/juliaup/julia-1.12.7+0.x64.linux.gnu/share/julia/stdlib/v1.12/REPL/src/REPL.jl:330
[17] repl_backend_loop(backend::REPL.REPLBackend, get_module::Function)
@ REPL /pfs/lustrep4/users/decristoforo/.julia/juliaup/julia-1.12.7+0.x64.linux.gnu/share/julia/stdlib/v1.12/REPL/src/REPL.jl:452
[18] start_repl_backend(backend::REPL.REPLBackend, consumer::Any; get_module::Function)
@ REPL /pfs/lustrep4/users/decristoforo/.julia/juliaup/julia-1.12.7+0.x64.linux.gnu/share/julia/stdlib/v1.12/REPL/src/REPL.jl:427
[19] start_repl_backend
@ /pfs/lustrep4/users/decristoforo/.julia/juliaup/julia-1.12.7+0.x64.linux.gnu/share/julia/stdlib/v1.12/REPL/src/REPL.jl:424 [inlined]
[20] run_repl(repl::REPL.AbstractREPL, consumer::Any; backend_on_current_task::Bool, backend::Any)
@ REPL /pfs/lustrep4/users/decristoforo/.julia/juliaup/julia-1.12.7+0.x64.linux.gnu/share/julia/stdlib/v1.12/REPL/src/REPL.jl:653
[21] run_repl(repl::REPL.AbstractREPL, consumer::Any)
@ REPL /pfs/lustrep4/users/decristoforo/.julia/juliaup/julia-1.12.7+0.x64.linux.gnu/share/julia/stdlib/v1.12/REPL/src/REPL.jl:639
[22] run_std_repl(REPL::Module, quiet::Bool, banner::Symbol, history_file::Bool)
@ Base ./client.jl:478
[23] run_main_repl(interactive::Bool, quiet::Bool, banner::Symbol, history_file::Bool)
@ Base ./client.jl:499
[24] repl_main
@ ./client.jl:586 [inlined]
[25] _start()
@ Base ./client.jl:561
RESULT: FAIL (reproduced) after 2524 distinct FFT plan sizes
Contributor guide
No contributing guide indexed for this repository
First steps
- Read the whole issue, then the project's contributing guide.
- Comment on the issue to say you are picking it up — it saves two people doing the same work.
- Fork the repository and make your change on a branch.
- Open a pull request that references the issue number.
Research direction
Start by locating the IDLE_HANDLES cache in the rocFFT-related code and inspect how plans for distinct shapes are retained. Exercise the cache with distinct-shape plans and verify that completed plans are evicted and GPU memory no longer grows without bound.
Written by the indexing model from the issue text.
Assessment
- Tech stack
- julia
- Domain
- hpc
- Issue type
- Bug
- Difficulty
- 4/5
- Estimated time
- 3-5 days
- Activity status
- Active
- Clarity
- Needs clarification
- Newbie friendliness
- 45/100