JuliaGPU / JuliaGPU/AMDGPU.jl

`rocFFT`: plan handle cache (`IDLE_HANDLES`) never evicts distinct-shape plans, leaking GPU memory

Closed
#1,053 2 comments 1 reaction 0 assignees View on GitHub

Nobody has claimed this yet.

bug
Dominant language
Julia
Stars
344
Forks
79
Avg merge
2d 23h
Merged PRs (30d)
27

Description

Questionnaire

  1. Does ROCm work for you outside of Julia, e.g. C/C++/Python?
    yes

  2. Post output of rocminfo. (Agent 5 is about the GPU)

ROCk module version 6.10.5 is loaded
=====================    
HSA System Attributes    
=====================    
Runtime Version:         1.14
Runtime Ext Version:     1.6
System Timestamp Freq.:  1000.000000MHz
Sig. Max Wait Duration:  18446744073709551615 (0xFFFFFFFFFFFFFFFF) (timestamp count)
Machine Model:           LARGE                              
System Endianness:       LITTLE                             
Mwaitx:                  DISABLED
DMAbuf Support:          YES

==========               
HSA Agents               
==========               
*******                  
Agent 1                  
*******                  
  Name:                    AMD EPYC 7A53 64-Core Processor    
  Uuid:                    CPU-XX                             
  Marketing Name:          AMD EPYC 7A53 64-Core Processor    
  Vendor Name:             CPU                                
  Feature:                 None specified                     
  Profile:                 FULL_PROFILE                       
  Float Round Mode:        NEAR                               
  Max Queue Number:        0(0x0)                             
  Queue Min Size:          0(0x0)                             
  Queue Max Size:          0(0x0)                             
  Queue Type:              MULTI                              
  Node:                    0                                  
  Device Type:             CPU                                
  Cache Info:              
    L1:                      32768(0x8000) KB                   
  Chip ID:                 0(0x0)                             
  ASIC Revision:           0(0x0)                             
  Cacheline Size:          64(0x40)                           
  Max Clock Freq. (MHz):   2000                               
  BDFID:                   0                                  
  Internal Node ID:        0                                  
  Compute Unit:            32                                 
  SIMDs per CU:            0                                  
  Shader Engines:          0                                  
  Shader Arrs. per Eng.:   0                                  
  WatchPts on Addr. Ranges:1                                  
  Memory Properties:       
  Features:                None
  Pool Info:               
    Pool 1                   
      Segment:                 GLOBAL; FLAGS: FINE GRAINED        
      Size:                    131302256(0x7d38370) KB            
      Allocatable:             TRUE                               
      Alloc Granule:           4KB                                
      Alloc Recommended Granule:4KB                                
      Alloc Alignment:         4KB                                
      Accessible by all:       TRUE                               
    Pool 2                   
      Segment:                 GLOBAL; FLAGS: EXTENDED FINE GRAINED
      Size:                    131302256(0x7d38370) KB            
      Allocatable:             TRUE                               
      Alloc Granule:           4KB                                
      Alloc Recommended Granule:4KB                                
      Alloc Alignment:         4KB                                
      Accessible by all:       TRUE                               
    Pool 3                   
      Segment:                 GLOBAL; FLAGS: KERNARG, FINE GRAINED
      Size:                    131302256(0x7d38370) KB            
      Allocatable:             TRUE                               
      Alloc Granule:           4KB                                
      Alloc Recommended Granule:4KB                                
      Alloc Alignment:         4KB                                
      Accessible by all:       TRUE                               
    Pool 4                   
      Segment:                 GLOBAL; FLAGS: COARSE GRAINED      
      Size:                    131302256(0x7d38370) KB            
      Allocatable:             TRUE                               
      Alloc Granule:           4KB                                
      Alloc Recommended Granule:4KB                                
      Alloc Alignment:         4KB                                
      Accessible by all:       TRUE                               
  ISA Info:                
*******                  
Agent 2                  
*******                  
  Name:                    AMD EPYC 7A53 64-Core Processor    
  Uuid:                    CPU-XX                             
  Marketing Name:          AMD EPYC 7A53 64-Core Processor    
  Vendor Name:             CPU                                
  Feature:                 None specified                     
  Profile:                 FULL_PROFILE                       
  Float Round Mode:        NEAR                               
  Max Queue Number:        0(0x0)                             
  Queue Min Size:          0(0x0)                             
  Queue Max Size:          0(0x0)                             
  Queue Type:              MULTI                              
  Node:                    1                                  
  Device Type:             CPU                                
  Cache Info:              
    L1:                      32768(0x8000) KB                   
  Chip ID:                 0(0x0)                             
  ASIC Revision:           0(0x0)                             
  Cacheline Size:          64(0x40)                           
  Max Clock Freq. (MHz):   2000                               
  BDFID:                   0                                  
  Internal Node ID:        1                                  
  Compute Unit:            32                                 
  SIMDs per CU:            0                                  
  Shader Engines:          0                                  
  Shader Arrs. per Eng.:   0                                  
  WatchPts on Addr. Ranges:1                                  
  Memory Properties:       
  Features:                None
  Pool Info:               
    Pool 1                   
      Segment:                 GLOBAL; FLAGS: FINE GRAINED        
      Size:                    132110904(0x7dfda38) KB            
      Allocatable:             TRUE                               
      Alloc Granule:           4KB                                
      Alloc Recommended Granule:4KB                                
      Alloc Alignment:         4KB                                
      Accessible by all:       TRUE                               
    Pool 2                   
      Segment:                 GLOBAL; FLAGS: EXTENDED FINE GRAINED
      Size:                    132110904(0x7dfda38) KB            
      Allocatable:             TRUE                               
      Alloc Granule:           4KB                                
      Alloc Recommended Granule:4KB                                
      Alloc Alignment:         4KB                                
      Accessible by all:       TRUE                               
    Pool 3                   
      Segment:                 GLOBAL; FLAGS: KERNARG, FINE GRAINED
      Size:                    132110904(0x7dfda38) KB            
      Allocatable:             TRUE                               
      Alloc Granule:           4KB                                
      Alloc Recommended Granule:4KB                                
      Alloc Alignment:         4KB                                
      Accessible by all:       TRUE                               
    Pool 4                   
      Segment:                 GLOBAL; FLAGS: COARSE GRAINED      
      Size:                    132110904(0x7dfda38) KB            
      Allocatable:             TRUE                               
      Alloc Granule:           4KB                                
      Alloc Recommended Granule:4KB                                
      Alloc Alignment:         4KB                                
      Accessible by all:       TRUE                               
  ISA Info:                
*******                  
Agent 3                  
*******                  
  Name:                    AMD EPYC 7A53 64-Core Processor    
  Uuid:                    CPU-XX                             
  Marketing Name:          AMD EPYC 7A53 64-Core Processor    
  Vendor Name:             CPU                                
  Feature:                 None specified                     
  Profile:                 FULL_PROFILE                       
  Float Round Mode:        NEAR                               
  Max Queue Number:        0(0x0)                             
  Queue Min Size:          0(0x0)                             
  Queue Max Size:          0(0x0)                             
  Queue Type:              MULTI                              
  Node:                    2                                  
  Device Type:             CPU                                
  Cache Info:              
    L1:                      32768(0x8000) KB                   
  Chip ID:                 0(0x0)                             
  ASIC Revision:           0(0x0)                             
  Cacheline Size:          64(0x40)                           
  Max Clock Freq. (MHz):   2000                               
  BDFID:                   0                                  
  Internal Node ID:        2                                  
  Compute Unit:            32                                 
  SIMDs per CU:            0                                  
  Shader Engines:          0                                  
  Shader Arrs. per Eng.:   0                                  
  WatchPts on Addr. Ranges:1                                  
  Memory Properties:       
  Features:                None
  Pool Info:               
    Pool 1                   
      Segment:                 GLOBAL; FLAGS: FINE GRAINED        
      Size:                    132110908(0x7dfda3c) KB            
      Allocatable:             TRUE                               
      Alloc Granule:           4KB                                
      Alloc Recommended Granule:4KB                                
      Alloc Alignment:         4KB                                
      Accessible by all:       TRUE                               
    Pool 2                   
      Segment:                 GLOBAL; FLAGS: EXTENDED FINE GRAINED
      Size:                    132110908(0x7dfda3c) KB            
      Allocatable:             TRUE                               
      Alloc Granule:           4KB                                
      Alloc Recommended Granule:4KB                                
      Alloc Alignment:         4KB                                
      Accessible by all:       TRUE                               
    Pool 3                   
      Segment:                 GLOBAL; FLAGS: KERNARG, FINE GRAINED
      Size:                    132110908(0x7dfda3c) KB            
      Allocatable:             TRUE                               
      Alloc Granule:           4KB                                
      Alloc Recommended Granule:4KB                                
      Alloc Alignment:         4KB                                
      Accessible by all:       TRUE                               
    Pool 4                   
      Segment:                 GLOBAL; FLAGS: COARSE GRAINED      
      Size:                    132110908(0x7dfda3c) KB            
      Allocatable:             TRUE                               
      Alloc Granule:           4KB                                
      Alloc Recommended Granule:4KB                                
      Alloc Alignment:         4KB                                
      Accessible by all:       TRUE                               
  ISA Info:                
*******                  
Agent 4                  
*******                  
  Name:                    AMD EPYC 7A53 64-Core Processor    
  Uuid:                    CPU-XX                             
  Marketing Name:          AMD EPYC 7A53 64-Core Processor    
  Vendor Name:             CPU                                
  Feature:                 None specified                     
  Profile:                 FULL_PROFILE                       
  Float Round Mode:        NEAR                               
  Max Queue Number:        0(0x0)                             
  Queue Min Size:          0(0x0)                             
  Queue Max Size:          0(0x0)                             
  Queue Type:              MULTI                              
  Node:                    3                                  
  Device Type:             CPU                                
  Cache Info:              
    L1:                      32768(0x8000) KB                   
  Chip ID:                 0(0x0)                             
  ASIC Revision:           0(0x0)                             
  Cacheline Size:          64(0x40)                           
  Max Clock Freq. (MHz):   2000                               
  BDFID:                   0                                  
  Internal Node ID:        3                                  
  Compute Unit:            32                                 
  SIMDs per CU:            0                                  
  Shader Engines:          0                                  
  Shader Arrs. per Eng.:   0                                  
  WatchPts on Addr. Ranges:1                                  
  Memory Properties:       
  Features:                None
  Pool Info:               
    Pool 1                   
      Segment:                 GLOBAL; FLAGS: FINE GRAINED        
      Size:                    132055844(0x7df0324) KB            
      Allocatable:             TRUE                               
      Alloc Granule:           4KB                                
      Alloc Recommended Granule:4KB                                
      Alloc Alignment:         4KB                                
      Accessible by all:       TRUE                               
    Pool 2                   
      Segment:                 GLOBAL; FLAGS: EXTENDED FINE GRAINED
      Size:                    132055844(0x7df0324) KB            
      Allocatable:             TRUE                               
      Alloc Granule:           4KB                                
      Alloc Recommended Granule:4KB                                
      Alloc Alignment:         4KB                                
      Accessible by all:       TRUE                               
    Pool 3                   
      Segment:                 GLOBAL; FLAGS: KERNARG, FINE GRAINED
      Size:                    132055844(0x7df0324) KB            
      Allocatable:             TRUE                               
      Alloc Granule:           4KB                                
      Alloc Recommended Granule:4KB                                
      Alloc Alignment:         4KB                                
      Accessible by all:       TRUE                               
    Pool 4                   
      Segment:                 GLOBAL; FLAGS: COARSE GRAINED      
      Size:                    132055844(0x7df0324) KB            
      Allocatable:             TRUE                               
      Alloc Granule:           4KB                                
      Alloc Recommended Granule:4KB                                
      Alloc Alignment:         4KB                                
      Accessible by all:       TRUE                               
  ISA Info:                
*******                  
Agent 5                  
*******                  
  Name:                    gfx90a                             
  Uuid:                    GPU-9d4ab50ff31ca655               
  Marketing Name:          AMD Instinct MI250X                
  Vendor Name:             AMD                                
  Feature:                 KERNEL_DISPATCH                    
  Profile:                 BASE_PROFILE                       
  Float Round Mode:        NEAR                               
  Max Queue Number:        128(0x80)                          
  Queue Min Size:          64(0x40)                           
  Queue Max Size:          131072(0x20000)                    
  Queue Type:              MULTI                              
  Node:                    4                                  
  Device Type:             GPU                                
  Cache Info:              
    L1:                      16(0x10) KB                        
    L2:                      8192(0x2000) KB                    
  Chip ID:                 29704(0x7408)                      
  ASIC Revision:           1(0x1)                             
  Cacheline Size:          128(0x80)                          
  Max Clock Freq. (MHz):   1700                               
  BDFID:                   50688                              
  Internal Node ID:        4                                  
  Compute Unit:            110                                
  SIMDs per CU:            4                                  
  Shader Engines:          8                                  
  Shader Arrs. per Eng.:   1                                  
  WatchPts on Addr. Ranges:4                                  
  Coherent Host Access:    TRUE                               
  Memory Properties:       
  Features:                KERNEL_DISPATCH 
  Fast F16 Operation:      TRUE                               
  Wavefront Size:          64(0x40)                           
  Workgroup Max Size:      1024(0x400)                        
  Workgroup Max Size per Dimension:
    x                        1024(0x400)                        
    y                        1024(0x400)                        
    z                        1024(0x400)                        
  Max Waves Per CU:        32(0x20)                           
  Max Work-item Per CU:    2048(0x800)                        
  Grid Max Size:           4294967295(0xffffffff)             
  Grid Max Size per Dimension:
    x                        4294967295(0xffffffff)             
    y                        4294967295(0xffffffff)             
    z                        4294967295(0xffffffff)             
  Max fbarriers/Workgrp:   32                                 
  Packet Processor uCode:: 92                                 
  SDMA engine uCode::      9                                  
  IOMMU Support::          None                               
  Pool Info:               
    Pool 1                   
      Segment:                 GLOBAL; FLAGS: COARSE GRAINED      
      Size:                    67092480(0x3ffc000) KB             
      Allocatable:             TRUE                               
      Alloc Granule:           4KB                                
      Alloc Recommended Granule:2048KB                             
      Alloc Alignment:         4KB                                
      Accessible by all:       FALSE                              
    Pool 2                   
      Segment:                 GLOBAL; FLAGS: EXTENDED FINE GRAINED
      Size:                    67092480(0x3ffc000) KB             
      Allocatable:             TRUE                               
      Alloc Granule:           4KB                                
      Alloc Recommended Granule:2048KB                             
      Alloc Alignment:         4KB                                
      Accessible by all:       FALSE                              
    Pool 3                   
      Segment:                 GLOBAL; FLAGS: FINE GRAINED        
      Size:                    67092480(0x3ffc000) KB             
      Allocatable:             TRUE                               
      Alloc Granule:           4KB                                
      Alloc Recommended Granule:2048KB                             
      Alloc Alignment:         4KB                                
      Accessible by all:       FALSE                              
    Pool 4                   
      Segment:                 GROUP                              
      Size:                    64(0x40) KB                        
      Allocatable:             FALSE                              
      Alloc Granule:           0KB                                
      Alloc Recommended Granule:0KB                                
      Alloc Alignment:         0KB                                
      Accessible by all:       FALSE                              
  ISA Info:                
    ISA 1                    
      Name:                    amdgcn-amd-amdhsa--gfx90a:sramecc+:xnack-
      Machine Models:          HSA_MACHINE_MODEL_LARGE            
      Profiles:                HSA_PROFILE_BASE                   
      Default Rounding Mode:   NEAR                               
      Default Rounding Mode:   NEAR                               
      Fast f16:                TRUE                               
      Workgroup Max Size:      1024(0x400)                        
      Workgroup Max Size per Dimension:
        x                        1024(0x400)                        
        y                        1024(0x400)                        
        z                        1024(0x400)                        
      Grid Max Size:           4294967295(0xffffffff)             
      Grid Max Size per Dimension:
        x                        4294967295(0xffffffff)             
        y                        4294967295(0xffffffff)             
        z                        4294967295(0xffffffff)             
      FBarrier Max Size:       32                                 
*** Done ***             

  1. Post output of AMDGPU.versioninfo() if possible.
┌───────────┬──────────────────┬───────────┬─────────────────────────────────────────────────────────────────────────────────────────────────────────
│ Available │ Name             │ Version   │ Path                                                                                                   ⋯
├───────────┼──────────────────┼───────────┼─────────────────────────────────────────────────────────────────────────────────────────────────────────
│     +     │ LLD              │ -         │ /opt/rocm-6.3.4/lib/llvm/bin/ld.lld                                                                    ⋯
│     +     │ Device Libraries │ -         │ /projappl/project_462000008/decristoforo/.julia/artifacts/b46ab46ef568406312e5f500efb677511199c2f9/amd ⋯
│     +     │ HIP              │ 6.3.42134 │ /opt/rocm-6.3.4/lib/libamdhip64.so                                                                     ⋯
│     +     │ rocBLAS          │ 4.3.0     │ /opt/rocm-6.3.4/lib/librocblas.so                                                                      ⋯
│     +     │ rocSOLVER        │ 3.27.0    │ /opt/rocm-6.3.4/lib/librocsolver.so                                                                    ⋯
│     +     │ rocSPARSE        │ 3.3.0     │ /opt/rocm-6.3.4/lib/librocsparse.so                                                                    ⋯
│     +     │ rocRAND          │ 2.10.5    │ /opt/rocm-6.3.4/lib/librocrand.so                                                                      ⋯
│     +     │ rocFFT           │ 1.0.31    │ /opt/rocm-6.3.4/lib/librocfft.so                                                                       ⋯
│     +     │ MIOpen           │ 3.3.0     │ /opt/rocm-6.3.4/lib/libMIOpen.so                                                                       ⋯
└───────────┴──────────────────┴───────────┴─────────────────────────────────────────────────────────────────────────────────────────────────────────
                                                                                                                                     1 column omitted

[ Info: AMDGPU devices
┌────┬─────────────────────┬────────────────────────┬───────────┬────────────┬───────────────┐
│ Id │                Name │               GCN arch │ Wavefront │     Memory │ Shared Memory │
├────┼─────────────────────┼────────────────────────┼───────────┼────────────┼───────────────┤
│  1 │ AMD Instinct MI250X │ gfx90a:sramecc+:xnack- │        64 │ 63.984 GiB │    64.000 KiB │
└────┴─────────────────────┴────────────────────────┴───────────┴────────────┴───────────────┘


Reproducing the bug

  1. Describe what's not working.

AMDGPU.rocFFT's process-global plan cache (AMDGPU.rocFFT.IDLE_HANDLES, a HandleCache in src/cache.jl/src/fft/wrappers.jl) never releases a plan handle unless more than max_entries (32) idle handles share the exact same cache key ((context, fft_type, dims, eltype, inplace, region)). A workload that plans many FFTs of distinct shapes (rebuilding FFT plans for a sweep of different problem sizes in one process) accumulates one leaked rocfft_plan handle per shape forever, since no shape ever repeats 32+ times. This eventually exhausts GPU memory and a later, unrelated rocfft_plan_create call fails with ROCFFTError(rocfft_status_failure).

Note: cache.jl has a standing # TODO: https://github.com/JuliaGPU/AMDGPU.jl/blob/16f5974fe0ee0b1b63f249b6592caf5b335b6485/src/cache.jl#L6

  1. Provide MWE to reproduce it (if possible).
    ⚠️ DISCLAIMER: AI was involved when writing this scrit.
using AMDGPU

const N_SIZES = length(ARGS) >= 1 ? parse(Int, ARGS[1]) : 6000
const CHECKPOINT_EVERY = 50

# One real->complex forward plan + one complex->real inverse plan for a given length,
# exercised once and then dropped. This is exactly the access pattern
# SpeedyWeather.jl's SpectralTransform uses: one FFT plan per latitude ring (ring lengths
# differ across a Gaussian/octahedral grid), rebuilt from scratch for every distinct
# resolution in a benchmark sweep.
function run_one!(len::Int)
    x = ROCArray(rand(Float32, len))
    p_fwd = AMDGPU.rocFFT.plan_rfft(x, (1,))
    y = p_fwd * x
    p_inv = AMDGPU.rocFFT.plan_brfft(y, len, (1,))
    z = p_inv * y
    AMDGPU.synchronize()
    return nothing
end

function cache_stats()
    idle = AMDGPU.rocFFT.IDLE_HANDLES.idle_handles
    n_keys = length(idle)
    n_idle = isempty(idle) ? 0 : sum(length(v) for v in values(idle))
    return n_keys, n_idle
end

function main(n_sizes::Int)
    println("AMDGPU version: ", AMDGPU.pkgversion(AMDGPU))
    try
        AMDGPU.versioninfo()
    catch err
        println("AMDGPU.versioninfo() failed: ", err)
    end

    free0, total0 = AMDGPU.info()
    println("Initial GPU memory: ", Base.format_bytes(free0), " free / ", Base.format_bytes(total0), " total")
    println("Planning $n_sizes distinct FFT lengths (16, 18, 20, ... in steps of 2)...")

    n_done = 0
    for i in 1:n_sizes
        len = 16 + 2 * (i - 1)
        try
            run_one!(len)
        catch err
            bt = catch_backtrace()
            println()
            println("=== REPRODUCED after $i distinct plan sizes (last size = $len) ===")
            n_keys, n_idle = cache_stats()
            free, total = AMDGPU.info()
            println("cached_plan_shapes=$n_keys idle_handles=$n_idle free=$(Base.format_bytes(free))")
            showerror(stdout, err, bt)
            println()
            println("RESULT: FAIL (reproduced) after $i distinct FFT plan sizes")
            exit(1)
        end
        n_done = i

        if i % CHECKPOINT_EVERY == 0
            GC.gc()               # force finalizers (release_plan!) to run
            AMDGPU.synchronize()
            free, total = AMDGPU.info()
            n_keys, n_idle = cache_stats()
            println(
                "[$i/$n_sizes] free=$(Base.format_bytes(free)) " *
                    "cached_plan_shapes=$n_keys idle_handles=$n_idle"
            )
        end
    end

    free_end, _ = AMDGPU.info()
    println()
    println(
        "RESULT: did not reproduce with n_sizes=$n_sizes (free memory dropped from ",
        Base.format_bytes(free0), " to ", Base.format_bytes(free_end),
        "). Try a larger n_sizes argument, or run on a GPU with less free memory."
    )
    return
end

main(N_SIZES)

On LUMI this results in the following error:

julia> main(N_SIZES)
AMDGPU version: 2.8.0
AMDGPU versioninfo
AMDGPU.versioninfo() failed: ErrorException("could not load symbol \"hiptensorGetVersion\":\n/opt/rocm-6.3.4/lib/libhiptensor.so: undefined symbol: hiptensorGetVersion")
Initial GPU memory: 63.896 GiB free / 63.984 GiB total
Planning 6000 distinct FFT lengths (16, 18, 20, ... in steps of 2)...
[50/6000] free=63.168 GiB cached_plan_shapes=100 idle_handles=100
[100/6000] free=62.385 GiB cached_plan_shapes=200 idle_handles=200
[150/6000] free=61.555 GiB cached_plan_shapes=300 idle_handles=300
[200/6000] free=60.664 GiB cached_plan_shapes=400 idle_handles=400
[250/6000] free=59.768 GiB cached_plan_shapes=500 idle_handles=500
[300/6000] free=58.758 GiB cached_plan_shapes=600 idle_handles=600
[350/6000] free=57.760 GiB cached_plan_shapes=700 idle_handles=700
[400/6000] free=56.738 GiB cached_plan_shapes=800 idle_handles=800
[450/6000] free=55.705 GiB cached_plan_shapes=900 idle_handles=900
[500/6000] free=54.672 GiB cached_plan_shapes=1000 idle_handles=1000
[550/6000] free=53.580 GiB cached_plan_shapes=1100 idle_handles=1100
[600/6000] free=52.500 GiB cached_plan_shapes=1200 idle_handles=1200
[650/6000] free=51.396 GiB cached_plan_shapes=1300 idle_handles=1300
[700/6000] free=50.309 GiB cached_plan_shapes=1400 idle_handles=1400
[750/6000] free=49.213 GiB cached_plan_shapes=1500 idle_handles=1500
[800/6000] free=48.094 GiB cached_plan_shapes=1600 idle_handles=1600
[850/6000] free=47.000 GiB cached_plan_shapes=1700 idle_handles=1700
[900/6000] free=45.904 GiB cached_plan_shapes=1800 idle_handles=1800
[950/6000] free=44.785 GiB cached_plan_shapes=1900 idle_handles=1900
[1000/6000] free=43.674 GiB cached_plan_shapes=2000 idle_handles=2000
[1050/6000] free=42.561 GiB cached_plan_shapes=2100 idle_handles=2100
[1100/6000] free=41.449 GiB cached_plan_shapes=2200 idle_handles=2200
[1150/6000] free=40.338 GiB cached_plan_shapes=2300 idle_handles=2300
[1200/6000] free=39.211 GiB cached_plan_shapes=2400 idle_handles=2400
[1250/6000] free=38.092 GiB cached_plan_shapes=2500 idle_handles=2500
[1300/6000] free=36.973 GiB cached_plan_shapes=2600 idle_handles=2600
[1350/6000] free=35.861 GiB cached_plan_shapes=2700 idle_handles=2700
[1400/6000] free=34.750 GiB cached_plan_shapes=2800 idle_handles=2800
[1450/6000] free=33.631 GiB cached_plan_shapes=2900 idle_handles=2900
[1500/6000] free=32.496 GiB cached_plan_shapes=3000 idle_handles=3000
[1550/6000] free=31.367 GiB cached_plan_shapes=3100 idle_handles=3100
[1600/6000] free=30.240 GiB cached_plan_shapes=3200 idle_handles=3200
[1650/6000] free=29.113 GiB cached_plan_shapes=3300 idle_handles=3300
[1700/6000] free=28.002 GiB cached_plan_shapes=3400 idle_handles=3400
[1750/6000] free=26.875 GiB cached_plan_shapes=3500 idle_handles=3500
[1800/6000] free=25.756 GiB cached_plan_shapes=3600 idle_handles=3600
[1850/6000] free=24.613 GiB cached_plan_shapes=3700 idle_handles=3700
[1900/6000] free=23.486 GiB cached_plan_shapes=3800 idle_handles=3800
[1950/6000] free=22.359 GiB cached_plan_shapes=3900 idle_handles=3900
[2000/6000] free=21.225 GiB cached_plan_shapes=4000 idle_handles=4000
[2050/6000] free=19.939 GiB cached_plan_shapes=4100 idle_handles=4100
[2100/6000] free=17.883 GiB cached_plan_shapes=4200 idle_handles=4200
[2150/6000] free=15.824 GiB cached_plan_shapes=4300 idle_handles=4300
[2200/6000] free=13.848 GiB cached_plan_shapes=4400 idle_handles=4400
[2250/6000] free=11.781 GiB cached_plan_shapes=4500 idle_handles=4500
[2300/6000] free=9.734 GiB cached_plan_shapes=4600 idle_handles=4600
[2350/6000] free=7.656 GiB cached_plan_shapes=4700 idle_handles=4700
[2400/6000] free=5.615 GiB cached_plan_shapes=4800 idle_handles=4800
[2450/6000] free=3.584 GiB cached_plan_shapes=4900 idle_handles=4900
[2500/6000] free=1.533 GiB cached_plan_shapes=5000 idle_handles=5000

=== REPRODUCED after 2524 distinct plan sizes (last size = 5062) ===
cached_plan_shapes=5033 idle_handles=5033 free=576.000 MiB
ROCFFTError(code rocfft_status_failure, an operation failed)
Stacktrace:
  [1] check
    @ /projappl/project_462000008/decristoforo/.julia/packages/AMDGPU/HUKS4/src/fft/error.jl:38 [inlined]
  [2] macro expansion
    @ /projappl/project_462000008/decristoforo/.julia/packages/AMDGPU/HUKS4/src/utils.jl:239 [inlined]
  [3] rocfft_plan_create(plan::Base.RefValue{Ptr{AMDGPU.rocFFT.rocfft_plan_t}}, placement::AMDGPU.rocFFT.rocfft_result_placement_e, transform_type::AMDGPU.rocFFT.rocfft_transform_type_e, precision::AMDGPU.rocFFT.rocfft_precision_e, dimensions::Int64, lengths::Vector{Int64}, number_of_transforms::Int64, description::Ptr{Nothing})
    @ AMDGPU.rocFFT /projappl/project_462000008/decristoforo/.julia/packages/AMDGPU/HUKS4/src/fft/librocfft.jl:92
  [4] create_plan(xtype::AMDGPU.rocFFT.rocfft_transform_type_e, xdims::Tuple{Int64}, T::Type, inplace::Bool, region::Tuple{Int64})
    @ AMDGPU.rocFFT /projappl/project_462000008/decristoforo/.julia/packages/AMDGPU/HUKS4/src/fft/wrappers.jl:40
  [5] #get_plan##0
    @ /projappl/project_462000008/decristoforo/.julia/packages/AMDGPU/HUKS4/src/fft/wrappers.jl:11 [inlined]
  [6] pop!(f::AMDGPU.rocFFT.var"#get_plan##0#get_plan##1"{AMDGPU.rocFFT.rocfft_transform_type_e, Tuple{Int64}, Type{Float32}, Bool, Tuple{Int64}}, cache::HandleCache{Tuple{HIPContext, AMDGPU.rocFFT.rocfft_transform_type_e, NTuple{N, Int64} where N, Type, Bool, Any}, Tuple{Ptr{AMDGPU.rocFFT.rocfft_plan_t}, Int64}}, key::Tuple{HIPContext, AMDGPU.rocFFT.rocfft_transform_type_e, Tuple{Int64}, DataType, Bool, Tuple{Int64}})
    @ AMDGPU /projappl/project_462000008/decristoforo/.julia/packages/AMDGPU/HUKS4/src/cache.jl:49
  [7] get_plan
    @ /projappl/project_462000008/decristoforo/.julia/packages/AMDGPU/HUKS4/src/fft/wrappers.jl:11 [inlined]
  [8] plan_rfft(X::ROCArray{Float32, 1, AMDGPU.Runtime.Mem.HIPBuffer}, region::Tuple{Int64})
    @ AMDGPU.rocFFT /projappl/project_462000008/decristoforo/.julia/packages/AMDGPU/HUKS4/src/fft/fft.jl:149
  [9] run_one!(len::Int64)
    @ Main ./REPL[4]:8
 [10] main(n_sizes::Int64)
    @ Main ./REPL[6]:17
 [11] top-level scope
    @ REPL[7]:1
 [12] __repl_entry_eval_expanded_with_loc(mod::Module, ast::Any, toplevel_file::Ref{Ptr{UInt8}}, toplevel_line::Ref{Int32})
    @ REPL /pfs/lustrep4/users/decristoforo/.julia/juliaup/julia-1.12.7+0.x64.linux.gnu/share/julia/stdlib/v1.12/REPL/src/REPL.jl:301
 [13] toplevel_eval_with_hooks(mod::Module, ast::Any, toplevel_file::Any, toplevel_line::Any)
    @ REPL /pfs/lustrep4/users/decristoforo/.julia/juliaup/julia-1.12.7+0.x64.linux.gnu/share/julia/stdlib/v1.12/REPL/src/REPL.jl:308
 [14] toplevel_eval_with_hooks(mod::Module, ast::Any, toplevel_file::Any, toplevel_line::Any)
    @ REPL /pfs/lustrep4/users/decristoforo/.julia/juliaup/julia-1.12.7+0.x64.linux.gnu/share/julia/stdlib/v1.12/REPL/src/REPL.jl:312
 [15] toplevel_eval_with_hooks
    @ /pfs/lustrep4/users/decristoforo/.julia/juliaup/julia-1.12.7+0.x64.linux.gnu/share/julia/stdlib/v1.12/REPL/src/REPL.jl:305 [inlined]
 [16] eval_user_input(ast::Any, backend::REPL.REPLBackend, mod::Module)
    @ REPL /pfs/lustrep4/users/decristoforo/.julia/juliaup/julia-1.12.7+0.x64.linux.gnu/share/julia/stdlib/v1.12/REPL/src/REPL.jl:330
 [17] repl_backend_loop(backend::REPL.REPLBackend, get_module::Function)
    @ REPL /pfs/lustrep4/users/decristoforo/.julia/juliaup/julia-1.12.7+0.x64.linux.gnu/share/julia/stdlib/v1.12/REPL/src/REPL.jl:452
 [18] start_repl_backend(backend::REPL.REPLBackend, consumer::Any; get_module::Function)
    @ REPL /pfs/lustrep4/users/decristoforo/.julia/juliaup/julia-1.12.7+0.x64.linux.gnu/share/julia/stdlib/v1.12/REPL/src/REPL.jl:427
 [19] start_repl_backend
    @ /pfs/lustrep4/users/decristoforo/.julia/juliaup/julia-1.12.7+0.x64.linux.gnu/share/julia/stdlib/v1.12/REPL/src/REPL.jl:424 [inlined]
 [20] run_repl(repl::REPL.AbstractREPL, consumer::Any; backend_on_current_task::Bool, backend::Any)
    @ REPL /pfs/lustrep4/users/decristoforo/.julia/juliaup/julia-1.12.7+0.x64.linux.gnu/share/julia/stdlib/v1.12/REPL/src/REPL.jl:653
 [21] run_repl(repl::REPL.AbstractREPL, consumer::Any)
    @ REPL /pfs/lustrep4/users/decristoforo/.julia/juliaup/julia-1.12.7+0.x64.linux.gnu/share/julia/stdlib/v1.12/REPL/src/REPL.jl:639
 [22] run_std_repl(REPL::Module, quiet::Bool, banner::Symbol, history_file::Bool)
    @ Base ./client.jl:478
 [23] run_main_repl(interactive::Bool, quiet::Bool, banner::Symbol, history_file::Bool)
    @ Base ./client.jl:499
 [24] repl_main
    @ ./client.jl:586 [inlined]
 [25] _start()
    @ Base ./client.jl:561
RESULT: FAIL (reproduced) after 2524 distinct FFT plan sizes

Contributor guide

No contributing guide indexed for this repository

First steps

  1. Read the whole issue, then the project's contributing guide.
  2. Comment on the issue to say you are picking it up — it saves two people doing the same work.
  3. Fork the repository and make your change on a branch.
  4. Open a pull request that references the issue number.

Research direction

Start by locating the IDLE_HANDLES cache in the rocFFT-related code and inspect how plans for distinct shapes are retained. Exercise the cache with distinct-shape plans and verify that completed plans are evicted and GPU memory no longer grows without bound.

Written by the indexing model from the issue text.

Assessment

Tech stack
julia
Domain
hpc
Issue type
Bug
Difficulty
4/5
Estimated time
3-5 days
Activity status
Active
Clarity
Needs clarification
Newbie friendliness
45/100

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.