JuliaGPU / JuliaGPU/KernelAbstractions.jl
BUG: @synchronize inside loop in CPU is not working
Nobody has claimed this yet.
- Dominant language
- Julia
- Stars
- 523
- Forks
- 88
- Avg merge
- 1d 11h
- Merged PRs (30d)
- 25
Description
Hi! I realized that there is a bug using CPU backend (the same code is working in CUDA backend)
This is the example:
using KernelAbstractions, CUDA
const backend = CPU()
const BLOCK_SIZE = 4
@kernel function kk(A, B, C)
varr = 2
for i in 1:varr
for j in 1:varr
@synchronize()
C[i, j] = 0
for k in 1:varr
C[i, j] += A[i, k] * B[k, j]
end
end
end
end
function run_gpu()
m = 10
n = 20
#Inicializo las matrices en la GPU
A = KernelAbstractions.zeros(backend, Int, m, n)
B = KernelAbstractions.zeros(backend, Int, m, n)
C = KernelAbstractions.zeros(backend, Int, m, n)
#Calculo el tamaño de bloque
block_size = BLOCK_SIZE
mn = max(m, n)
if mn < BLOCK_SIZE
block_size = mn
end
#Calculo el número de bloques
total_blocks = (mn + block_size - 1) ÷ block_size
#Anti-diagonal loop
@time @inbounds for diag in 0:(2*total_blocks-1)
#Número de bloques a lanzar en la anti-diagonal
num_blocks_diagonal = min(diag + 1, 2 * total_blocks - diag - 1)
kernel! = kk(backend)
kernel!(A, B, C, ndrange = (block_size * block_size, num_blocks_diagonal), workgroupsize = block_size * block_size)
KernelAbstractions.synchronize(backend)
end
end
run_gpu()
and this is the error:
ERROR: LoadError: UndefVarError: `varr` not defined
Stacktrace:
[1] cpu_kk
@ ~/.julia/packages/KernelAbstractions/Zcyra/src/macros.jl:288 [inlined]
[2] cpu_kk(__ctx__::KernelAbstractions.CompilerMetadata{KernelAbstractions.NDIteration.DynamicSize, KernelAbstractions.NDIteration.NoDynamicCheck, CartesianIndex{2}, CartesianIndices{2, Tuple{Base.OneTo{Int64}, Base.OneTo{Int64}}}, KernelAbstractions.NDIteration.NDRange{2, KernelAbstractions.NDIteration.DynamicSize, KernelAbstractions.NDIteration.DynamicSize, CartesianIndices{2, Tuple{Base.OneTo{Int64}, Base.OneTo{Int64}}}, CartesianIndices{2, Tuple{Base.OneTo{Int64}, Base.OneTo{Int64}}}}}, A::Matrix{Int64}, B::Matrix{Int64}, C::Matrix{Int64})
@ Main ./none:0
[3] __thread_run(tid::Int64, len::Int64, rem::Int64, obj::KernelAbstractions.Kernel{CPU, KernelAbstractions.NDIteration.DynamicSize, KernelAbstractions.NDIteration.DynamicSize, typeof(cpu_kk)}, ndrange::Tuple{Int64, Int64}, iterspace::KernelAbstractions.NDIteration.NDRange{2, KernelAbstractions.NDIteration.DynamicSize, KernelAbstractions.NDIteration.DynamicSize, CartesianIndices{2, Tuple{Base.OneTo{Int64}, Base.OneTo{Int64}}}, CartesianIndices{2, Tuple{Base.OneTo{Int64}, Base.OneTo{Int64}}}}, args::Tuple{Matrix{Int64}, Matrix{Int64}, Matrix{Int64}}, dynamic::KernelAbstractions.NDIteration.NoDynamicCheck)
@ KernelAbstractions ~/.julia/packages/KernelAbstractions/Zcyra/src/cpu.jl:115
[4] __run(obj::KernelAbstractions.Kernel{CPU, KernelAbstractions.NDIteration.DynamicSize, KernelAbstractions.NDIteration.DynamicSize, typeof(cpu_kk)}, ndrange::Tuple{Int64, Int64}, iterspace::KernelAbstractions.NDIteration.NDRange{2, KernelAbstractions.NDIteration.DynamicSize, KernelAbstractions.NDIteration.DynamicSize, CartesianIndices{2, Tuple{Base.OneTo{Int64}, Base.OneTo{Int64}}}, CartesianIndices{2, Tuple{Base.OneTo{Int64}, Base.OneTo{Int64}}}}, args::Tuple{Matrix{Int64}, Matrix{Int64}, Matrix{Int64}}, dynamic::KernelAbstractions.NDIteration.NoDynamicCheck, static_threads::Bool)
@ KernelAbstractions ~/.julia/packages/KernelAbstractions/Zcyra/src/cpu.jl:82
[5] (::KernelAbstractions.Kernel{CPU, KernelAbstractions.NDIteration.DynamicSize, KernelAbstractions.NDIteration.DynamicSize, typeof(cpu_kk)})(::Matrix{Int64}, ::Vararg{Matrix{Int64}}; ndrange::Tuple{Int64, Int64}, workgroupsize::Int64)
@ KernelAbstractions ~/.julia/packages/KernelAbstractions/Zcyra/src/cpu.jl:44
[6] Kernel
@ ~/.julia/packages/KernelAbstractions/Zcyra/src/cpu.jl:37 [inlined]
[7] macro expansion
@ /app/manuel/test3.jl:48 [inlined]
[8] macro expansion
@ ./timing.jl:279 [inlined]
[9] run_gpu()
@ Main /app/manuel/test3.jl:44
[10] top-level scope
@ /app/manuel/test3.jl:55
So, it's like after @synchronize, all local variables are lost and this only happens in the CPU backend.
Any thoughts on this? Thank you very much!
Contributor guide
No contributing guide indexed for this repository
First steps
- Read the whole issue, then the project's contributing guide.
- Comment on the issue to say you are picking it up — it saves two people doing the same work.
- Fork the repository and make your change on a branch.
- Open a pull request that references the issue number.
Research direction
Start by reproducing the example and inspect the CPU backend stack frames in src/cpu.jl and the generated kernel code around macros.jl:288. Compare how @synchronize preserves local variables in the CUDA backend, then verify that variables remain available after synchronization in the CPU backend with a regression test based on the reported example.
Written by the indexing model from the issue text.
Assessment
- Tech stack
- julia
- Domain
- backend
- Issue type
- Bug
- Difficulty
- 4/5
- Estimated time
- 3-5 days
- Activity status
- Stale
- Clarity
- Mostly clear
- Newbie friendliness
- 38/100