mcabbott / mcabbott/TensorCast.jl
Performance of nested reductions
Open
Nobody has claimed this yet.
- Dominant language
- Julia
- Stars
- 142
- Forks
- 12
- PR merge metrics
- No merged PRs in 30d
Description
What can be done to improve TensorCast's performance on the following nested reductions:
using TensorCast, BenchmarkTools
M = [i+j for i=1:4, j=0:4:12]
B = [M[i:i+1, j:j+1] for i in 1:2:size(M,1), j in 1:2:size(M,2)]
M2 = reduce(hcat, reduce.(vcat, eachcol(B)))
@cast M3[i⊗k,j⊗l] |= B[k,l][i,j] # \otimes<tab>;
M == M2 == M3 # true
@btime M2 = reduce(hcat, reduce.(vcat, eachcol($B))) # 392 ns (4 allocs: 512 bytes)
@btime @cast M3[i⊗k,j⊗l] |= $B[k,l][i,j] # 1.250 μs (15 allocs: 640 bytes)
Cheers.
Contributor guide
No contributing guide indexed for this repository
First steps
- Read the whole issue, then the project's contributing guide.
- Comment on the issue to say you are picking it up — it saves two people doing the same work.
- Fork the repository and make your change on a branch.
- Open a pull request that references the issue number.
Research direction
Start with the nested-reduction benchmark in the issue and inspect the expansion and execution of the @cast expression for B[k,l][i,j]. Compare it with the reduce(hcat, reduce.(vcat, eachcol(B))) baseline; done means the cast form produces the same result with improved timing and allocation results.
Written by the indexing model from the issue text.
Assessment
- Tech stack
- julia
- Domain
- performance
- Issue type
- Bug
- Difficulty
- 4/5
- Estimated time
- 3-5 days
- Activity status
- Stale
- Clarity
- Needs clarification
- Newbie friendliness
- 35/100