Accelerate resample_by_picking through better ordering
- Dominant language
- Python
- Stars
- 34
- Forks
- 25
- Avg merge
- 3h 37m
- Merged PRs (30d)
- 2
Description
`resample_by_picking` routinely shows up at the top of our GPU profiles. Here's an example from a run (mirgecom wave-eager, `nel_1d = 24`, 3D, order 3):
```
GPU activities: 15.95% 1.72898s 28160 61.398us 3.4240us 133.22us resample_by_picking
14.57% 1.57934s 26499 59.599us 4.4480us 117.47us multiply
14.41% 1.56223s 9698 161.09us 16.000us 543.23us diff
12.79% 1.38640s 21601 64.182us 4.4150us 118.56us axpbyz
11.66% 1.26342s 5283 239.15us 120.54us 529.85us grudge_assign_0
8.81% 954.48ms 23428 40.741us 1.6640us 81.888us axpb
7.93% 859.20ms 10560 81.363us 60.831us 135.04us resample_by_mat
7.84% 849.96ms 1760 482.93us 481.79us 541.57us face_mass
2.16% 233.96ms 62 3.7735ms 1.3440us 11.178ms [CUDA memcpy DtoH]
1.58% 171.51ms 2235 76.738us 1.6640us 87.391us [CUDA memcpy DtoD]
1.19% 128.99ms 3523 36.612us 19.328us 37.952us [CUDA memset]
0.49% 53.375ms 440 121.31us 120.67us 136.42us grudge_assign_1
0.49% 53.364ms 440 121.28us 120.67us 136.19us grudge_assign_2
0.07% 7.2597ms 127 57.162us 1.1520us 847.90us [CUDA memcpy HtoD]
0.03% 3.5939ms 12 299.49us 13.696us 529.66us nodes_0
0.01% 1.4677ms 6 244.62us 14.112us 528.73us actx_special_sqrt
0.00% 359.81us 6 59.967us 5.0880us 115.14us divide
0.00% 136.06us 1 136.06us 136.06us 136.06us actx_special_exp
```
It's especially striking that it's at the top of the list because it touches lower-dimensional data. (surface vs volume) `multiply` has a similar number of calls, but it touches volume data, and it completes more quickly.
I think there are two opportunities here that we could try:
- Currently, the kernel has an indirection on the read and the write end (see the (very simple) [source](https://github.com/inducer/meshmode/blob/3af83a45f134014c1506bd46430cafc9641de9c9/meshmode/discretization/connection/direct.py#L296-L297)). For surjective/onto connections, we can do away with the indirection on write by appropriately sorting the source indices.
- Even for non-surjective connections, it's likely that we would benefit by sorting by the write index, to try to keep the writes as coalesced as possible.
IMO it's likely that this will have a benefit (but I obviously can't guarantee it). I think it's worth trying.
cc @lukeolson
Contributor guide
No contributing guide indexed for this repository
Assessment
This issue has not been assessed yet.