Deformable convolution best practice?
Nobody has claimed this yet.
- Dominant language
- Python
- Stars
- 17.9k
- Forks
- 7.3k
- Avg merge
- 1d 15h
- Merged PRs (30d)
- 13
Description
❓ Questions and Help
Would appreciate it if anyone has some insight on how to use deformable convolution correctly.
Deformable convolution is tricky as even the official implementation is different from what's described in the paper. The paper claims to use 2N offset size instead of 2 x ks x ks.
Anyway, we're using the 2 x ks x ks offset here, but I always got poor performance. Accuracy drops in CIFAR10 and YOLACT. Anything wrong with my usage?
from torchvision.ops import DeformConv2d
class DConv(nn.Module):
def __init__(self, inplanes, planes, kernel_size=3, stride=1, padding=1, bias=False):
super(DConv, self).__init__()
self.conv1 = nn.Conv2d(inplanes, 2 * kernel_size * kernel_size, kernel_size=kernel_size,
stride=stride, padding=padding, bias=bias)
self.conv2 = DeformConv2d(inplanes, planes, kernel_size=kernel_size, stride=stride, padding=padding, bias=bias)
def forward(self, x):
out = self.conv1(x)
out = self.conv2(x, out)
return out
Contributor guide
First steps
- Read the whole issue, then the project's contributing guide.
- Comment on the issue to say you are picking it up — it saves two people doing the same work.
- Fork the repository and make your change on a branch.
- Open a pull request that references the issue number.
Research direction
Start with the shown DConv example and the torchvision.ops.DeformConv2d entry point, then compare the offset shape described in the paper with the 2 × ks × ks setup used here. Done means establishing whether this usage is correct and explaining or reproducing the CIFAR10 and YOLACT accuracy drop.
Written by the indexing model from the issue text.
Assessment
- Tech stack
- python, pytorch
- Domain
- computer-vision, machine-learning
- Issue type
- Bug
- Difficulty
- 5/5
- Estimated time
- Over a week
- Activity status
- Stale
- Clarity
- Needs clarification
- Newbie friendliness
- 20/100