lucidrains / lucidrains/vector-quantize-pytorch

Why do I get almost the same codes after the 1st batch?

Open
#131 1 comment 2 reactions 0 assignees View on GitHub

Nobody has claimed this yet.

Dominant language
Python
Stars
4k
Forks
338
PR merge metrics
No merged PRs in 30d

Description

Hi there, I am trying to quantize my input feature sparse_feat with the following codes in my network.

class MyModel(nn.Module):
    def __init__ (self):
        super(MyModel, self).__init__()
        self.residual_vq = ResidualVQ(
                    dim = 256,
                    codebook_size = 128,
                    num_quantizers = 3,
                    threshold_ema_dead_code = 2
                )
    def forward(self, sparse_feat):
        quantized_sparse, codes_sparse, commit_loss_sparse = self.residual_vq(sparse_feat)


model = MyModel()
optimizer = torch.optim.Adam(model.parameters(), lr=args.lr)

loop_inner = tqdm(enumerate(dataloader, 0), total=len(dataloader), leave=True)
for idx, (x, y) in loop_inner:
    quantized_sparse, codes_sparse, commit_loss_sparse = model(x)
    optimizer.zero_grad()
    loss.backward()
    optimizer.step()

However, I've observed that, in the first batch, the codes I got were uniformly distributed. In the second and following batches, the codes and features within codes_sparse and quantized_sparse were almost the same. The following codes were I got for the second batch

tensor([[ 44,  13,  82],
        [ 44,  13,  82],
        [ 44,  13,  82],
        [ 44,  13,  82],
        [ 44,  13,  82],
        [ 44,  13,  82],
        [ 44, 111,  82],
        [ 44,  13,  82],
        X 16 times,
        [ 44, 111,  82],
        [ 44,  13,  82],
        X 11 times,
        [ 44, 111,  82],
        [ 44,  13,  82],
        X 5 times,
        [ 44, 111,  82],
        [ 44,  13,  82],
         X 11 times
        [ 44, 111,  82],
        [ 44, 111,  82],
        [ 44,  13,  82],
        X 8 times])

Note that the input features within sparse_feat are not alike, so I am wondering what could go wrong here? Am I not configuring the quantization procedure right?
Looking forward to your helpful suggestions. Appreciated in advance.

Contributor guide

No contributing guide indexed for this repository

First steps

  1. Read the whole issue, then the project's contributing guide.
  2. Comment on the issue to say you are picking it up — it saves two people doing the same work.
  3. Fork the repository and make your change on a branch.
  4. Open a pull request that references the issue number.

Research direction

Start by reproducing the shown MyModel and training loop with ResidualVQ, then compare sparse_feat, codes_sparse, and quantized_sparse across the first batches. Check whether the repeated codes follow from the quantizer configuration or training behavior; done means identifying the cause and documenting a reproducible correction or confirming expected behavior.

Written by the indexing model from the issue text.

Assessment

Tech stack
python, pytorch
Domain
machine-learning
Issue type
Bug
Difficulty
4/5
Estimated time
3-5 days
Activity status
Stale
Clarity
Mostly clear
Newbie friendliness
25/100

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.