ai-forever / ai-forever/Kandinsky-3

MoVQ implementation question

Aperta
#10 2 commenti 3 reazioni 0 assegnatari Vedi su GitHub
Lingua principale
Python
Stelle
396
Fork
40
Metriche di merge delle PR
Nessuna PR unita negli ultimi 30g

Descrizione

I have some questions regarding the implementation of MoVQ and would appreciate your clarification.

From the original MoVQ paper, it is mentioned that a multi-channle VQ is adopted.
![image](https://github.com/ai-forever/Kandinsky-3/assets/29164878/ec8c5fb9-e019-48b3-bb6a-60a436232b24)

However, the implementation of kandinsky3 does not involve any vector quantization operation:
```python
class MoVQ(nn.Module):

def __init__(self, generator_params):
super().__init__()
z_channels = generator_params["z_channels"]
self.encoder = Encoder(**generator_params)
self.quant_conv = torch.nn.Conv2d(z_channels, z_channels, 1)
self.post_quant_conv = torch.nn.Conv2d(z_channels, z_channels, 1)
self.decoder = Decoder(zq_ch=z_channels, **generator_params)

@torch.no_grad()
def encode(self, x):
h = self.encoder(x)
h = self.quant_conv(h)
return h

@torch.no_grad()
def decode(self, quant):
decoder_input = self.post_quant_conv(quant)
decoded = self.decoder(decoder_input, quant)
return decoded
```

May I ask if it is a misunderstanding on my part regarding MoVQ, or if Kandinsky has made some modifications to the implementation of MoVQ?

Guida per i contributori

Nessuna guida per i contributori indicizzata per questo repository

Valutazione

Questa issue non è ancora stata valutata.

Ricevi le nuove issue nella tua casella

Un breve riepilogo di issue GitHub adatte ai principianti.