AdityaNG / AdityaNG/kan-gpt

CUDA out of memory

Aperta
#18 1 commento 0 reazioni 0 assegnatari Vedi su GitHub
bug help wanted
Lingua principale
Python
Stelle
726
Fork
54
Metriche di merge delle PR
Nessuna PR unita negli ultimi 30g

Descrizione

```
class KanMLP(nn.Module):
"""Some Information about KanLinear"""
def __init__(self,
in_features=1152,
hidden_features = None,
out_features = None,
drop=0.
):
super().__init__()

approx_gelu = lambda: nn.GELU(approximate="tanh")

out_features = out_features or in_features
hidden_features = hidden_features or in_features
self.mlp = nn.ModuleDict(
dict(
c_fc=KAN(width=[in_features, hidden_features]),
c_proj=KAN(width=[hidden_features, out_features]),
act=NewGELU(),
dropout=nn.Dropout(0.0),
)
)
m = self.mlp
self.mlpf = lambda x: m.dropout(
m.c_proj(m.act(m.c_fc(x)))
) # MLP forward


def forward(self, x):
x = self.mlpf(x)
return x

net = KanMLP(1152,1152*4).to("cuda")
x = torch.rand(size=(4,4096*4,1152)).to("cuda")
nex(x)

When the number of tokens reaches a certain size, the following situation will occur

CUDA out of memory.

Guida per i contributori

Apri la guida per i contributori

Valutazione

Questa issue non è ancora stata valutata.

Ricevi le nuove issue nella tua casella

Un breve riepilogo di issue GitHub adatte ai principianti.