graykode / graykode/nlp-tutorial
The comment in the Bi-LSTM (Attention) model has an issue.
Aperta
- Lingua principale
- Jupyter Notebook
- Stelle
- 14.9k
- Fork
- 3.9k
- Metriche di merge delle PR
- Nessuna PR unita negli ultimi 30g
Descrizione
The comment `# output : [batch_size, len_seq, n_hidden]` should indeed be corrected to `# output : [batch_size, len_seq, n_hidden*2]` because the Bi-LSTM model is bidirectional. In a bidirectional LSTM, the hidden size is effectively doubled, as it concatenates the forward and backward hidden states. Therefore, the correct shape of the `output` after permutation is `[batch_size, len_seq, n_hidden * 2]`.
Guida per i contributori
Apri la guida per i contributori
Valutazione
Questa issue non è ancora stata valutata.