graykode / graykode/nlp-tutorial
BiLstm(tf) maybe have mistake
- Dominant language
- Jupyter Notebook
- Stars
- 14.9k
- Forks
- 3.9k
- PR merge metrics
- No merged PRs in 30d
Description
calculate attention_score
`
# Attention
outputs = tf.concat([output[0], output[1]], 2) # output[0] : lstm_fw, output[1] : lstm_bw
outputs = tf.transpose(outputs, [1, 0, 2]) # [n_step, batch_size, n_hidden]
# 只用了最后一个步长的输出
final_hidden_state = outputs[-1]
output_all = tf.concat([output[0], output[1]], 2)
final_hidden_state = tf.expand_dims(final_hidden_state, 2)
attn_weights = tf.squeeze(tf.matmul(output_all, final_hidden_state), 2) `
Contributor guide
Research direction
Start by locating the BiLstm(tf) implementation and reviewing the shown attention-score calculation, especially the final hidden state and tensor dimensions. Verify whether the resulting attention weights match the intended bidirectional LSTM attention behavior; done means resolving and correcting the suspected mistake, if confirmed.
Written by the indexing model from the issue text.
Assessment
- Tech stack
- tensorflow
- Domain
- machine-learning
- Issue type
- Bug
- Difficulty
- 4/5
- Estimated time
- 3-5 days
- Activity status
- Stale
- Clarity
- Needs clarification
- Newbie friendliness
- 20/100