Feedback about NLP From Scratch: Generating Names with a Character-Level RNN
Nobody has claimed this yet.
- Dominant language
- Python
- Stars
- 9.3k
- Forks
- 4.4k
- Avg merge
- 1d 21h
- Merged PRs (30d)
- 4
Description
There is the following issue on this page: https://docs.pytorch.org/tutorials/intermediate/char_rnn_generation_tutorial.html.
There is no non-linearity wrapper while passing the hidden state onto the next iteration, and on output layer that predicts distribution. Results are coming relatively fine though, maybe because nn.LogSoftmax is implicitly acting as non-linearity. But for learning hidden states, its bad, $h_t$ is literally linear combination of $h_{t-1}$, all past $x$ s and $Cat$.
Simple Fix:
class RNN(nn.Module):
def __init__(self,C_in:int, C_hid:int, C_out:int):
super().__init__()
self.C_in = C_in
self.C_hid = C_hid
self.C_out = C_out
self.C_cat = data1.C_cat
self.i2h = nn.Linear(data1.C_cat + C_in + C_hid,C_hid)
self.i2o1 = nn.Linear(data1.C_cat + C_in + C_hid,C_out)
self.o12o2 = nn.Linear(C_out + C_hid,C_hid)
self.o22o3 = nn.Linear(C_hid, C_out)
self.tanh = nn.Tanh()
self.todist = nn.LogSoftmax(dim=1)
def forward(self,x_C:tensor, x_X:tensor, x_H:tensor):
# Inputs = (B=1,C_cat), (B=1, C_in), (B=1,C_out),
# Output = (B=1, C_out), (B=1,C_hid)
x = torch.cat((x_C,x_X,x_H),dim=1)
h = self.tanh(self.i2h(x))
x = self.i2o1(x)
x = torch.cat((x,h),dim=1)
x = self.tanh(self.o12o2(x))
x = self.o22o3(x)
return self.todist(x), h
Edit1: nn.tanh 's range is between $[-1,1]$. Compressing logits before softmax can really impact the result. Using one more layer before projection gave me similar error of 2.2127 , and nice samples.
Thank you.
Contributor guide
First steps
- Read the whole issue, then the project's contributing guide.
- Comment on the issue to say you are picking it up — it saves two people doing the same work.
- Fork the repository and make your change on a branch.
- Open a pull request that references the issue number.
Research direction
Start with the linked NLP From Scratch tutorial page and inspect the character-level RNN implementation and its training example. Reproduce the reported behavior, evaluate whether the proposed activation changes are correct, and update the tutorial only if the issue is confirmed while preserving working training and generated samples.
Written by the indexing model from the issue text.
Assessment
- Tech stack
- python
- Domain
- machine-learning
- Issue type
- Bug
- Difficulty
- 4/5
- Estimated time
- 3-5 days
- Activity status
- Quiet
- Clarity
- Mostly clear
- Newbie friendliness
- 52/100