google-deepmind / google-deepmind/emergent_in_context_learning
Have you experiment with smaller size model
Open
- Dominant language
- Python
- Stars
- 88
- Forks
- 23
- PR merge metrics
- No merged PRs in 30d
Description
Hi, in the paper, a transformer model with 12 layers and embdding size 64 is used to validate your hypothesis. Did you do any trial experiment on smaller sized model ? and if you did, what's the result ?
Contributor guide
Assessment
This issue has not been assessed yet.