google-deepmind / google-deepmind/emergent_in_context_learning

Have you experiment with smaller size model

Open
#4 0 comments 0 reactions 0 assignees View on GitHub
Dominant language
Python
Stars
88
Forks
23
PR merge metrics
No merged PRs in 30d

Description

Hi, in the paper, a transformer model with 12 layers and embdding size 64 is used to validate your hypothesis. Did you do any trial experiment on smaller sized model ? and if you did, what's the result ?

Contributor guide

Open the contributing guide

Assessment

This issue has not been assessed yet.

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.