bigscience-workshop / bigscience-workshop/bigscience

Sharing the 1.3B-Pile@300B model

Open
#46 0 comments 0 reactions 0 assignees View on GitHub
Dominant language
Shell
Stars
1k
Forks
102
PR merge metrics
No merged PRs in 30d

Description

The 1.3B-Pile@300B model is quite strong:
https://docs.google.com/spreadsheets/d/1CI8Q9RCblLRzUOPJ6ViqBmo284-8ojluQ-CmaEuhuv0/edit#gid=1295801165

lambada 0.6088 piqa 0.7160 hellaswag 0.5209 --> these are all better than gpt-neo 1.3B.

Could you share the model? Thank you.

Contributor guide

Open the contributing guide

Assessment

This issue has not been assessed yet.

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.