allenai / allenai/olmpool

Request for 10B / 50B long-ctx training configs or hyperparameters

Offen
#1 0 Kommentare 0 Reaktionen 0 zugewiesene Personen Auf GitHub ansehen
Vorherrschende Sprache
Python
Sterne
7
Forks
1
PR-Merge-Kennzahlen
Keine gemergten PRs in 30 T.

Beschreibung

Hi OlmPool team,

Thanks for releasing OlmPool! We are very interested in this project and are currently looking into reproducing and extending some of the experiments.

I was wondering whether you could provide the training configs (or a description of the key hyperparameters) used for the **10B- and 50B-token long-ctx experiments**.

In particular, I'm curious about how the training horizon and learning-rate schedule were configured. For example, were both runs configured similarly to pretraining with the schedule defined for **50B tokens**, while the 10B run was obtained by a **hard stop at 10B tokens**? Or were separate schedules/hyperparameters used for the 10B and 50B runs?

Having the corresponding configs, or even just a brief description of the differences from the released configs, would be very helpful for reproduction.

Thanks again for releasing this work!

Beitragsleitfaden

Für dieses Repository ist kein Beitragsleitfaden indexiert

Bewertung

Dieses Issue wurde noch nicht bewertet.

Neue Issues direkt in Ihr Postfach

Eine kurze Übersicht über anfängerfreundliche GitHub-Issues.