Reproduction of Olmo3
Aperta
- Lingua principale
- Python
- Stelle
- 395
- Fork
- 105
- Metriche di merge delle PR
- Nessuna PR unita negli ultimi 30g
Descrizione
Hi Dear Author,
I'm trying to reproduce Olmo3's result, but for intermediate checkpoints, whenever I wanted to eval on mmlu(mmlu:cot::olmo3:midtrain), it seems the memory just gradually increases until OOM, seems the COT is too long?
the model I was trying to reproduce is: https://huggingface.co/allenai/Olmo-3-7B-Think-SFT
Much appreciated!
Guida per i contributori
Nessuna guida per i contributori indicizzata per questo repository
Valutazione
Questa issue non è ancora stata valutata.