allenai / allenai/OLMo

Differences and exact configs of allenai/OLMo-1B-0724-hf vs allenai/OLMo-1B-hf

未關閉
#901 1 則留言 0 個 reaction 已指派 0 人 在 GitHub 檢視
type/question
主要語言
Python
星號
6.7k
分支
797
PR 合併指標
30 天內沒有已合併 PR

描述

### ❓ The question

Hi! I'm trying to understand the differences between these two models. Looking at the HF configs and model pages I found the following ones:

- Difference training process, 1 stage vs 2 stages and different version of dolma v1_5 vs v_1_7
- Context length for training, 2048 vs 4096
- embedding weight tying , tied vs untied
- clip_qkv, null vs 8.0

I also looked at the config at `configs/official-0724/OLMo-1B.yaml` but can't figure out it is supposed to be the 0724 version or the first one. Are there training runs/exact configs for these exact models?

Thanks for any info!

貢獻指南

開啟貢獻指南

評估

這個 Issue 還沒有評估資料。

把新 issue 寄到你的電子郵件信箱

精選適合新手參與的 GitHub issue 摘要。