Questions: Will OLMo4 use MoE, Linear Attention, BitNet etc.?
Abierto
- Lenguaje dominante
- Python
- Estrellas
- 1.5k
- Forks
- 315
- Merge medio
- 1 d 9 h
- PR fusionados (30 d)
- 11
Descripción
There are many techniques used by established firms that seems to accelerate LLM power, but OLMo is not using any of them.
P.S. Also there are decay-free training, evolutionary model merging, debiased RL methods, self-play etc. for improving existing models. Wonder if these will be investigated in a later time as well
Guía de contribución
Evaluación
Este issue todavía no se ha evaluado.