[Feature Request] 3B Model
未關閉
- 主要語言
- Python
- 星號
- 19.5k
- 分支
- 1.6k
- PR 合併指標
- 30 天內沒有已合併 PR
描述
### 🚀 The feature, motivation and pitch
I have seen that many users have constrained hardware choices/or are converting pdfs at large scales. Would it be beneficial to train a 3B model (based on Qwen2.5-vl)?
Besides the speedup from switching to a smaller model, we can also further tune for speculative decoding methods such as EAGLE which in turn would make inference of the base model faster.
If that sounds reasonable, I would love to help and contribute to creating this.
### Alternatives
_No response_
### Additional context
(For me specifically, I have to convert ~1M pdfs and I don't have enough compute to do this in a reasonable timeframe)
貢獻指南
評估
這個 Issue 還沒有評估資料。