[Model Submission] Nanbeige4-3B-Thinking-2511: Open-sourced 3B model with solid reasoning capability
Nobody has claimed this yet.
- Dominant language
- Python
- Stars
- 39.5k
- Forks
- 4.8k
- PR merge metrics
- No merged PRs in 30d
Description
We are pleased to submit our recently open-sourced model, Nanbeige4-3B-Thinking-2511, for evaluation in the Chatbot Arena. This model has achieved a score of 60 points on the Arena-Hard V2 benchmark via advanced distillation techniques and reinforcement learning optimization.
Model Details:
- Organization: Nanbeige LLM Lab
- Model Name: Nanbeige4-3B-Thinking-2511
- HuggingFace Link: https://huggingface.co/Nanbeige/Nanbeige4-3B-Thinking-2511
Despite its small size (3B parameters), Nanbeige4-3B-Thinking-2511 exhibits competitive performance across a broad range of tasks, making it a compelling case study for efficient language modeling. We believe its inclusion in Chatbot Arena will provide valuable insights into the capabilities of parameter-efficient models in real-world, human-preference-driven evaluations.
We are happy to provide any additional support required for evaluation, such as API access.
Thank you for your consideration. We look forward to your feedback and the opportunity to contribute to the community.
Contributor guide
No contributing guide indexed for this repository
First steps
- Read the whole issue, then the project's contributing guide.
- Comment on the issue to say you are picking it up — it saves two people doing the same work.
- Fork the repository and make your change on a branch.
- Open a pull request that references the issue number.
Research direction
Start with the Chatbot Arena evaluation process and the linked Hugging Face model page for Nanbeige4-3B-Thinking-2511. Determine how model submissions are evaluated and what additional support or integration details are required; done means the model has been evaluated or included in Chatbot Arena.
Written by the indexing model from the issue text.
Assessment
- Tech stack
- huggingface
- Domain
- machine-learning
- Issue type
- Feature
- Difficulty
- 5/5
- Estimated time
- Over a week
- Activity status
- Stale
- Clarity
- Needs clarification
- Newbie friendliness
- 20/100