[ARCH-PROP] Root arch: Olmo3 & Qwen3
- Dominant language
- No language data
- Stars
- 85
- Forks
- 6
- PR merge metrics
- No merged PRs in 30d
Description
### Architecture Name
Root
### Parent issue
None
### Motivations
OLMo 3 and Qwen3 should be added as the initial root architectures in ArchSpace because they are exemplars of open large language model development.
### Proposed Architecture
N/A
### Preliminary Results (if any)
- The OLMo 3 technical report presents OLMo 3 as a family of state-of-the-art, fully open language models with 7B and 32B parameter scales. It emphasizes the release of the full model flow, including stages, checkpoints, data points, and dependencies. The official Ai2 blog introduces OLMo 3 as a release intended to make the model-development process open and inspectable, including Base, Think, Instruct, and RL Zero variants.
- The Qwen3 technical report describes Qwen3 as a family of dense and MoE models ranging from 0.6B to 235B parameters. It reports hybrid thinking and non-thinking modes, support for 119 languages and dialects, and public availability under Apache 2.0. The official Qwen3 blog lists two MoE models, Qwen3-235B-A22B and Qwen3-30B-A3B, plus dense models from 0.6B to 32B. It also reports approximately 36 trillion pretraining tokens, long-context variants up to 128K context length, and a post-training pipeline involving long chain-of-thought cold start, reasoning reinforcement learning, thinking-mode fusion, and general reinforcement learning.
These two model families already provide substantial public evidence and documentation, making them appropriate root references for future ArchSpace architecture proposals.
References:
- https://arxiv.org/abs/2512.13961
- https://allenai.org/blog/olmo3
- https://arxiv.org/abs/2505.09388
- https://qwenlm.github.io/blog/qwen3/
### Experiments Plan
N/A
Contributor guide
No contributing guide indexed for this repository
Research direction
No implementation files, tests, or entry points are named. Start by reading the linked OLMo 3 and Qwen3 technical reports and official blog posts, then inspect how ArchSpace represents existing root architectures; done means both model families are added as documented root references with their supporting evidence.
Written by the indexing model from the issue text.
Assessment
- Tech stack
- machine-learning
- Domain
- ai, machine-learning
- Issue type
- Feature
- Difficulty
- 4/5
- Estimated time
- 3-5 days
- Activity status
- Quiet
- Clarity
- Needs clarification
- Newbie friendliness
- 35/100