google / google/gemma.cpp

Feature Request / Proposal: "E12B" (Effective 12B) Model Class for Gemma 5 (Optimized for 12GB VRAM / Desktop Users)

Open
#990 0 comments 0 reactions 0 assignees View on GitHub
Dominant language
C++
Stars
7k
Forks
660
Avg merge
20h 43m
Merged PRs (30d)
33

Description

**Hi Gemma Team,**

First off, thank you for the incredible work on the Gemma 4 series! The E2B and E4B models have been game-changers for light/mobile workloads.

However, there is a massive gap in the desktop consumer landscape: the **12GB VRAM bracket** (e.g., RTX 4070 / 3060 12GB users), which represents a huge portion of local LLM enthusiasts and independent developers.

While standard 12B Dense models run well, they often trade off reasoning capability compared to 27B–35B models. On the other hand, running 27B+ models locally requires offloading or heavy quantization that severely degrades speed.

**Proposal for Gemma 5:**
Introducing an **E12B (Effective 12B)** architecture (e.g., a sparse/MoE or active-parameter architecture that yields 27B–35B intelligence while maintaining a 12B active VRAM/RAM footprint).

This would hit the absolute "**sweet spot**" for desktop hardware, allowing high-speed, high-context generation without forcing users into massive VRAM upgrades during current hardware/RAM market constraints.

Bringing the "Effective" architecture scaling up to the 12B class in Gemma 5 would empower millions of local deployment users.

Thanks for considering this feedback!

Contributor guide

Open the contributing guide

Research direction

No files, tests, or entry points are named in the issue. The proposal needs maintainer-defined architecture and implementation scope before a contributor can identify a concrete change or a clear completion check.

Written by the indexing model from the issue text.

Assessment

Tech stack
cpp
Domain
ai, machine-learning, performance
Issue type
Feature
Difficulty
5/5
Estimated time
Over a week
Activity status
Active
Clarity
Needs clarification
Newbie friendliness
25/100

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.