ROCm / ROCm/FastFlowLM

Feature Request: Support for Bonsai-8B 1-bit Quantized Model

Open
#474 3 comments 4 reactions 0 assignees View on GitHub

Nobody has claimed this yet.

model request
Dominant language
C++
Stars
1.9k
Forks
152
Avg merge
4h 14m
Merged PRs (30d)
11

Description

Feature Request: Support for Bonsai-8B 1-bit Quantized Model

Hello,

I'd like to request support for Prism-ML's Bonsai-8B, a groundbreaking 1-bit quantized model available on Hugging Face:

Model URL: https://huggingface.co/prism-ml/Bonsai-8B-gguf

Why This Matters

Bonsai-8B represents a significant advancement in efficient LLM deployment:

  • Compact Size: An 8B parameter model compressed to just ~1.15 GB
  • Impressive Performance: Achieves an MMLU-R of 65.7, competitive with larger models
  • Energy Efficient: Reported to be 5x more energy efficient than comparable 8B models
  • Dramatically Smaller: 14x smaller footprint than typical 8B models
NPU Potential

The 1-bit quantization approach makes Bonsai-8B an ideal candidate for NPU acceleration. The dramatically reduced memory footprint and simplified computations could enable:

  • Smooth inference on edge devices with NPUs
  • Significantly reduced power consumption
  • Real-time AI capabilities without cloud dependency
  • Privacy-preserving local inference
Request

I'd love to see FastFlowLM add support for this model. Given its efficiency characteristics, it seems like a perfect fit for demonstrating NPU performance advantages with quantized models.

Thank you for considering this request!

Best regards

Contributor guide

Open the contributing guide

First steps

  1. Read the whole issue, then the project's contributing guide.
  2. Comment on the issue to say you are picking it up — it saves two people doing the same work.
  3. Fork the repository and make your change on a branch.
  4. Open a pull request that references the issue number.

Research direction

Start by reviewing how FastFlowLM currently supports quantized models and then inspect the Bonsai-8B-gguf model at the linked Hugging Face URL. Done should mean FastFlowLM supports loading and running this model, including the requested NPU inference path, but the issue does not identify files, tests, or an implementation entry point.

Written by the indexing model from the issue text.

Assessment

Tech stack
cpp, huggingface
Domain
ai, machine-learning
Issue type
Feature
Difficulty
5/5
Estimated time
Over a week
Activity status
Active
Clarity
Needs clarification
Newbie friendliness
35/100

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.