Meta drops 30B agentic model Muse Glimmer under Apache 2.0 — runs on a single GPU
MarkTechPost · Aug 10, 2026 · 2 min read
Meta just open-sourced Muse Glimmer, a 30-billion-parameter model purpose-built for always-on local agent workflows. It...
11 articles
Explore our coverage of quantization — 11 curated articles, summaries, and related resources from the Inblix archive.
MarkTechPost · Aug 10, 2026 · 2 min read
Meta just open-sourced Muse Glimmer, a 30-billion-parameter model purpose-built for always-on local agent workflows. It...
Machine Learning Mastery · Jul 29, 2026 · 2 min read
If you've spent any time running small language models locally, you've probably noticed the elephant in the room: Ollam...
MarkTechPost · Jul 23, 2026 · 3 min read
Three products, three different answers to the question "what is a model," and zero cases telling us which one a court...
The Decoder · Jul 15, 2026 · 3 min read
PrismML, a startup founded by Caltech researchers, just released Bonsai 27B — a 27-billion-parameter model that runs lo...
Hugging Face Blog · Oct 15, 2025 · 2 min read
Running AI on your own hardware usually means buying a pricey GPU. But a new workflow from Hugging Face and Intel lets...
Hugging Face Blog · Aug 5, 2025 · 3 min read
OpenAI just did something it rarely does: it open-sourced a family of models. Meet GPT-OSS, a pair of reasoning-focused...
Hugging Face Blog · Apr 29, 2025 · 2 min read
Here's a number that makes other quantization researchers wince: 2.1x higher relative accuracy at 2-bit precision. That...
Hugging Face Blog · Sep 20, 2024 · 2 min read
Running a model like Meta-Llama-3.1-8B on a laptop or edge device is moving from 'technically possible' to genuinely pr...
Hugging Face Blog · Sep 18, 2024 · 2 min read
Microsoft Research’s BitNet architecture promised a world where LLM parameters are just -1, 0, or 1, slashing memory an...
Hugging Face Blog · Aug 13, 2024 · 2 min read
The library that quietly powers llama.cpp and Ollama is worth a close look for any developer who's ever winced at a PyT...
Hugging Face Blog · Jul 30, 2024 · 2 min read
Running Stable Diffusion 3 in FP16 eats 18.765 GB of GPU memory. That's a problem for anyone without a datacenter card....