AI Pulse by Inblix

IBM's new Granite Nano models pack agentic smarts into a 350M parameter footprint

Hugging Face Blog · Oct 28, 2025 · 2 min read · Read original article →

Curated by the Inblix editorial team


Featured image for article: IBM's new Granite Nano models pack agentic smarts into a 350M parameter footprint

IBM isn’t trying to win the race to a trillion parameters. With the new Granite 4.0 Nano family, they’re sprinting in the complete opposite direction. The release covers four instruct models and their base counterparts, ranging from a 350-million-parameter lightweight up to a roughly 1.5-billion-parameter dense model. The real technical talking point is the hybrid-SSM architecture in the ‘H’ variants, which IBM says delivers a significant performance leap over traditional transformers of a similar size.

These aren’t just toys for simple classification tasks. IBM is positioning these sub-billion-parameter models for serious agentic workloads. On the Berkeley Function Calling Leaderboard v3 (BFCLv3) and the IFEval instruction-following benchmark, the Nano models are posting numbers that outpace similarly sized offerings from Alibaba’s Qwen, Google’s Gemma, and LiquidAI. It’s a direct shot across the bow of the small language model market, suggesting you don’t need a massive model to handle tool calling if the training data and architecture are optimized ruthlessly.

The governance angle is a quiet but crucial part of this launch. All the Granite 4.0 models carry IBM’s ISO 42001 certification for responsible model development. In a landscape where enterprise adoption often stalls on compliance and risk, having an internationally recognized stamp for responsible AI governance baked in from day one is a differentiator that data scientists might overlook but general counsels won’t. The models are available under the permissive Apache 2.0 license and have native support baked into vLLM, llama.cpp, and MLX, lowering the friction for developers who just want to pull and run something on a MacBook or a Raspberry Pi.

Looking ahead, IBM is framing this as just one slice of a broader Granite 4.0 expansion. The 15 trillion-token training data diet and refined pipelines hint that the efficiency gains here will ripple into whatever they ship next. For developers tired of burning cloud credits on oversized models just to get reliable JSON output from a function call, a 350M-parameter engine that actually works on-device is a welcome arrival.

💡 Key Takeaways

  1. The Granite 4.0 H Nano models use a hybrid-SSM architecture that outperforms standard transformers in the sub-2B parameter class on key benchmarks.
  2. IBM's small models are specifically optimized for agentic tasks, beating competitors like Gemma and Qwen on the BFCLv3 function calling and IFEval instruction-following metrics.
  3. The inclusion of ISO 42001 certification for responsible AI development gives enterprise users a concrete governance advantage that most open-weight small models lack.

Keep reading: See related articles below for more coverage on this topic.

Get smarter about AI

The sharpest AI news, curated daily. Delivered free to your inbox.

Learn more

Glossary terms

← Back to all articles