AI Pulse by Inblix

AMD buys Taalas to bake Llama 3.1 into silicon at 16,000 tokens per second

The Decoder · Aug 7, 2026 · 2 min read · Read original article →

Curated by the Inblix editorial team


Featured image for article: AMD buys Taalas to bake Llama 3.1 into silicon at 16,000 tokens per second

AMD is acquiring Toronto-based startup Taalas, a company that came out of stealth just last February with a genuinely weird idea: instead of designing a general-purpose AI accelerator, they hard-wire a specific model’s architecture and trained weights directly into the silicon. The result is a chip that can’t run anything else, but runs its one model absurdly fast. The company’s demo silicon reportedly cranked out over 16,000 tokens per second per user on Llama 3.1-8B — a figure that makes most inference benchmarks look like they’re standing still.

It’s a bet that the market is heading toward specialization at the hardware layer, not just the software layer. If you know exactly which model you’re deploying at scale, Taalas’s approach eliminates the overhead that comes from general-purpose matrix multipliers and memory hierarchies designed for flexibility. You’re left with pure throughput. Google appears to be thinking along similar lines for Gemini, though the search giant hasn’t shared performance numbers publicly.

AMD plans to fold the technology into its broader accelerator roadmap, offering it alongside Instinct GPUs as a system-level option rather than a standalone product. That’s the right call — nobody is swapping out their entire inference fleet for a fixed-function chip, but dedicated hardware for a workhorse model like Llama could slot neatly into high-volume production pipelines. Vamsi Boppana, SVP of AMD’s AI group, framed it as strengthening the portfolio, while Taalas co-founder Ljubica Bajic pointed to the scale AMD can provide.

The deal is still subject to regulatory approvals, which shouldn’t be a roadblock given Taalas’s size and the non-controversial nature of the tech. The real question is whether model-hardened silicon ages well when the models themselves get updated every few months. A chip that’s locked to Llama 3.1-8B is a depreciating asset the moment Llama 4 ships. For AMD, that tension between specialization and obsolescence will define whether this acquisition looks prescient or just premature.

💡 Key Takeaways

  1. Taalas embeds a model's architecture and trained weights directly into silicon, making inference extremely fast but locking the chip to a single model forever.
  2. A demo chip running Llama 3.1-8B hit over 16,000 tokens per second per user, vastly outpacing conventional GPU-based inference.
  3. AMD plans to offer this fixed-function hardware alongside its Instinct GPUs as a system-level solution rather than replacing its general-purpose accelerator line.

Keep reading: See related articles below for more coverage on this topic.

Get smarter about AI

The sharpest AI news, curated daily. Delivered free to your inbox.

Learn more

Glossary terms

← Back to all articles