AI Pulse by Inblix

AMD buys Taalas to bake AI models into silicon, hitting 17,000 tokens/sec

The Register AI · Aug 6, 2026 · 2 min read · Read original article →

Curated by the Inblix editorial team


Featured image for article: AMD buys Taalas to bake AI models into silicon, hitting 17,000 tokens/sec

AMD just made a fascinating bet on the future of AI hardware, acquiring a startup called Taalas that’s pursuing a radical idea: etching entire AI models directly into silicon. Forget general-purpose GPUs running software—Taalas builds model-specific integrated circuits. Early tech demos are reportedly spitting out up to 17,000 tokens a second. That’s not a typo. If those numbers hold up in production, we’re talking about inference speeds that make today’s GPU clusters look like they’re standing still.

The logic is straightforward, even if the engineering is brutal. When you hard-wire a model like Llama or a specific diffusion transformer into silicon, you eliminate the massive overhead of fetching instructions and shuttling data between memory and compute. It’s the same reason an ASIC miner crushes a GPU for Bitcoin. The catch, of course, is flexibility. A chip built for Llama 3 70B is a paperweight if Llama 4 changes its architecture. That’s the wager AMD is making: that the market is large enough, and the performance gains are dramatic enough, to justify what amounts to custom silicon for frozen models.

AMD’s move feels like a direct counterpunch to the custom chip ambitions at hyperscalers. Google has its TPUs. Amazon has Trainium and Inferentia. Microsoft and Meta are both reportedly developing their own inference ASICs. By acquiring Taalas, AMD can offer a similar one-two punch it’s used in CPUs and GPUs: sell off-the-shelf Instinct GPUs for training and experimentation, then pitch a Taalas-derived ASIC when a customer’s model is stable and needs to serve billions of queries cheaply.

Whether this pans out depends entirely on execution. The history of AI chip startups is a graveyard—Graphcore, Wave Computing, and the famously overhyped Nervana come to mind. But AMD has the balance sheet and the customer relationships to absorb a moonshot. If Taalas’s technology can shorten the design cycle for a model-specific chip from years to months, it could reshape what the inference market looks like three years from now. That’s a big “if,” and I’ll believe it when I see it in a customer’s data center rather than a press release.

💡 Key Takeaways

  1. Taalas builds chips that hard-wire specific AI models into silicon, a stark departure from programmable GPUs.
  2. Early demos claim 17,000 tokens per second, a speed that could slash the cost and latency of deploying large models at scale.
  3. This acquisition pits AMD's custom-silicon strategy directly against in-house chips from Google, Amazon, and other cloud giants.
  4. The trade-off is brutal: a chip customized for one model version becomes obsolete the moment the model architecture changes.

Keep reading: See related articles below for more coverage on this topic.

Get smarter about AI

The sharpest AI news, curated daily. Delivered free to your inbox.

Learn more

Glossary terms

← Back to all articles