AI Pulse by Inblix

AMD buys AI chip startup Taalas to etch models directly into silicon, hits 17,000 tokens/sec

The Register AI · Aug 10, 2026 · 2 min read · Read original article →

Curated by the Inblix editorial team


Featured image for article: AMD buys AI chip startup Taalas to etch models directly into silicon, hits 17,000 tokens/sec

AMD isn’t just shopping for faster chips. It’s betting that the future of AI inference looks less like a general-purpose GPU and more like a custom-made suit. The company has acquired Taalas, a startup that etches entire AI models directly into silicon, creating model-specific integrated circuits that one early demo showed churning out up to 17,000 tokens per second.

That number is staggering. For context, it’s an order of magnitude beyond what even the most optimized software stacks on top-tier hardware can typically manage for large models. The trade-off, of course, is flexibility. A chip etched for Llama doesn’t run Mistral. But for cloud providers and enterprises serving a single model at planetary scale—think a ChatGPT-sized inference workload—the efficiency gains could rewrite the economics of deployment. This isn’t just about speed; it’s about the cost per token plummeting.

Details on the deal’s price tag are conspicuously absent from the announcement, but the strategic logic fits neatly into AMD’s aggressive push against Nvidia’s data center dominance. Nvidia’s Blackwell platform leans on brute-force programmability. AMD, with this acquisition, is signaling an alternative path: if you know exactly what model you’re running, why carry the overhead of a programmable pipeline? It’s a bet on maturity. As foundation models stabilize, the industry’s need for bespoke, ultra-efficient inference hardware grows.

This move also parallels a broader trend where hardware is becoming a direct extension of the software stack. We’ve seen similar thinking from Groq and Cerebras, but Taalas’s approach pushes the concept to its logical extreme by physically baking the model’s weights into the chip’s architecture. The gamble is that the velocity of model innovation will slow enough to make a fixed-function chip economical before it’s rendered obsolete by the next architecture release. Given how much money hyperscalers are pouring into single-model deployments, it’s a bet that might just pay off.

💡 Key Takeaways

  1. Early demos of Taalas's model-specific chips reached 17,000 tokens per second, a speed that dramatically outpaces conventional GPU-based inference.
  2. The acquisition signals AMD is pursuing a bifurcated strategy: general-purpose GPUs for training and custom-etched silicon for massive-scale, single-model inference workloads.
  3. Baking model weights directly into silicon eliminates programmable overhead, potentially slashing per-token costs for enterprises serving a single model at scale.
  4. The approach hinges on the assumption that major foundation models will stabilize long enough for fixed-function chips to be economical before they are outdated.

Keep reading: See related articles below for more coverage on this topic.

Get smarter about AI

The sharpest AI news, curated daily. Delivered free to your inbox.

Learn more

Glossary terms

← Back to all articles