AI Pulse by Inblix

Topic: inference

15 articles

Explore our coverage of inference — 15 curated articles, summaries, and related resources from the Inblix archive.

Related Terms

AMD buys AI chip startup Taalas to etch models directly into silicon, hits 17,000 tokens/sec — Inblix summary
Industry

AMD buys AI chip startup Taalas to etch models directly into silicon, hits 17,000 tokens/sec

The Register AI · Aug 10, 2026 · 2 min read

AMD isn't just shopping for faster chips. It's betting that the future of AI inference looks less like a general-purpos...

AMD buys Taalas to bake Llama 3.1 into silicon at 16,000 tokens per second — Inblix summary
AI News

AMD buys Taalas to bake Llama 3.1 into silicon at 16,000 tokens per second

The Decoder · Aug 7, 2026 · 2 min read

AMD is acquiring Toronto-based startup Taalas, a company that came out of stealth just last February with a genuinely w...

AMD buys Taalas to bake AI models into silicon, hitting 17,000 tokens/sec — Inblix summary
Industry

AMD buys Taalas to bake AI models into silicon, hitting 17,000 tokens/sec

The Register AI · Aug 6, 2026 · 2 min read

AMD just made a fascinating bet on the future of AI hardware, acquiring a startup called Taalas that's pursuing a radic...

Runware's shipping-container data pods slash AI inference costs without water — Inblix summary
Startups

Runware's shipping-container data pods slash AI inference costs without water

TechCrunch AI · Aug 4, 2026 · 2 min read

Runware just threw a shipping container-sized wrench into the AI infrastructure race. The company unveiled its Sonic In...

AMD and Cerebras team up to take on Nvidia's Groq LPUs together — Inblix summary
Industry

AMD and Cerebras team up to take on Nvidia's Groq LPUs together

The Register AI · Jul 23, 2026 · 2 min read

The enemy of my enemy is my friend. That's the logic driving a new partnership between AMD and Cerebras, two companies...

Etched hits $10.3B valuation on $300M raise, claiming its AI chips run faster at lower voltage — Inblix summary
Startups

Etched hits $10.3B valuation on $300M raise, claiming its AI chips run faster at lower voltage

TechCrunch AI · Jul 23, 2026 · 2 min read

Etched, the AI chip startup founded by three Harvard dropouts, just closed a $300 million Series C led by Sequoia, doub...

Google's 'Frozen v2' chip hardcodes Gemini's structure for 10x faster AI — Inblix summary
AI News

Google's 'Frozen v2' chip hardcodes Gemini's structure for 10x faster AI

The Decoder · Jul 20, 2026 · 2 min read

Google is building a server chip that doesn't just run its Gemini AI model — it has the model's architecture physically...

The chatbot era is dead: Open Responses aims to fix the agent mess — Inblix summary
Research

The chatbot era is dead: Open Responses aims to fix the agent mess

Hugging Face Blog · Jan 15, 2026 · 2 min read

The Chat Completion API is a zombie. It's still shuffling around as the de facto standard for AI inference, but it was...

NVIDIA's single NIM container now deploys over 100,000 Hugging Face LLMs without manual config — Inblix summary
Research

NVIDIA's single NIM container now deploys over 100,000 Hugging Face LLMs without manual config

Hugging Face Blog · Jul 21, 2025 · 2 min read

NVIDIA is making a hard play to become the default engine for the open-source AI boom. The company announced that its N...

Hugging Face’s async robot brains cut task time in half by never waiting to think — Inblix summary
Research

Hugging Face’s async robot brains cut task time in half by never waiting to think

Hugging Face Blog · Jul 10, 2025 · 2 min read

The dirty secret of most robot learning models isn't their failure rate—it's how much time they spend just sitting ther...

SGLang taps Hugging Face transformers as a backend, instantly unlocking models like Kyutai's Helium — Inblix summary
Research

SGLang taps Hugging Face transformers as a backend, instantly unlocking models like Kyutai's Helium

Hugging Face Blog · Jun 23, 2025 · 2 min read

SGLang, the inference engine built for high-throughput and low-latency AI, just closed a major gap in its ecosystem. It...

Groq's LPU chips land on Hugging Face, bringing sub-100ms inference to open models — Inblix summary
Research

Groq's LPU chips land on Hugging Face, bringing sub-100ms inference to open models

Hugging Face Blog · Jun 16, 2025 · 2 min read

The wait for truly fast, on-demand open-source AI just got shorter. Groq, the inference provider that famously eschews...

Hugging Face offloads VAE decoding to the cloud, slashing GPU memory for Flux and video models — Inblix summary
Research

Hugging Face offloads VAE decoding to the cloud, slashing GPU memory for Flux and video models

Hugging Face Blog · Feb 24, 2025 · 2 min read

Running a cutting-edge image or video model on a consumer GPU often hits a wall—not during the diffusion process itself...

Hugging Face's TGI adds vLLM and TRT-LLM backends to end the serving wars — Inblix summary
Research

Hugging Face's TGI adds vLLM and TRT-LLM backends to end the serving wars

Hugging Face Blog · Jan 16, 2025 · 2 min read

Hugging Face is finally addressing the fragmentation that has turned LLM serving into a choose-your-own-headache exerci...

Hugging Face kills HUGS microservice just 2 years after launch — Inblix summary
Research

Hugging Face kills HUGS microservice just 2 years after launch

Hugging Face Blog · Oct 23, 2024 · 2 min read

Hugging Face has quietly pulled the plug on HUGS, its Generative AI Services platform for deploying open models, markin...