AI Pulse by Inblix

Topic: H100

3 articles

Explore our coverage of H100 — 3 curated articles, summaries, and related resources from the Inblix archive.

TNG runs 5,000+ LLM inferences per hour on 24 H100s—here's the prefill/decode trade-off — Inblix summary
Research

TNG runs 5,000+ LLM inferences per hour on 24 H100s—here's the prefill/decode trade-off

Hugging Face Blog · Apr 16, 2025 · 2 min read

Most teams treat LLM inference as a black box. TNG doesn't have that luxury. They're running over 5,000 inferences per...

Hugging Face taps FriendliAI, ranked fastest GPU inference, for 1-click H100 deployment — Inblix summary
Research

Hugging Face taps FriendliAI, ranked fastest GPU inference, for 1-click H100 deployment

Hugging Face Blog · Jan 22, 2025 · 2 min read

Hugging Face just gave its users a faster on-ramp to production AI. A new partnership with FriendliAI—ranked by Artific...

Google Cloud A3 nodes can now run Meta's 405B Llama 3.1, but you'll need 8 H100s — Inblix summary
Research

Google Cloud A3 nodes can now run Meta's 405B Llama 3.1, but you'll need 8 H100s

Hugging Face Blog · Aug 19, 2024 · 2 min read

If you want to run Meta's monster 405-billion-parameter Llama 3.1 model without it melting your hardware, Google Cloud'...