AI Pulse by Inblix

Topic: long-context

7 articles

Explore our coverage of long-context — 7 curated articles, summaries, and related resources from the Inblix archive.

Pokee's 28B model holds 10 million tokens on a single GPU — and stays inside your firewall — Inblix summary
AI News

Pokee's 28B model holds 10 million tokens on a single GPU — and stays inside your firewall

MarkTechPost · Aug 8, 2026 · 2 min read

Pokee AI just shipped a model that solves a problem most enterprises are too busy working around to admit they have. Po...

GLM-5.2 crushes open-source coding leaders, lands within 4 points of Claude Opus 4.8 — Inblix summary
Research

GLM-5.2 crushes open-source coding leaders, lands within 4 points of Claude Opus 4.8

Hugging Face Blog · Jun 17, 2026 · 3 min read

The open-source model race just got a lot more interesting. The team behind GLM has dropped GLM-5.2, and it's not just...

Snowflake’s Ulysses trick trains transformers on million-token contexts across GPUs — Inblix summary
Research

Snowflake’s Ulysses trick trains transformers on million-token contexts across GPUs

Hugging Face Blog · Mar 9, 2026 · 2 min read

The quadratic cost of attention has been the brick wall for training transformers on truly long sequences. You can’t ju...

Falcon-H1-Arabic Hits 256K Context, Outperforms Bigger Models with Hybrid Design — Inblix summary
Research

Falcon-H1-Arabic Hits 256K Context, Outperforms Bigger Models with Hybrid Design

Hugging Face Blog · Jan 5, 2026 · 2 min read

The team behind Falcon-Arabic just dropped a sequel that isn't a minor tune-up — it's a ground-up rebuild. Falcon-H1-Ar...

Meta drops Llama 4: 10M context window and a bold bet on NoPE layers — Inblix summary
Research

Meta drops Llama 4: 10M context window and a bold bet on NoPE layers

Hugging Face Blog · Apr 5, 2025 · 2 min read

Forget the slow drip of model releases. Meta just dumped two Llama 4 variants on Hugging Face—Maverick and Scout—and th...

NVIDIA’s KVPress toolkit slashes 1M-token Llama 3 memory from 330GB to fit on a single GPU — Inblix summary
Research

NVIDIA’s KVPress toolkit slashes 1M-token Llama 3 memory from 330GB to fit on a single GPU

Hugging Face Blog · Jan 23, 2025 · 3 min read

If you’ve ever tried to run a model with a million-token context window, you already know the math is brutal. For Llama...

Falcon Mamba: The First 7B Model to Ditch Attention Without Performance Loss — Inblix summary
Research

Falcon Mamba: The First 7B Model to Ditch Attention Without Performance Loss

Hugging Face Blog · Aug 12, 2024 · 2 min read

The team behind Falcon has released Falcon Mamba, a 7-billion-parameter model that stands as the first general-purpose,...