Grok 4.6 ties GPT-5.6 Sol at 61 points while costing 60% less
The Decoder · Aug 12, 2026 · 2 min read
SpaceXAI's Grok 4.6 just landed on the Artificial Analysis Intelligence Index with a score of 61 — dead even with OpenA...
71 articles
Explore our coverage of benchmarks — 71 curated articles, summaries, and related resources from the Inblix archive.
The Decoder · Aug 12, 2026 · 2 min read
SpaceXAI's Grok 4.6 just landed on the Artificial Analysis Intelligence Index with a score of 61 — dead even with OpenA...
The Decoder · Aug 12, 2026 · 2 min read
Microsoft has rolled out MAI Code 1.1 Flash, a new code-generation model purpose-built for GitHub Copilot, but the laun...
The Decoder · Aug 8, 2026 · 2 min read
xAI just dropped Imagine Image 2.0, and the Arena leaderboard numbers tell a clear story: it's fast, it's capable, and...
The Decoder · Aug 6, 2026 · 2 min read
Alibaba's Qwen3.8 Max just posted a 56 on the Artificial Analysis Intelligence Index — a 10-point leap over its predece...
The Decoder · Aug 6, 2026 · 2 min read
Meta just dropped Muse Spark 1.2, a coding-focused upgrade that doubles as the company's first dedicated coding agent....
MarkTechPost · Aug 4, 2026 · 3 min read
The standard Python visualization stack hits a wall at a few hundred thousand points. Reflex AI just open-sourced XY un...
The Decoder · Aug 3, 2026 · 2 min read
Alibaba just dropped a model that doesn't just answer prompts — it clocks in for a multi-day shift. Qwen3.8-Max, a 2.4-...
MarkTechPost · Aug 3, 2026 · 3 min read
Alibaba's Qwen team just flipped the switch on general availability for Qwen3.8-Max and confirmed that open weights for...
MarkTechPost · Aug 1, 2026 · 2 min read
Supabase just gave every AI-coding-pessimist a receipt. The company open-sourced supabase/evals, an Apache-2.0 benchmar...
The Decoder · Jul 30, 2026 · 2 min read
The ARC-AGI-3 leaderboard just got messy. OpenAI claims GPT-5.6 Sol scored 38.3 percent on the tough reasoning benchmar...
AI News · Jul 29, 2026 · 3 min read
OpenAI published a field report this week that sounds like a greatest-hits reel for coding agents in research. It track...
The Decoder · Jul 27, 2026 · 2 min read
Moonshot AI just put real weight behind its frontier-model ambitions. The Beijing-based company dropped the model weigh...
MarkTechPost · Jul 26, 2026 · 2 min read
Kuaishou’s KwaiKAT team just dropped KAT-Coder-V2.5, and the most interesting part isn't the model weights—it's the tra...
The Decoder · Jul 26, 2026 · 2 min read
Anthropic's Claude Opus 5 didn't just edge past the competition on the ARC-AGI-3 benchmark—it obliterated the previous...
MarkTechPost · Jul 26, 2026 · 2 min read
Sakana AI dropped Fugu-Cyber on Monday, and no, it’s not a new frontier model. It’s a cybersecurity-tuned endpoint that...
MarkTechPost · Jul 25, 2026 · 2 min read
Datalab dropped Marker 2 this week, and the numbers are worth paying attention to. It is a ground-up rewrite of their o...
MarkTechPost · Jul 24, 2026 · 2 min read
Anthropic dropped Claude Opus 5 today, and the biggest shock isn't the benchmark scores — it's that your existing API i...
The Decoder · Jul 24, 2026 · 2 min read
Sakana AI just dropped an update to Fugu Ultra, its AI model router, and the claim is genuinely weird: version 1.1 can...
MarkTechPost · Jul 22, 2026 · 2 min read
The fine-tuning wars aren't about which library you use anymore. They're about where each project places its engineerin...
MarkTechPost · Jul 22, 2026 · 2 min read
Poolside dropped Laguna S 2.1 on Thursday, a 118-billion-parameter open-weight coding model that punches wildly above i...
The Decoder · Jul 21, 2026 · 2 min read
Alibaba just claimed the top spot on a major text-to-speech leaderboard with a model that sounds great — so long as you...
OpenAI Blog · Jul 21, 2026 · 2 min read
OpenAI just pulled back the curtain on its technical roadmap, and it's refreshingly candid about one thing: they don't...
MarkTechPost · Jul 19, 2026 · 3 min read
Most text-to-SQL systems treat the task like translation: turn a question into a query and hope for the best. The probl...
OpenAI Blog · Jul 19, 2026 · 2 min read
The era of training reinforcement learning agents entirely in comfortable, consequence-free simulators might be coming...
The Decoder · Jul 19, 2026 · 2 min read
Moonshot's Kimi K3 just pulled off something no Chinese model has managed before: it grabbed the number one spot on the...
OpenAI Blog · Jul 19, 2026 · 2 min read
OpenAI just dropped a truth bomb wrapped in 16 pixelated environments. Their new Procgen Benchmark isn't just another s...
OpenAI Blog · Jul 19, 2026 · 2 min read
OpenAI has formally introduced Codex, the code-generating descendant of GPT-3 that's been powering GitHub Copilot behin...
OpenAI Blog · Jul 18, 2026 · 2 min read
Here's a finding that should give anyone deploying large language models serious pause: the bigger the model, the more...
OpenAI Blog · Jul 18, 2026 · 2 min read
The classic left-to-right way of training language models might finally have a permanent co-pilot. A new paper from res...
OpenAI Blog · Jul 16, 2026 · 2 min read
OpenAI just dropped an early version of a new model, o1-preview, that doesn't just incrementally improve on GPT-4o — it...
The Decoder · Jul 16, 2026 · 3 min read
Thinking Machines Lab, the startup Mira Murati built after leaving OpenAI, just shipped its first model — and it's a st...
OpenAI Blog · Jul 16, 2026 · 2 min read
SWE-bench has been the gold standard for proving your AI can code like a real engineer. Top models barely scraped 20% o...
OpenAI Blog · Jul 16, 2026 · 2 min read
OpenAI isn't calling it GPT-5. They're resetting the counter to 1. The company just released o1-preview, the first in a...
OpenAI Blog · Jul 15, 2026 · 2 min read
OpenAI just open-sourced a new factuality benchmark, and the results are humbling. The dataset, called SimpleQA, is des...
OpenAI Blog · Jul 15, 2026 · 2 min read
OpenAI pulled back the curtain on Computer-Using Agent, the model powering its new Operator research preview. CUA ditch...
OpenAI Blog · Jul 15, 2026 · 2 min read
The question isn't whether AI can code anymore. It's whether it can get paid for it. OpenAI's new SWE-Lancer benchmark...
OpenAI Blog · Jul 14, 2026 · 2 min read
OpenAI just announced the Pioneers Program, a hands-on initiative that pairs its research teams directly with startups...
OpenAI Blog · Jul 14, 2026 · 3 min read
OpenAI just dropped a pair of new models—o3 and o4-mini—and for the first time, the chain-of-thought process includes a...
OpenAI Blog · Jul 14, 2026 · 2 min read
OpenAI isn't just releasing new models. With o3 and o4-mini, they're fundamentally changing how their AI reasons about...
OpenAI Blog · Jul 13, 2026 · 3 min read
OpenAI has a new theory for why chatbots keep making stuff up, and it points the finger squarely at the people building...
OpenAI Blog · Jul 12, 2026 · 2 min read
OpenAI just admitted something that anyone working in multilingual AI has known for years: the benchmarks we use to mea...
Hugging Face Blog · Jun 30, 2026 · 2 min read
AI-assisted enterprise modernization just hit a reality check. ScarfBench, a new open benchmark designed to stress-test...
Hugging Face Blog · Jun 30, 2026 · 2 min read
The mess of AI evaluation just got a little easier to navigate. The EvalEval Coalition's EEE project, launched in Febru...
Hugging Face Blog · Jun 24, 2026 · 2 min read
The assumption that a speech recognition model that aces a clean benchmark will work in your kitchen or car is, frankly...
Hugging Face Blog · Jun 17, 2026 · 3 min read
The open-source model race just got a lot more interesting. The team behind GLM has dropped GLM-5.2, and it's not just...
Hugging Face Blog · Apr 1, 2026 · 2 min read
The team behind Falcon Perception asked a question most of the field has been dodging: why are vision systems still sti...
Hugging Face Blog · Feb 12, 2026 · 2 min read
Meta and Hugging Face have released OpenEnv, an open-source framework that hooks AI agents up to real-world tools inste...
Hugging Face Blog · Feb 4, 2026 · 2 min read
Hugging Face just took a swing at one of AI’s dirtiest open secrets: benchmark scores you can’t trust. The platform ann...
Hugging Face Blog · Feb 3, 2026 · 2 min read
H Company just dropped a preview of its largest UI localization model yet, and the numbers are worth paying attention t...
Hugging Face Blog · Nov 21, 2025 · 2 min read
The Open ASR Leaderboard just got a major expansion, adding multilingual and long-form transcription tracks that reveal...
Hugging Face Blog · Oct 7, 2025 · 2 min read
Forget reading code and guessing if it works. BigCodeArena, a new platform launched in February, has already logged ove...
Hugging Face Blog · Oct 1, 2025 · 2 min read
Cohere just dropped a retrieval benchmark that's designed to embarrass models that have been gaming the system. The Ret...
Hugging Face Blog · Sep 22, 2025 · 2 min read
The original GAIA benchmark is practically cooked. Two years after its 2023 release, the top models are acing the harde...
Hugging Face Blog · Jul 16, 2025 · 2 min read
For years, the industry has guessed whether bidirectional encoders or causal decoders were better for non-generative ta...
Hugging Face Blog · Jul 4, 2025 · 2 min read
The blank stare of a loss curve is a familiar frustration for anyone who has trained a large language model. In the ear...
Hugging Face Blog · Jun 6, 2025 · 2 min read
Evaluating AI that can actually use a computer like a person — by looking at the screen — remains surprisingly hard to...
Hugging Face Blog · Jun 3, 2025 · 2 min read
H Company just released Holo1, an open-source family of action VLMs that makes web automation dramatically cheaper with...
Hugging Face Blog · Mar 26, 2025 · 2 min read
The newest model from DeepSeek isn't a flashy launch. It's a silent update that speaks volumes. DeepSeek-V3-0324, an up...
Hugging Face Blog · Mar 20, 2025 · 2 min read
The open-source model OlympicCoder 7B is now topping the LiveCodeBench leaderboard, outperforming the daily-driver codi...
Hugging Face Blog · Mar 12, 2025 · 2 min read
Google just dropped Gemma 3, and the numbers are genuinely surprising. The new open-weight model family runs from a tin...
Hugging Face Blog · Mar 11, 2025 · 2 min read
The Open R1 project just dropped a genuinely surprising result: a fine-tuned 32-billion-parameter model that out-codes...
Hugging Face Blog · Mar 4, 2025 · 2 min read
Cohere For AI just dropped a pair of open-weight vision-language models that punch well above their weight class — and...
Hugging Face Blog · Feb 14, 2025 · 2 min read
Here's something you don't see every day: a three-line code change completely reshuffled the top 20 rankings on the mos...
Hugging Face Blog · Feb 4, 2025 · 2 min read
If you're a data analyst waiting for AI to take the boring parts of your job off your plate, you're going to be waiting...
Hugging Face Blog · Feb 4, 2025 · 2 min read
OpenAI’s new Deep Research tool is undeniably slick—it browses the web, synthesizes multi-step answers, and just scored...
Hugging Face Blog · Dec 20, 2024 · 2 min read
Give a state-of-the-art model a logic puzzle in writing, and it aces it. Read the same puzzle aloud, and that performan...
Hugging Face Blog · Dec 17, 2024 · 2 min read
The Technology Innovation Institute just dropped the Falcon 3 family, and the headliner is a 10-billion-parameter model...
Hugging Face Blog · Nov 20, 2024 · 2 min read
Evaluating Japanese large language models has been a mess. The language's unique mix of kanji, hiragana, katakana, and...
Hugging Face Blog · Nov 19, 2024 · 2 min read
Atla just dropped Judge Arena, a platform that flips the script on how we evaluate AI evaluators. Instead of relying so...
Hugging Face Blog · Oct 23, 2024 · 2 min read
The original CinePile dataset launched in May 2024 with a genuinely impressive stat: humans outperformed the best comme...
Hugging Face Blog · Oct 1, 2024 · 2 min read
The best open-source large language models still stumble badly over Czech, and now we have the receipts. A new evaluati...