AI Pulse by Inblix

Topic: benchmarks

71 articles

Explore our coverage of benchmarks — 71 curated articles, summaries, and related resources from the Inblix archive.

Grok 4.6 ties GPT-5.6 Sol at 61 points while costing 60% less — Inblix summary
AI News

Grok 4.6 ties GPT-5.6 Sol at 61 points while costing 60% less

The Decoder · Aug 12, 2026 · 2 min read

SpaceXAI's Grok 4.6 just landed on the Artificial Analysis Intelligence Index with a score of 61 — dead even with OpenA...

Microsoft's MAI Code 1.1 Flash gets crushed by DeepSeek on benchmarks and cost — Inblix summary
AI News

Microsoft's MAI Code 1.1 Flash gets crushed by DeepSeek on benchmarks and cost

The Decoder · Aug 12, 2026 · 2 min read

Microsoft has rolled out MAI Code 1.1 Flash, a new code-generation model purpose-built for GitHub Copilot, but the laun...

xAI's Imagine 2.0 scores 1,439 Elo, nipping at OpenAI's heels in image benchmarks — Inblix summary
AI News

xAI's Imagine 2.0 scores 1,439 Elo, nipping at OpenAI's heels in image benchmarks

The Decoder · Aug 8, 2026 · 2 min read

xAI just dropped Imagine Image 2.0, and the Arena leaderboard numbers tell a clear story: it's fast, it's capable, and...

Qwen3.8 Max catches Claude Opus 4.8 but burns 15x the tokens to do it — Inblix summary
AI News

Qwen3.8 Max catches Claude Opus 4.8 but burns 15x the tokens to do it

The Decoder · Aug 6, 2026 · 2 min read

Alibaba's Qwen3.8 Max just posted a 56 on the Artificial Analysis Intelligence Index — a 10-point leap over its predece...

Meta's new Muse Spark 1.2 undercuts rivals at 20 cents—if you hand over your data — Inblix summary
AI News

Meta's new Muse Spark 1.2 undercuts rivals at 20 cents—if you hand over your data

The Decoder · Aug 6, 2026 · 2 min read

Meta just dropped Muse Spark 1.2, a coding-focused upgrade that doubles as the company's first dedicated coding agent....

Reflex's XY charts 100M points in 0.08s, shrinks HTML exports by 1,000x — Inblix summary
AI News

Reflex's XY charts 100M points in 0.08s, shrinks HTML exports by 1,000x

MarkTechPost · Aug 4, 2026 · 3 min read

The standard Python visualization stack hits a wall at a few hundred thousand points. Reflex AI just open-sourced XY un...

Alibaba’s Qwen3.8-Max autonomously ran a business for a year and beat 87% of human coders — Inblix summary
AI News

Alibaba’s Qwen3.8-Max autonomously ran a business for a year and beat 87% of human coders

The Decoder · Aug 3, 2026 · 2 min read

Alibaba just dropped a model that doesn't just answer prompts — it clocks in for a multi-day shift. Qwen3.8-Max, a 2.4-...

Alibaba's Qwen3.8-Max hits 86.6 on Terminal-Bench, open weights drop next week — Inblix summary
AI News

Alibaba's Qwen3.8-Max hits 86.6 on Terminal-Bench, open weights drop next week

MarkTechPost · Aug 3, 2026 · 3 min read

Alibaba's Qwen team just flipped the switch on general availability for Qwen3.8-Max and confirmed that open weights for...

Supabase open-sources an AI agent benchmark — and 78% of coding agents flub auth without a skill file — Inblix summary
AI News

Supabase open-sources an AI agent benchmark — and 78% of coding agents flub auth without a skill file

MarkTechPost · Aug 1, 2026 · 2 min read

Supabase just gave every AI-coding-pessimist a receipt. The company open-sourced supabase/evals, an Apache-2.0 benchmar...

GPT-5.6 Sol hits 38.3% on ARC-AGI-3, but only after OpenAI tweaks the rules — Inblix summary
AI News

GPT-5.6 Sol hits 38.3% on ARC-AGI-3, but only after OpenAI tweaks the rules

The Decoder · Jul 30, 2026 · 2 min read

The ARC-AGI-3 leaderboard just got messy. OpenAI claims GPT-5.6 Sol scored 38.3 percent on the tough reasoning benchmar...

OpenAI’s agents cut scientific code runtimes 60x, but the real bottleneck is verification — Inblix summary
Industry

OpenAI’s agents cut scientific code runtimes 60x, but the real bottleneck is verification

AI News · Jul 29, 2026 · 3 min read

OpenAI published a field report this week that sounds like a greatest-hits reel for coding agents in research. It track...

Kimi K3 delivers 2.5x compute efficiency but flunks UK cyber test — Inblix summary
AI News

Kimi K3 delivers 2.5x compute efficiency but flunks UK cyber test

The Decoder · Jul 27, 2026 · 2 min read

Moonshot AI just put real weight behind its frontier-model ambitions. The Beijing-based company dropped the model weigh...

Kuaishou's KAT-Coder cracks 100K repo tasks by fixing a 16% infrastructure error — Inblix summary
AI News

Kuaishou's KAT-Coder cracks 100K repo tasks by fixing a 16% infrastructure error

MarkTechPost · Jul 26, 2026 · 2 min read

Kuaishou’s KwaiKAT team just dropped KAT-Coder-V2.5, and the most interesting part isn't the model weights—it's the tra...

Claude Opus 5 quadruples ARC-AGI-3 reasoning record, hitting 30.2% — Inblix summary
AI News

Claude Opus 5 quadruples ARC-AGI-3 reasoning record, hitting 30.2%

The Decoder · Jul 26, 2026 · 2 min read

Anthropic's Claude Opus 5 didn't just edge past the competition on the ARC-AGI-3 benchmark—it obliterated the previous...

Sakana’s Fugu-Cyber nudges past GPT-5.5 on one security benchmark—but it’s still gated — Inblix summary
AI News

Sakana’s Fugu-Cyber nudges past GPT-5.5 on one security benchmark—but it’s still gated

MarkTechPost · Jul 26, 2026 · 2 min read

Sakana AI dropped Fugu-Cyber on Monday, and no, it’s not a new frontier model. It’s a cybersecurity-tuned endpoint that...

Marker 2 converts PDFs 5× faster than MinerU while scoring higher on AI2's benchmark — Inblix summary
AI News

Marker 2 converts PDFs 5× faster than MinerU while scoring higher on AI2's benchmark

MarkTechPost · Jul 25, 2026 · 2 min read

Datalab dropped Marker 2 this week, and the numbers are worth paying attention to. It is a ground-up rewrite of their o...

Claude Opus 5 ships with thinking on by default, breaks old API calls — Inblix summary
AI News

Claude Opus 5 ships with thinking on by default, breaks old API calls

MarkTechPost · Jul 24, 2026 · 2 min read

Anthropic dropped Claude Opus 5 today, and the biggest shock isn't the benchmark scores — it's that your existing API i...

Sakana's Fugu router now beats a model it can't even access, claiming 7.9-point gain — Inblix summary
AI News

Sakana's Fugu router now beats a model it can't even access, claiming 7.9-point gain

The Decoder · Jul 24, 2026 · 2 min read

Sakana AI just dropped an update to Fugu Ultra, its AI model router, and the claim is genuinely weird: version 1.1 can...

Unsloth crushes MoE training with 7.3x speedup, Axolotl hits 1.45x on Qwen3.5 — Inblix summary
AI News

Unsloth crushes MoE training with 7.3x speedup, Axolotl hits 1.45x on Qwen3.5

MarkTechPost · Jul 22, 2026 · 2 min read

The fine-tuning wars aren't about which library you use anymore. They're about where each project places its engineerin...

Poolside's Laguna S 2.1 crushes DeepSeek on DeepSWE with one-sixth the active parameters — Inblix summary
AI News

Poolside's Laguna S 2.1 crushes DeepSeek on DeepSWE with one-sixth the active parameters

MarkTechPost · Jul 22, 2026 · 2 min read

Poolside dropped Laguna S 2.1 on Thursday, a 118-billion-parameter open-weight coding model that punches wildly above i...

Alibaba's Qwen TTS edges past rivals in quality, but crawls at 16 chars/second — Inblix summary
AI News

Alibaba's Qwen TTS edges past rivals in quality, but crawls at 16 chars/second

The Decoder · Jul 21, 2026 · 2 min read

Alibaba just claimed the top spot on a major text-to-speech leaderboard with a model that sounds great — so long as you...

OpenAI reveals 3 moonshot projects to measure 'true' AI — Inblix summary
Product

OpenAI reveals 3 moonshot projects to measure 'true' AI

OpenAI Blog · Jul 21, 2026 · 2 min read

OpenAI just pulled back the curtain on its technical roadmap, and it's refreshingly candid about one thing: they don't...

Feyn's SQRL-35B model edges Claude Opus by inspecting databases before writing SQL — Inblix summary
AI News

Feyn's SQRL-35B model edges Claude Opus by inspecting databases before writing SQL

MarkTechPost · Jul 19, 2026 · 3 min read

Most text-to-SQL systems treat the task like translation: turn a question into a query and hope for the best. The probl...

Why your next RL agent might need a safety net — Inblix summary
Product

Why your next RL agent might need a safety net

OpenAI Blog · Jul 19, 2026 · 2 min read

The era of training reinforcement learning agents entirely in comfortable, consequence-free simulators might be coming...

Kimi K3 tops frontend coding but flunks advanced math — Inblix summary
AI News

Kimi K3 tops frontend coding but flunks advanced math

The Decoder · Jul 19, 2026 · 2 min read

Moonshot's Kimi K3 just pulled off something no Chinese model has managed before: it grabbed the number one spot on the...

OpenAI's 16-Game Gauntlet Exposes RL's Memorization Problem — Inblix summary
Product

OpenAI's 16-Game Gauntlet Exposes RL's Memorization Problem

OpenAI Blog · Jul 19, 2026 · 2 min read

OpenAI just dropped a truth bomb wrapped in 16 pixelated environments. Their new Procgen Benchmark isn't just another s...

OpenAI's Codex solves 70% of coding puzzles by guessing 100 times — Inblix summary
Product

OpenAI's Codex solves 70% of coding puzzles by guessing 100 times

OpenAI Blog · Jul 19, 2026 · 2 min read

OpenAI has formally introduced Codex, the code-generating descendant of GPT-3 that's been powering GitHub Copilot behin...

Bigger AI models lie more, new TruthfulQA benchmark finds — Inblix summary
Product

Bigger AI models lie more, new TruthfulQA benchmark finds

OpenAI Blog · Jul 18, 2026 · 2 min read

Here's a finding that should give anyone deploying large language models serious pause: the bigger the model, the more...

Fill-in-the-middle training doesn't hurt, it helps: new defaults for LLMs — Inblix summary
Product

Fill-in-the-middle training doesn't hurt, it helps: new defaults for LLMs

OpenAI Blog · Jul 18, 2026 · 2 min read

The classic left-to-right way of training language models might finally have a permanent co-pilot. A new paper from res...

OpenAI's o1 model beats PhDs and cracks top 500 in Math Olympiad — Inblix summary
Product

OpenAI's o1 model beats PhDs and cracks top 500 in Math Olympiad

OpenAI Blog · Jul 16, 2026 · 2 min read

OpenAI just dropped an early version of a new model, o1-preview, that doesn't just incrementally improve on GPT-4o — it...

Murati's Thinking Machines drops Inkling: US best, but trails China — Inblix summary
AI News

Murati's Thinking Machines drops Inkling: US best, but trails China

The Decoder · Jul 16, 2026 · 3 min read

Thinking Machines Lab, the startup Mira Murati built after leaving OpenAI, just shipped its first model — and it's a st...

OpenAI Finds SWE-bench Is Broken—So They're Fixing It — Inblix summary
Product

OpenAI Finds SWE-bench Is Broken—So They're Fixing It

OpenAI Blog · Jul 16, 2026 · 2 min read

SWE-bench has been the gold standard for proving your AI can code like a real engineer. Top models barely scraped 20% o...

OpenAI drops o1-preview: A PhD-level reasoner that thinks first — Inblix summary
Product

OpenAI drops o1-preview: A PhD-level reasoner that thinks first

OpenAI Blog · Jul 16, 2026 · 2 min read

OpenAI isn't calling it GPT-5. They're resetting the counter to 1. The company just released o1-preview, the first in a...

OpenAI's SimpleQA Exposes GPT-4o's 40% Factuality Score — Inblix summary
Product

OpenAI's SimpleQA Exposes GPT-4o's 40% Factuality Score

OpenAI Blog · Jul 15, 2026 · 2 min read

OpenAI just open-sourced a new factuality benchmark, and the results are humbling. The dataset, called SimpleQA, is des...

OpenAI's CUA sets benchmark records but still fumbles 62% of computer tasks — Inblix summary
Product

OpenAI's CUA sets benchmark records but still fumbles 62% of computer tasks

OpenAI Blog · Jul 15, 2026 · 2 min read

OpenAI pulled back the curtain on Computer-Using Agent, the model powering its new Operator research preview. CUA ditch...

OpenAI's SWE-Lancer tests if AI can actually earn $1M on Upwork — Inblix summary
Product

OpenAI's SWE-Lancer tests if AI can actually earn $1M on Upwork

OpenAI Blog · Jul 15, 2026 · 2 min read

The question isn't whether AI can code anymore. It's whether it can get paid for it. OpenAI's new SWE-Lancer benchmark...

OpenAI's New Pioneers Program Aims to Fix AI's Benchmark Obsession — Inblix summary
Product

OpenAI's New Pioneers Program Aims to Fix AI's Benchmark Obsession

OpenAI Blog · Jul 14, 2026 · 2 min read

OpenAI just announced the Pioneers Program, a hands-on initiative that pairs its research teams directly with startups...

OpenAI's o3 and o4-mini can now 'think' with images, not just see them — Inblix summary
Product

OpenAI's o3 and o4-mini can now 'think' with images, not just see them

OpenAI Blog · Jul 14, 2026 · 3 min read

OpenAI just dropped a pair of new models—o3 and o4-mini—and for the first time, the chain-of-thought process includes a...

OpenAI drops o3 and o4-mini: Reasoning models that actually use tools — Inblix summary
Product

OpenAI drops o3 and o4-mini: Reasoning models that actually use tools

OpenAI Blog · Jul 14, 2026 · 2 min read

OpenAI isn't just releasing new models. With o3 and o4-mini, they're fundamentally changing how their AI reasons about...

OpenAI Says the Way We Score AI Actually Rewards Lying — Inblix summary
Product

OpenAI Says the Way We Score AI Actually Rewards Lying

OpenAI Blog · Jul 13, 2026 · 3 min read

OpenAI has a new theory for why chatbots keep making stuff up, and it points the finger squarely at the people building...

OpenAI admits its own benchmarks are useless for 80% of the world — Inblix summary
Product

OpenAI admits its own benchmarks are useless for 80% of the world

OpenAI Blog · Jul 12, 2026 · 2 min read

OpenAI just admitted something that anyone working in multilingual AI has known for years: the benchmarks we use to mea...

Claude Code claims 97% build success on Java migrations — but 24% of those apps were quietly broken — Inblix summary
Research

Claude Code claims 97% build success on Java migrations — but 24% of those apps were quietly broken

Hugging Face Blog · Jun 30, 2026 · 2 min read

AI-assisted enterprise modernization just hit a reality check. ScarfBench, a new open benchmark designed to stress-test...

229,000 benchmark scores now link directly to Hugging Face model cards — Inblix summary
Research

229,000 benchmark scores now link directly to Hugging Face model cards

Hugging Face Blog · Jun 30, 2026 · 2 min read

The mess of AI evaluation just got a little easier to navigate. The EvalEval Coalition's EEE project, launched in Febru...

ASR models fail hard in real rooms — Treble and Hugging Face just proved it — Inblix summary
Research

ASR models fail hard in real rooms — Treble and Hugging Face just proved it

Hugging Face Blog · Jun 24, 2026 · 2 min read

The assumption that a speech recognition model that aces a clean benchmark will work in your kitchen or car is, frankly...

GLM-5.2 crushes open-source coding leaders, lands within 4 points of Claude Opus 4.8 — Inblix summary
Research

GLM-5.2 crushes open-source coding leaders, lands within 4 points of Claude Opus 4.8

Hugging Face Blog · Jun 17, 2026 · 3 min read

The open-source model race just got a lot more interesting. The team behind GLM has dropped GLM-5.2, and it's not just...

A 0.6B model just beat SAM 3 on grounding—here's the simple trick that made it work — Inblix summary
Research

A 0.6B model just beat SAM 3 on grounding—here's the simple trick that made it work

Hugging Face Blog · Apr 1, 2026 · 2 min read

The team behind Falcon Perception asked a question most of the field has been dodging: why are vision systems still sti...

Meta and Hugging Face’s OpenEnv proves AI agents still can’t handle a simple calendar — Inblix summary
Research

Meta and Hugging Face’s OpenEnv proves AI agents still can’t handle a simple calendar

Hugging Face Blog · Feb 12, 2026 · 2 min read

Meta and Hugging Face have released OpenEnv, an open-source framework that hooks AI agents up to real-world tools inste...

Hugging Face kills black-box leaderboards with decentralized evals — Inblix summary
Research

Hugging Face kills black-box leaderboards with decentralized evals

Hugging Face Blog · Feb 4, 2026 · 2 min read

Hugging Face just took a swing at one of AI’s dirtiest open secrets: benchmark scores you can’t trust. The platform ann...

H Company's Holo2 nails 78.5% on tough UI benchmark by looking twice — Inblix summary
Research

H Company's Holo2 nails 78.5% on tough UI benchmark by looking twice

Hugging Face Blog · Feb 3, 2026 · 2 min read

H Company just dropped a preview of its largest UI localization model yet, and the numbers are worth paying attention t...

NVIDIA's Conformer-LLM hybrid tops ASR accuracy, but open-source still trails in long-form audio — Inblix summary
Research

NVIDIA's Conformer-LLM hybrid tops ASR accuracy, but open-source still trails in long-form audio

Hugging Face Blog · Nov 21, 2025 · 2 min read

The Open ASR Leaderboard just got a major expansion, adding multilingual and long-form transcription tracks that reveal...

14,000 coding battles reveal o3-mini rules the BigCodeArena leaderboard — Inblix summary
Research

14,000 coding battles reveal o3-mini rules the BigCodeArena leaderboard

Hugging Face Blog · Oct 7, 2025 · 2 min read

Forget reading code and guessing if it works. BigCodeArena, a new platform launched in February, has already logged ove...

Cohere's new RTEB benchmark catches models faking their retrieval scores — Inblix summary
Research

Cohere's new RTEB benchmark catches models faking their retrieval scores

Hugging Face Blog · Oct 1, 2025 · 2 min read

Cohere just dropped a retrieval benchmark that's designed to embarrass models that have been gaming the system. The Ret...

Meta's GAIA 2.0 benchmark humbles GPT-5 and Grok-4 with noisy, time-sensitive agent tasks — Inblix summary
Research

Meta's GAIA 2.0 benchmark humbles GPT-5 and Grok-4 with noisy, time-sensitive agent tasks

Hugging Face Blog · Sep 22, 2025 · 2 min read

The original GAIA benchmark is practically cooked. Two years after its 2023 release, the top models are acing the harde...

Ettin benchmarks prove encoder models crush decoders at 4X the efficiency — Inblix summary
Research

Ettin benchmarks prove encoder models crush decoders at 4X the efficiency

Hugging Face Blog · Jul 16, 2025 · 2 min read

For years, the industry has guessed whether bidirectional encoders or causal decoders were better for non-generative ta...

NeurIPS 2025 bets $16K that benchmarks, not loss curves, can judge baby LLMs — Inblix summary
Research

NeurIPS 2025 bets $16K that benchmarks, not loss curves, can judge baby LLMs

Hugging Face Blog · Jul 4, 2025 · 2 min read

The blank stare of a loss curve is a familiar frustration for anyone who has trained a large language model. In the ear...

Hugging Face's ScreenSuite stress-tests 5 AI models on pure vision — no DOM cheating — Inblix summary
Research

Hugging Face's ScreenSuite stress-tests 5 AI models on pure vision — no DOM cheating

Hugging Face Blog · Jun 6, 2025 · 2 min read

Evaluating AI that can actually use a computer like a person — by looking at the screen — remains surprisingly hard to...

H Company drops open-source Holo1 VLMs that slash web automation costs to $0.13 per task — Inblix summary
Research

H Company drops open-source Holo1 VLMs that slash web automation costs to $0.13 per task

Hugging Face Blog · Jun 3, 2025 · 2 min read

H Company just released Holo1, an open-source family of action VLMs that makes web automation dramatically cheaper with...

DeepSeek drops V3-0324 with MIT license and a 19.8-point math leap — Inblix summary
Research

DeepSeek drops V3-0324 with MIT license and a 19.8-point math leap

Hugging Face Blog · Mar 26, 2025 · 2 min read

The newest model from DeepSeek isn't a flashy launch. It's a silent update that speaks volumes. DeepSeek-V3-0324, an up...

OlympicCoder 7B, a tiny open model, just beat GPT-4o and Claude 3.7 at live coding — Inblix summary
Research

OlympicCoder 7B, a tiny open model, just beat GPT-4o and Claude 3.7 at live coding

Hugging Face Blog · Mar 20, 2025 · 2 min read

The open-source model OlympicCoder 7B is now topping the LiveCodeBench leaderboard, outperforming the daily-driver codi...

Gemma 3's 27B model punches into top 10 of Chatbot Arena, beating Gemini 1.5 Pro — Inblix summary
Research

Gemma 3's 27B model punches into top 10 of Chatbot Arena, beating Gemini 1.5 Pro

Hugging Face Blog · Mar 12, 2025 · 2 min read

Google just dropped Gemma 3, and the numbers are genuinely surprising. The new open-weight model family runs from a tin...

A 32B model just crushed Claude 3.7 Sonnet on Olympiad coding problems — Inblix summary
Research

A 32B model just crushed Claude 3.7 Sonnet on Olympiad coding problems

Hugging Face Blog · Mar 11, 2025 · 2 min read

The Open R1 project just dropped a genuinely surprising result: a fine-tuned 32-billion-parameter model that out-codes...

Cohere's Aya Vision beats models twice its size in 23-language image tests — Inblix summary
Research

Cohere's Aya Vision beats models twice its size in 23-language image tests

Hugging Face Blog · Mar 4, 2025 · 2 min read

Cohere For AI just dropped a pair of open-weight vision-language models that punch well above their weight class — and...

A 3-line code fix just tripled scores for DeepSeek and doubled Qwen on the math leaderboard — Inblix summary
Research

A 3-line code fix just tripled scores for DeepSeek and doubled Qwen on the math leaderboard

Hugging Face Blog · Feb 14, 2025 · 2 min read

Here's something you don't see every day: a three-line code change completely reshuffled the top 20 rankings on the mos...

AI can't do your data analyst job yet: Hugging Face benchmark scores just 16% — Inblix summary
Research

AI can't do your data analyst job yet: Hugging Face benchmark scores just 16%

Hugging Face Blog · Feb 4, 2025 · 2 min read

If you're a data analyst waiting for AI to take the boring parts of your job off your plate, you're going to be waiting...

Hugging Face tries to clone OpenAI’s Deep Research in 24 hours with open code agent — Inblix summary
Research

Hugging Face tries to clone OpenAI’s Deep Research in 24 hours with open code agent

Hugging Face Blog · Feb 4, 2025 · 2 min read

OpenAI’s new Deep Research tool is undeniably slick—it browses the web, synthesizes multi-step answers, and just scored...

GPT-4o's Reasoning Plummets 26 Points When Listening Instead of Reading — Inblix summary
Research

GPT-4o's Reasoning Plummets 26 Points When Listening Instead of Reading

Hugging Face Blog · Dec 20, 2024 · 2 min read

Give a state-of-the-art model a logic puzzle in writing, and it aces it. Read the same puzzle aloud, and that performan...

Falcon 3's 10B model beats everything under 13B parameters — here's how — Inblix summary
Research

Falcon 3's 10B model beats everything under 13B parameters — here's how

Hugging Face Blog · Dec 17, 2024 · 2 min read

The Technology Innovation Institute just dropped the Falcon 3 family, and the headliner is a 10-billion-parameter model...

Japan's LLMs get their first real report card with 20+ open benchmarks — Inblix summary
Research

Japan's LLMs get their first real report card with 20+ open benchmarks

Hugging Face Blog · Nov 20, 2024 · 2 min read

Evaluating Japanese large language models has been a mess. The language's unique mix of kanji, hiragana, katakana, and...

Judge Arena lets you pick the best AI evaluator by voting in blind, head-to-head battles — Inblix summary
Research

Judge Arena lets you pick the best AI evaluator by voting in blind, head-to-head battles

Hugging Face Blog · Nov 19, 2024 · 2 min read

Atla just dropped Judge Arena, a platform that flips the script on how we evaluate AI evaluators. Instead of relying so...

CinePile 2.0 uses adversarial AI to fix broken video questions instead of trashing them — Inblix summary
Research

CinePile 2.0 uses adversarial AI to fix broken video questions instead of trashing them

Hugging Face Blog · Oct 23, 2024 · 2 min read

The original CinePile dataset launched in May 2024 with a genuinely impressive stat: humans outperformed the best comme...

BenCzechMark exposes 25 LLMs: most still mangle Czech grammar and culture — Inblix summary
Research

BenCzechMark exposes 25 LLMs: most still mangle Czech grammar and culture

Hugging Face Blog · Oct 1, 2024 · 2 min read

The best open-source large language models still stumble badly over Czech, and now we have the receipts. A new evaluati...