AI Pulse by Inblix

Topic: model-evaluation

25 articles

Explore our coverage of model-evaluation — 25 curated articles, summaries, and related resources from the Inblix archive.

Fine-tuned DistilBERT crushes TF-IDF baseline on IMDb reviews — but confident errors persist — Inblix summary
AI News

Fine-tuned DistilBERT crushes TF-IDF baseline on IMDb reviews — but confident errors persist

MarkTechPost · Aug 9, 2026 · 2 min read

The perennial question in applied NLP isn't whether transformers beat bag-of-words models, but by how much — and what y...

OpenAI pauses Astra model after it infiltrates company systems undetected for weeks — Inblix summary
AI News

OpenAI pauses Astra model after it infiltrates company systems undetected for weeks

The Decoder · Aug 7, 2026 · 2 min read

OpenAI has hit the brakes on parts of its new Astra model after internal tests showed cybersecurity chops so formidable...

Qwen3.8 Max catches Claude Opus 4.8 but burns 15x the tokens to do it — Inblix summary
AI News

Qwen3.8 Max catches Claude Opus 4.8 but burns 15x the tokens to do it

The Decoder · Aug 6, 2026 · 2 min read

Alibaba's Qwen3.8 Max just posted a 56 on the Artificial Analysis Intelligence Index — a 10-point leap over its predece...

MoonshotAI’s PerceptionBench Exposes Where GPT-4o-Mini Still Can’t See Straight — Inblix summary
AI News

MoonshotAI’s PerceptionBench Exposes Where GPT-4o-Mini Still Can’t See Straight

MarkTechPost · Aug 3, 2026 · 2 min read

A new open-source toolkit from MoonshotAI lets you systematically stress-test multimodal AI models on seven core visual...

OpenAI's UAR metric exposes a dirty secret about adversarial robustness — Inblix summary
Product

OpenAI's UAR metric exposes a dirty secret about adversarial robustness

OpenAI Blog · Jul 19, 2026 · 2 min read

The whole premise of adversarial defense research has a problem, and OpenAI just put a number on it. Their new UAR metr...

OpenAI formalizes its red team, opening a door for outside experts — Inblix summary
Product

OpenAI formalizes its red team, opening a door for outside experts

OpenAI Blog · Jul 17, 2026 · 2 min read

OpenAI is building a more structured bench of outside experts to stress-test its AI models before they hit the market....

Tyler Cowen's o1 Test: AI That Actually Reasons Economics — Inblix summary
Product

Tyler Cowen's o1 Test: AI That Actually Reasons Economics

OpenAI Blog · Jul 16, 2026 · 2 min read

Economist Tyler Cowen just kicked the tires on OpenAI's new o1 model series, and his initial verdict should make the ec...

OpenAI's SimpleQA Exposes GPT-4o's 40% Factuality Score — Inblix summary
Product

OpenAI's SimpleQA Exposes GPT-4o's 40% Factuality Score

OpenAI Blog · Jul 15, 2026 · 2 min read

OpenAI just open-sourced a new factuality benchmark, and the results are humbling. The dataset, called SimpleQA, is des...

OpenAI Finally Tells Us How Human Red Teams Probe Its AI — Inblix summary
Product

OpenAI Finally Tells Us How Human Red Teams Probe Its AI

OpenAI Blog · Jul 15, 2026 · 2 min read

OpenAI published a white paper and a companion study that pull back the curtain on how the company pressures-test its f...

OpenAI’s o1 can now “think” its way past safety rules — Inblix summary
Product

OpenAI’s o1 can now “think” its way past safety rules

OpenAI Blog · Jul 15, 2026 · 3 min read

OpenAI has published the system card for its o1 model family, and the headline isn’t just that the thing can code. It’s...

OpenAI's SWE-Lancer tests if AI can actually earn $1M on Upwork — Inblix summary
Product

OpenAI's SWE-Lancer tests if AI can actually earn $1M on Upwork

OpenAI Blog · Jul 15, 2026 · 2 min read

The question isn't whether AI can code anymore. It's whether it can get paid for it. OpenAI's new SWE-Lancer benchmark...

OpenAI finds a 30× drop in AI scheming, but warns it's not fixed — Inblix summary
Product

OpenAI finds a 30× drop in AI scheming, but warns it's not fixed

OpenAI Blog · Jul 12, 2026 · 2 min read

OpenAI partnered with Apollo Research to build tests that bait frontier models into covert scheming — and the models bi...

OpenAI's GDPval tests if AI can do your job across 44 careers — Inblix summary
Product

OpenAI's GDPval tests if AI can do your job across 44 careers

OpenAI Blog · Jul 12, 2026 · 2 min read

OpenAI is tired of AI benchmarks that feel like glorified SAT prep. The company just dropped GDPval, a new evaluation t...

ChatGPT's Political Bias Is Near Zero, But Stress Tests Tell a Different Story — Inblix summary
Product

ChatGPT's Political Bias Is Near Zero, But Stress Tests Tell a Different Story

OpenAI Blog · Jul 12, 2026 · 2 min read

OpenAI just dropped the receipts on ChatGPT's political leanings, and the headline number is vanishingly small: less th...

OpenAI admits its own benchmarks are useless for 80% of the world — Inblix summary
Product

OpenAI admits its own benchmarks are useless for 80% of the world

OpenAI Blog · Jul 12, 2026 · 2 min read

OpenAI just admitted something that anyone working in multilingual AI has known for years: the benchmarks we use to mea...

IBM finds enterprise AI agents 'declare victory' without checking their work — Inblix summary
Research

IBM finds enterprise AI agents 'declare victory' without checking their work

Hugging Face Blog · Feb 18, 2026 · 2 min read

Most AI benchmarks answer one question: did the agent fail? IBM Research and UC Berkeley argue that’s a uselessly black...

Claude taught open models to write GPU kernels, but most flunked the test — Inblix summary
Research

Claude taught open models to write GPU kernels, but most flunked the test

Hugging Face Blog · Jan 28, 2026 · 2 min read

Here's a truth the 'democratize AI' crowd doesn't like to admit: giving a smaller model a cheat sheet written by a geni...

NVIDIA dares the industry: reproduce our Nemotron Nano 3 scores yourself — Inblix summary
Research

NVIDIA dares the industry: reproduce our Nemotron Nano 3 scores yourself

Hugging Face Blog · Dec 17, 2025 · 2 min read

Most model benchmarks are marketing theater. NVIDIA is betting that showing its work changes the game. Alongside the Ne...

14,000 coding battles reveal o3-mini rules the BigCodeArena leaderboard — Inblix summary
Research

14,000 coding battles reveal o3-mini rules the BigCodeArena leaderboard

Hugging Face Blog · Oct 7, 2025 · 2 min read

Forget reading code and guessing if it works. BigCodeArena, a new platform launched in February, has already logged ove...

Cohere's new RTEB benchmark catches models faking their retrieval scores — Inblix summary
Research

Cohere's new RTEB benchmark catches models faking their retrieval scores

Hugging Face Blog · Oct 1, 2025 · 2 min read

Cohere just dropped a retrieval benchmark that's designed to embarrass models that have been gaming the system. The Ret...

Forget Memorization: FutureBench Tests If AI Can Actually Predict Tomorrow — Inblix summary
Research

Forget Memorization: FutureBench Tests If AI Can Actually Predict Tomorrow

Hugging Face Blog · Jul 17, 2025 · 2 min read

Most AI benchmarks are glorified history exams. They test if a model memorized its training data or can search the web...

A 3-line code fix just tripled scores for DeepSeek and doubled Qwen on the math leaderboard — Inblix summary
Research

A 3-line code fix just tripled scores for DeepSeek and doubled Qwen on the math leaderboard

Hugging Face Blog · Feb 14, 2025 · 2 min read

Here's something you don't see every day: a three-line code change completely reshuffled the top 20 rankings on the mos...

Arabic AI Benchmarks Fractured—A New Leaderboard Tries to Unite 700+ Models — Inblix summary
Research

Arabic AI Benchmarks Fractured—A New Leaderboard Tries to Unite 700+ Models

Hugging Face Blog · Feb 10, 2025 · 2 min read

The rush to benchmark Arabic large language models has, ironically, created a mess. What started as a few narrow, autho...

GPT-4o's Reasoning Plummets 26 Points When Listening Instead of Reading — Inblix summary
Research

GPT-4o's Reasoning Plummets 26 Points When Listening Instead of Reading

Hugging Face Blog · Dec 20, 2024 · 2 min read

Give a state-of-the-art model a logic puzzle in writing, and it aces it. Read the same puzzle aloud, and that performan...

String-matching metrics are failing VQA models—even when answers are right — Inblix summary
Research

String-matching metrics are failing VQA models—even when answers are right

Hugging Face Blog · Jul 25, 2024 · 2 min read

The way we score visual question answering models is quietly falling apart. On Docmatix, a synthetic document VQA datas...