AI Pulse by Inblix

Topic: computer-vision

17 articles

Explore our coverage of computer-vision — 17 curated articles, summaries, and related resources from the Inblix archive.

PSNR and SSIM reward blurry erasures — Xiaomi's PROVE metric finally catches it — Inblix summary
AI News

PSNR and SSIM reward blurry erasures — Xiaomi's PROVE metric finally catches it

MarkTechPost · Aug 12, 2026 · 3 min read

The metrics we've been using to judge AI object removal are quietly misleading us. A team from MiLM Plus at Xiaomi drop...

ByteDance's SeedRealtime AI interrupted a bad espresso shot — and that changes everything — Inblix summary
AI News

ByteDance's SeedRealtime AI interrupted a bad espresso shot — and that changes everything

MarkTechPost · Aug 10, 2026 · 2 min read

ByteDance's Seed team just showed what happens when you give an AI permission to speak first. Their new model, SeedReal...

PixelRAG ditches HTML parsing, retrieves documents as screenshots with 7-query Recall@k eval — Inblix summary
AI News

PixelRAG ditches HTML parsing, retrieves documents as screenshots with 7-query Recall@k eval

MarkTechPost · Aug 4, 2026 · 2 min read

Most retrieval pipelines treat every web page as a bag of parsed text, but that assumption breaks the moment you encoun...

LingBot-Map Demo: 3D Scene Reconstruction Tunes Itself to Your GPU's VRAM — Inblix summary
AI News

LingBot-Map Demo: 3D Scene Reconstruction Tunes Itself to Your GPU's VRAM

MarkTechPost · Jul 31, 2026 · 2 min read

A new end-to-end streaming 3D reconstruction pipeline, LingBot-Map, is making waves for a clever bit of engineering tha...

A single doctored photo can now fool AI from any angle — Inblix summary
Product

A single doctored photo can now fool AI from any angle

OpenAI Blog · Jul 20, 2026 · 2 min read

Remember last week's reassuring claim—that self-driving cars couldn't easily be tricked by altered street signs because...

DeepMind: Video generators are the world models CV has been waiting for — Inblix summary
AI News

DeepMind: Video generators are the world models CV has been waiting for

The Decoder · Jul 19, 2026 · 2 min read

Google DeepMind just dropped a paper arguing that the universal training task computer vision has been missing was hidi...

OpenAI Puts Neural Networks Under a Microscope — Inblix summary
Product

OpenAI Puts Neural Networks Under a Microscope

OpenAI Blog · Jul 19, 2026 · 2 min read

OpenAI has released a new tool called Microscope, a collection of visualizations that peers into the inner workings of...

GPT-2 trained on pixels rivals top CNNs without labels — Inblix summary
Product

GPT-2 trained on pixels rivals top CNNs without labels

OpenAI Blog · Jul 19, 2026 · 2 min read

Here's a sentence I didn't expect to write: a raw GPT-2 language model, fed nothing but pixel sequences and told to pre...

CLIP drops the training wheels: Zero-shot vision matches ResNet-50 — Inblix summary
Product

CLIP drops the training wheels: Zero-shot vision matches ResNet-50

OpenAI Blog · Jul 19, 2026 · 2 min read

What if a vision model never saw a single labeled example from ImageNet, yet matched the accuracy of a fully-supervised...

OpenAI finds CLIP has 'Halle Berry' neurons like the human brain — Inblix summary
Product

OpenAI finds CLIP has 'Halle Berry' neurons like the human brain

OpenAI Blog · Jul 19, 2026 · 2 min read

OpenAI researchers have discovered that CLIP, their vision model matching ResNet-50 performance, contains multimodal ne...

OpenAI lets you fine-tune GPT-4o with images, and yes, it's a big deal — Inblix summary
Product

OpenAI lets you fine-tune GPT-4o with images, and yes, it's a big deal

OpenAI Blog · Jul 15, 2026 · 2 min read

OpenAI finally plugged the glaring hole in its fine-tuning API. Starting today, you can feed images—not just text—into...

Grab fine-tunes GPT-4o vision on just 100 images, maps SE Asia — Inblix summary
Product

Grab fine-tunes GPT-4o vision on just 100 images, maps SE Asia

OpenAI Blog · Jul 15, 2026 · 2 min read

Mapping Southeast Asia has always been a nightmare for conventional providers. The streets are narrow, the signage is c...

A 0.6B model just beat SAM 3 on grounding—here's the simple trick that made it work — Inblix summary
Research

A 0.6B model just beat SAM 3 on grounding—here's the simple trick that made it work

Hugging Face Blog · Apr 1, 2026 · 2 min read

The team behind Falcon Perception asked a question most of the field has been dodging: why are vision systems still sti...

Hugging Face trains a 2.2B model to see GUIs, hitting 65% on ScreenSpot — Inblix summary
Research

Hugging Face trains a 2.2B model to see GUIs, hitting 65% on ScreenSpot

Hugging Face Blog · Sep 23, 2025 · 2 min read

Turning a vision-language model with zero grounding ability into a GUI agent that can click and type is a monumental da...

Google's SigLIP 2 adds a decoder and self-distillation to build vision encoders that actually know where things are — Inblix summary
Research

Google's SigLIP 2 adds a decoder and self-distillation to build vision encoders that actually know where things are

Hugging Face Blog · Feb 21, 2025 · 2 min read

Google just released SigLIP 2, a family of multilingual vision-language encoders that makes the original SigLIP look li...

Hugging Face bridges its entire pipeline to timm's 200K daily users via TimmWrapper — Inblix summary
Research

Hugging Face bridges its entire pipeline to timm's 200K daily users via TimmWrapper

Hugging Face Blog · Jan 16, 2025 · 2 min read

The wall between two of PyTorch's most popular libraries just came down. A new integration called TimmWrapper lets you...

InstantMesh now converts raw 3D meshes into textured models in 10 seconds — Inblix summary
Research

InstantMesh now converts raw 3D meshes into textured models in 10 seconds

Hugging Face Blog · Sep 30, 2024 · 2 min read

Generative 3D tools like InstantMesh have gotten remarkably fast at creating geometry from a single image, but the outp...