AI Pulse by Inblix

OpenAI lets you fine-tune GPT-4o with images, and yes, it's a big deal

OpenAI Blog · Jul 15, 2026 · 2 min read · Read original article →

Curated by the Inblix editorial team


Featured image for article: OpenAI lets you fine-tune GPT-4o with images, and yes, it's a big deal

OpenAI finally plugged the glaring hole in its fine-tuning API. Starting today, you can feed images—not just text—into GPT-4o’s custom training runs. For the hundreds of thousands of developers already tweaking models for text-only tasks, this opens a door they’ve been rattling for months. The process mirrors the existing text workflow. Prepare a dataset with the proper formatting, upload it, and you’re off. The real kicker? You can see meaningful performance gains with as few as 100 images.

Some early partner results are genuinely eyebrow-raising. Grab, the Southeast Asian rideshare and delivery giant, used vision fine-tuning to automate mapping data extraction from street-level imagery. With just 100 examples, they boosted lane count accuracy by 20% and speed limit sign localization by 13% over the base model. Enterprise automation company Automat saw an even more dramatic swing: their RPA agent’s success rate at locating UI elements from screenshots jumped from a paltry 16.6% to 61.67%. That’s a 272% uplift, the kind of number that makes you sit up and pay attention. Coframe, which builds AI tools for generating website variations, used image-and-code fine-tuning to improve visual consistency in generated websites by 26%.

It’s not just about throwing more data at the problem. These performance leaps come from relatively tiny datasets, which upends the assumption that you need massive labeled image libraries to move the needle. This matters for startups and mid-size companies that can’t afford annotation sweatshops. The pricing structure is straightforward: image inputs get tokenized based on size and then billed at the same per-token rate as text. There’s a promotional window—1 million training tokens per day free until October 31, 2024—after which training runs $25 per million tokens, with inference at $3.75 per million input tokens and $15 per million output tokens.

The deprecation notice at the top of the announcement deserves a moment of silence. OpenAI is winding down the fine-tuning platform entirely for new users as of May 2026. Existing users get a grace period to keep creating training jobs, and fine-tuned models survive until their base models are deprecated. It’s a strange bit of forward-looking honesty tucked into a product launch—celebrating a new capability while quietly signaling its eventual retirement. For now, though, the tool is live on the latest GPT-4o snapshot (gpt-4o-2024-08-06) and available to anyone on a paid usage tier.

💡 Key Takeaways

  1. Vision fine-tuning on GPT-4o delivers dramatic performance improvements with surprisingly small datasets—Grab needed only 100 images to boost lane count accuracy by 20%.
  2. Automat's RPA agent saw a 272% performance uplift on UI element detection, moving from 16.6% to 61.67% success rate after fine-tuning on screenshots.
  3. OpenAI is simultaneously launching this feature and announcing the entire fine-tuning platform will close to new users by May 2026, so the clock is ticking.

Keep reading: See related articles below for more coverage on this topic.

Get smarter about AI

The sharpest AI news, curated daily. Delivered free to your inbox.

Learn more

Glossary terms

← Back to all articles