AI Pulse by Inblix

Stop counting tokens: how to track what AI actually gets done

OpenAI Blog · Jul 14, 2026 · 2 min read · Read original article →

Curated by the Inblix editorial team


Featured image for article: Stop counting tokens: how to track what AI actually gets done

OpenAI’s latest models are undeniably cheaper. From GPT-4 to GPT-5.4, the price per million tokens cratered by 97%, and GPT-5.6 now uses 54% fewer output tokens while finishing tasks 57% faster. A good story, but token price is a dangerous metric for leaders to obsess over. It tells you nothing about whether the work produced is actually useful. The smarter play, and the thrust of OpenAI’s latest enterprise guidance, is to start measuring ‘useful work per dollar.’ That means tracking completed tasks, time saved, and decisions improved, not just credits consumed.

As teams shift from simple chat to long-running, multi-step ‘agentic’ workflows, the visibility problem gets worse. A rising bill could signal critical adoption or costly thrashing. OpenAI’s updated admin console aims to solve this by letting leaders slice usage data by user, product, and model to spot the difference between a power user’s critical workflow and someone burning credits on dead ends. The company’s advice is blunt: the cheapest model rarely wins on total cost if it fails and retries constantly. A pricier, more capable model can be cheaper if it nails the task on the first try with less human review.

This focus on outcomes over inputs is the core of OpenAI’s five-step investment playbook. It pushes for evaluating models on real tasks with real edge cases, then tracking ‘cost per accepted outcome’—like a resolved support ticket or a code change that passes review. The guidance also stresses that model choice is only one lever. Tighter instructions, reusable context, and explicit stopping conditions in a workflow can often cut waste more effectively than simply swapping models. The real challenge is governance, not just selecting GPT-5.6 over GPT-5.4.

That governance layer is what separates a science project from a scalable deployment. OpenAI is telling leaders to define precisely what context and tools an AI can access, what actions it’s allowed to take, and who approves riskier steps before they happen. This becomes non-negotiable when connecting AI to enterprise systems through plugins or ‘Computer Use’ capabilities. The end goal is a portfolio approach: broad access for everyday productivity, targeted workflows for specific functions, and a few strategic bets built on proprietary data. It’s a framework that sounds sensible, but the hard part—defining what ‘good enough’ looks like for your specific business—is still the work no vendor can do for you.

💡 Key Takeaways

  1. The 97% drop in token costs from GPT-4 to GPT-5.4 is a distraction; leaders should instead track 'cost per accepted outcome' to measure real AI value.
  2. A cheaper model can inflate total costs through failures and retries, making a pricier but more reliable model the more economical choice for complex tasks.
  3. Governance over which tools AI can access and what actions it can take is the operational layer that determines whether a workflow can scale safely.

Keep reading: See related articles below for more coverage on this topic.

Get smarter about AI

The sharpest AI news, curated daily. Delivered free to your inbox.

Learn more

← Back to all articles