SafetyKit hits 95% accuracy on 16B daily tokens using GPT-5
Curated by the Inblix editorial team
SafetyKit, which builds multimodal AI agents for fraud and policy enforcement, is now processing 16 billion tokens daily—a staggering leap from 200 million just six months ago. The company credits the scale-up to an aggressive model-matching strategy that routes every piece of content to the specific OpenAI model best suited for the job. Some tasks require the deep reasoning of GPT-5, while high-volume workflows lean on GPT-4.1 for speed and reliability.
The approach is producing measurable results. SafetyKit reports a review accuracy exceeding 95% across 100% of customer content, a benchmark that covers everything from scam images with embedded phone numbers to missing legal disclaimers on product pages. When OpenAI dropped the o3 model, SafetyKit deployed it the same day to boost edge-case performance. GPT-5 followed shortly after, delivering a more than 10-point gain on the company’s hardest vision tasks within days of release.
CEO Ryan Graunke frames this as a design philosophy rather than a model loyalty play. “We think of our agents as purpose-built workflows,” he explains. “Some tasks require deep reasoning, others need multimodal context. OpenAI is the only stack that delivers reliable performance across both.” A scam detection agent might use GPT-4.1 to parse a QR code in a product image, then hand off the compliance call to GPT-5. The goal is to catch what legacy keyword-trigger systems miss, like whether a wellness product’s listing includes a region-specific disclaimer that’s actually required by law.
What’s notable here isn’t just the accuracy numbers—it’s the cadence. SafetyKit has engineered its infrastructure to absorb new model releases as instant product upgrades rather than multi-month integration projects. That operational tempo, paired with tools like reinforcement fine-tuning and Computer Using Agent for automating manual reviews, suggests the real moat isn’t any single model. It’s the system designed to swap them out the moment something better arrives.
💡 Key Takeaways
- SafetyKit scaled from processing 200M to 16B tokens daily in six months by routing specific tasks to different OpenAI models rather than using one model for everything.
- Deploying GPT-5 immediately after launch improved benchmark scores by more than 10 points on the company’s most difficult vision-based policy tasks.
- Purpose-built agent workflows allow SafetyKit to catch nuanced violations—like missing region-specific legal disclaimers—that legacy keyword-based systems routinely miss.
Keep reading: See related articles below for more coverage on this topic.
Get smarter about AI
The sharpest AI news, curated daily. Delivered free to your inbox.