Stripe finds GPT-4 outperforms humans at classifying businesses
Curated by the Inblix editorial team
Stripe took an unusual approach to generative AI: they told 100 employees to stop their regular work and just experiment with GPT-4. The result was 50 potential use cases, 15 strong prototypes, and one genuinely surprising finding — the model beat human reviewers at figuring out what a business actually does.
Product lead Eugene Mann described the moment his team realized the humans had it wrong. “When we started hand-checking the results, we realized, ‘Wait a minute, the humans were wrong and the model was right.’” The task involved scanning sparse, mysterious websites — nightclubs were a prime example — and returning accurate summaries. GPT-4 didn’t just match human performance. It exceeded it.
The payment giant had already been using GPT-3 for routing support tickets and summarizing user questions, but Mann says GPT-4 “was a game changer” that opened up entirely new areas. Engineers from support, onboarding, risk, and docs teams all contributed ideas. The 15 prototypes that survived vetting focused on three main areas: understanding user businesses, answering technical documentation questions, and spotting fraud in community forums like Discord.
On the documentation front, GPT-4 acts as a near-instant virtual assistant. It digests Stripe’s extensive technical docs, identifies the relevant section, and summarizes solutions without requiring fine-tuning. “GPT just works out of the box,” Mann said. “That’s not how you’d expect software for large language models to work.” For fraud detection, the model analyzes post syntax on community platforms to flag potential bad actors coordinating activity or trying to regain credibility after being removed. The team is now exploring GPT-4 as a business coach that could advise on revenue models and strategy — a canvas Mann says changes daily.
💡 Key Takeaways
- GPT-4 outperformed human reviewers at classifying what businesses do, catching errors that human teams had missed — a rare case where the model wasn't just cheaper but actually superior.
- Stripe generated 50 AI use cases by pulling 100 employees off their normal jobs to experiment, proving that dedicated exploration time yields practical prototypes, not just hype.
- The fraud detection application in Discord communities shows LLMs can identify malicious intent from syntax alone, creating a new layer of security that doesn't rely on known bad actors.
- Stripe's documentation assistant required zero fine-tuning — GPT-4 understood technical docs and user questions out of the box, which Mann described as unexpected for enterprise software.
Keep reading: See related articles below for more coverage on this topic.
Get smarter about AI
The sharpest AI news, curated daily. Delivered free to your inbox.