Intercom spent $100M to go AI-first: 3 lessons from the trenches
Curated by the Inblix editorial team
When ChatGPT hit the scene in 2022, most companies held a meeting. Intercom started building. Within hours of GPT-3.5’s release, the customer service platform was experimenting, and just four months later, they shipped Fin, an AI agent now resolving millions of queries monthly. The speed wasn’t a fluke—leadership canceled non-AI projects and committed $100 million to replatform the entire business around AI.
That bet forced massive internal changes: product teams got reorganized, a new AI-first helpdesk strategy took shape, and the underlying architecture was rebuilt to handle high-volume, complex customer interactions. But the real story isn’t the cash. It’s the process that lets them ship new capabilities in days, not quarters.
Intercom’s SVP of Engineering, Jordan Neill, told me their early GPT-3.5 experiments showed “glimpses of magic” but not enough reliability. Because they’d already mapped the model’s limitations, the moment GPT-4 arrived, they knew it was production-ready and launched immediately. That fluency with model capabilities also let them scrap a more complex reasoning stack for Fin Tasks after discovering GPT-4.1 could handle the same workflows with lower latency. Sometimes the simpler answer is the right one.
Underneath all this speed is a rigorous evaluation engine. Principal ML Scientist Pedro Tabacof said that when GPT-4.1 dropped, they had eval results in 48 hours and a rollout plan right after. They benchmark against real support transcripts, testing for multi-step instruction handling, brand voice, and tool call accuracy. That system isn’t just for text—they’ve extended it to voice, assessing tone, interruption handling, and background noise. The architecture itself is on its third iteration, deliberately modular so models can be swapped without reengineering the whole system. The lesson is blunt: if your AI platform isn’t designed to be ripped apart and rebuilt, you’ve already fallen behind.
💡 Key Takeaways
- Intercom scrapped a planned reasoning-model architecture after evaluation showed GPT-4.1 could handle complex workflows like refunds more reliably and with lower latency.
- A rigorous evaluation pipeline using real customer transcripts allowed Intercom to benchmark GPT-4.1 and roll it out across production systems within 48 hours of its release.
- Intercom's AI architecture is deliberately modular and on its third major iteration, designed to swap underlying models without reengineering the core system.
Keep reading: See related articles below for more coverage on this topic.
Get smarter about AI
The sharpest AI news, curated daily. Delivered free to your inbox.