How Blue J scaled AI tax research to 3,000 firms with a 1-in-700 error rate
Curated by the Inblix editorial team
Tax research is a grind. It’s not uncommon for a professional to burn hours, days, or weeks wading through statutes, case law, and rulings to answer a single complex question — and still get burned by an outdated interpretation. Blue J, a company founded by tax law professors in 2015, saw the ChatGPT moment as a chance to torch that old model entirely.
“We moved quickly because we already knew the problem and how to solve it,” said CTO Brett Janssen. The result is a tax research engine powered by OpenAI’s GPT-4.1 and a proprietary RAG system that draws on millions of curated primary sources and expert commentaries. It spits out detailed, fully-cited answers that feel less like a model hallucinating and more like guidance from a senior partner. They launched their first product just six months after ChatGPT debuted and have since rolled out across the US, Canada, and the UK, landing over 3,000 firms.
The key to this speed wasn’t just the model — it was an obsessive feedback loop. Every answer has a ‘disagree’ button. When a user flags a bad response, it’s not just a support ticket; it gets categorized by issue type and root cause. GPT-4.1 itself acts as a triage layer, clustering thousands of feedback points so the team can spot systemic problems. Is partnership tax law generating nonsense? They see the pattern and fix it. That grind has pushed their disagree rate below 1 in every 700 answers, and more than 70% of users now log in weekly.
That trust is the whole ballgame in a domain where a small mistake can trigger an audit. The system’s responsiveness was battle-tested when a sweeping U.S. tax bill passed in 2025. Blue J’s team had spent six weeks prepping the codebase, so users saw updated answers within hours of the bill being signed. For Janssen, the choice of model is a gating function, not just a feature. His team benchmarks every new release against 350 prompts across three countries, checking for instruction adherence and source alignment. If a new model can’t follow a complex tax instruction perfectly, it doesn’t ship. In an industry drowning in documents, that ruthless focus on precision is what separates a useful tool from a liability.
💡 Key Takeaways
- Blue J's RAG system combines GPT-4.1 with millions of proprietary tax documents, turning weeks of research into seconds.
- A closed-loop feedback system where users flag bad answers allows GPT-4.1 to triage issues, driving the disagree rate to fewer than 1 in 700 responses.
- Model selection is treated as a strict gating function; Blue J uses a 350-prompt legal benchmark to test every new release before deployment.
- Deep domain preparation allowed Blue J to deploy updated answers across its platform within hours of the 2025 U.S. tax bill being signed.
Keep reading: See related articles below for more coverage on this topic.
Get smarter about AI
The sharpest AI news, curated daily. Delivered free to your inbox.