Inside Morgan Stanley’s relentless AI eval framework
Curated by the Inblix editorial team
Most enterprise AI stories are heavy on hype and light on the boring-but-critical infrastructure that actually makes a deployment work. Morgan Stanley’s collaboration with OpenAI is the opposite — a case study in sweating the details. The bank didn’t just drop GPT-4 into its wealth management division and pray. It built a rigorous evaluation framework that tests, measures, and refines every AI interaction before an advisor ever sees it. The result: their internal chatbot, AI @ Morgan Stanley Assistant, has hit over 98% daily adoption across advisor teams. That’s not a typo. Nearly everyone is using it.
The evals started with three concrete goals — speed up document retrieval, automate summarization of research reports, and surface insights tailored to individual clients. Human prompt engineers and advisors graded the model’s outputs for accuracy and coherence, then fed that feedback directly into prompt refinement. “We went from being able to answer 7,000 questions to a place where we can now effectively answer any question from a corpus of 100,000 documents,” says David Wu, the firm’s Head of Firmwide AI Product & Architecture Strategy. That’s a 14x expansion in the knowledge base, but it only works because the underlying retrieval methods were tuned in lockstep with OpenAI.
The framework isn’t static. It evolved to include translation evals for multilingual clients and, for the newer meeting-summary tool called AI @ Morgan Stanley Debrief, custom datasets representing different meeting types. Powered by Whisper and GPT-4, Debrief turns consented Zoom recordings into CRM-ready client notes and draft follow-ups. The kicker? Advisors review and tweak everything before it’s finalized. Automation with a human kill switch. One executive noted follow-ups that once took days now happen within hours, and advisors are having conversations on topics they previously couldn’t address because the friction between knowledge and communication “has gone to zero.”
Compliance is woven into the evals, not bolted on afterward. Morgan Stanley runs a daily regression suite of sample questions to sniff out weaknesses and works with OpenAI to adjust retrieval methods for greater accuracy. OpenAI’s zero data retention policy sealed the trust deal for a firm handling extraordinarily sensitive financial data. Document access across the organization jumped from 20% to 80% — a number that hints at just how much institutional knowledge was effectively locked away before the chatbot made it searchable. The lesson for other enterprises is blunt: if you’re not evaluating your models as obsessively as Morgan Stanley evaluates theirs, you’re not ready for prime time.
💡 Key Takeaways
- Morgan Stanley didn’t just deploy GPT-4; it built a continuously evolving evaluation framework that tests every use case and refines prompts based on direct feedback from human advisors and engineers.
- The AI @ Morgan Stanley Assistant expanded from answering 7,000 questions to effectively handling queries across a corpus of 100,000 documents, a leap that required close collaboration with OpenAI on retrieval methods.
- The new Debrief tool uses Whisper and GPT-4 to turn consented Zoom meetings into draft follow-ups and CRM notes, but every output is reviewed and adjusted by a human advisor before it’s finalized.
- Document access soared from 20% to 80% across the firm, revealing how much institutional knowledge was previously trapped in unstructured formats that AI-powered search finally unlocked.
Keep reading: See related articles below for more coverage on this topic.
Get smarter about AI
The sharpest AI news, curated daily. Delivered free to your inbox.