OpenAI’s internal AI data agent speeds insights across teams
Curated by the Inblix editorial team
OpenAI built a custom AI data agent to help its 3,500+ internal employees quickly navigate a massive data platform with over 600 petabytes across 70,000 datasets. The agent, powered by GPT‑5.2, allows teams in Engineering, Data Science, Finance, and Research to ask questions in natural language and get answers in minutes instead of days. It combines Codex for table-level knowledge with organizational context, and uses a self-learning memory system that improves with every interaction. The tool addresses common pain points like finding the right tables, avoiding SQL bugs, and handling complex joins. OpenAI built it using the same tools available to external developers, including Codex, the Evals API, and the Embeddings API. Why it matters: This shows how even the creators of frontier models need to solve real data bottlenecks internally, proving that AI’s biggest impact often comes from reducing friction in everyday knowledge work.
💡 Key Takeaways
- OpenAI’s internal AI data agent lets employees query 600+ petabytes of data using natural language, cutting analysis time from days to minutes.
- The agent combines Codex-enriched table context with organizational knowledge and a self-improving memory system to reduce SQL errors and table-searching fatigue.
- OpenAI used the same developer tools it offers externally (Codex, GPT‑5, Evals API, Embeddings API) to build this bespoke solution.
Keep reading: See related articles below for more coverage on this topic.
Get smarter about AI
The sharpest AI news, curated daily. Delivered free to your inbox.