Grab fine-tunes GPT-4o vision on just 100 images, maps SE Asia
Curated by the Inblix editorial team
Mapping Southeast Asia has always been a nightmare for conventional providers. The streets are narrow, the signage is chaotic, and everything changes faster than a satellite can blink. Grab decided the only way to solve it was to build something entirely its own—and it just got a serious upgrade from OpenAI.
The company’s GrabMaps service now uses GPT-4o with vision fine-tuning to turn millions of street-level images into hyperlocal mapping data. Those images come from a fleet of motorbike drivers and pedestrian partners equipped with 360-degree cameras, capturing the controlled chaos of cities across eight countries. The initial experiment was tightly scoped: match speed limit signs to their correct roads. With a training set of only 100 sample cases, the team fine-tuned the model through multiple hyperparameter adjustments. Baseline accuracy sat at a mediocre 67 percent. After two rounds of fine-tuning, it jumped to 80 percent—a 13-percentage-point gain that materially reduced the need for humans to untangle complex scenarios like elevated roads and visual occlusions.
Adrian Margin, Grab’s Head of Data Science for Geo Mapping, put it bluntly: the fine-tuned model handles complex geometries effectively enough to cut manual interventions and operational costs. The downstream numbers are what make procurement teams pay attention. Lane count accuracy rose by 20 percent. Speed limit sign localization got that 13-point bump. Fewer errors in map outputs means a more reliable platform, not just for Grab’s 42 million monthly users and 6 million driver-partners, but for the enterprise customers now buying into GrabMaps’ location intelligence capabilities.
What’s notable here isn’t just the accuracy gains—it’s the data efficiency. A hundred examples. That’s a shockingly small dataset to produce a production-ready improvement, and it signals how quickly domain-specific vision tasks can now be solved without massive labeling budgets. Grab is already pushing further into AI territory, building a multilingual voice assistant for visually impaired users and an advanced support chatbot designed to parse dense SOPs. CPO Philipp Kandal frames the OpenAI partnership as an acceleration play, but the subtext is clearer: in markets Western tech giants still can’t map properly, owning the data pipeline and fine-tuning bespoke models is becoming the competitive moat.
💡 Key Takeaways
- Fine-tuning GPT-4o on just 100 labeled examples improved speed limit sign matching from 67% to 80% accuracy, proving that production-grade vision tasks no longer require enormous training datasets.
- Lane count accuracy jumped 20% after fine-tuning, directly cutting manual mapping costs and reducing errors in scenarios like elevated signs that previously stumped automated systems.
- Grab is leveraging its driver network as a distributed data collection fleet, giving it a proprietary image pipeline that Western mapping providers can't replicate across Southeast Asia.
- The company is expanding AI beyond maps into multilingual voice assistance and SOP-trained chatbots, signaling that Grab sees fine-tuned models as a platform-wide advantage, not a one-off experiment.
Keep reading: See related articles below for more coverage on this topic.
Get smarter about AI
The sharpest AI news, curated daily. Delivered free to your inbox.