OpenAI launches Sora video model built on visual 'patches'
Curated by the Inblix editorial team
OpenAI has finally pulled back the curtain on Sora, its text-to-video model that turns prompts, still images, or existing footage into high-definition clips. The model can generate up to 20 seconds of 1080p video, an output window that puts it squarely in the realm of short-form creative work rather than full-length features—at least for now. Like DALL·E before it, Sora comes with feeds for Featured and Recent creations, a clear nod to building a community showcase rather than just shipping a utility.
The technical guts are where things get genuinely interesting. Sora is a diffusion model that starts with static noise and refines it into coherent motion, but the real magic is in its architecture. Borrowing a page from large language models, OpenAI treats video as ‘visual patches’—spacetime chunks decompressed from a lower-dimensional latent space. This transformer-based approach means the model can keep a subject consistent even when it temporarily leaves the frame, solving one of the most stubborn problems in AI video generation. It’s the same recaptioning technique from DALL·E 3 that helps Sora follow text prompts with unusual fidelity.
OpenAI didn’t just train this on whatever was lying around. The training data mix includes publicly available web crawls, proprietary content licensed through a partnership with Shutterstock Pond5, and custom datasets built in-house. Human feedback from AI trainers and red teamers also fed into the process. All of it passed through pre-training filters designed to strip out explicit, violent, or otherwise toxic material—an extension of the same safety playbook used for DALL·E 2 and 3. The company frames Sora as a stepping stone toward models that genuinely understand and simulate the real world, calling that capability ‘an important milestone for achieving AGI.’
But video generation opens a nastier can of worms than still images ever did. OpenAI acknowledges the risks are novel: likeness misuse, deepfake-style deception, explicit content generated at scale. The safety stack builds on lessons from deploying DALL·E in ChatGPT and the API, layered with external red teaming and ongoing research. Starting in February 2024, the company also gathered input from hundreds of visual artists, designers, and filmmakers across more than 60 countries. Whether that feedback loop is enough to keep Sora from becoming the next misinformation machine is the question nobody can answer yet.
💡 Key Takeaways
- Sora converts video into 'visual patches'—a transformer-friendly representation that helps maintain subject consistency across frames, solving a problem that tripped up earlier models.
- OpenAI partnered with Shutterstock Pond5 for proprietary training data, signaling a shift toward licensed content deals rather than relying solely on public web crawls.
- The model was tested with feedback from hundreds of creative professionals across 60+ countries, but the company still frames safety largely in terms of iterative mitigation rather than solved challenges.
Keep reading: See related articles below for more coverage on this topic.
Get smarter about AI
The sharpest AI news, curated daily. Delivered free to your inbox.