Video & Audio
Sora generates realistic and imaginative video scenes up to a minute long from text descriptions or still images, with strong understanding of physics, motion, and scene continuity. It is available to ChatGPT Plus and Pro subscribers and supports storyboarding, remixing, and video editing workflows. Sora's model architecture treats video as patches of visual data, enabling high temporal consistency. It is one of the most capable publicly available text-to-video models.