Sora And Healthcare: Revolutionising Healthcare With AI Video Generation
Text-to-video generators will be excellent for patient education – among other use cases -, and present new challenges regarding privacy and credibility.

Key Takeaways
OpenAI recently revealed Sora, its new AI model that can generate short videos from text prompts.
Sora is not available to the public currently and is being stress-tested to counter potential nefarious activities.
Sora is not the only text-to-video AI generator announced and such technologies can be applied to the healthcare setting but will also generate new concerns.
“Historical footage of California during the gold rush.” From this simple prompt, OpenAI’s new text-to-video model Sora could generate a 25-second footage of what one could at a glance indeed think is real footage from that era. While such abilities of the company’s new artificial intelligence model were demonstrated in videos, Sora is not currently publicly available.
Nevertheless, its potential to create videos from simple text is bound to have consequences for the film and advertising industries. This has also led us to contemplate its possible medical and healthcare implications when it will be publicly available later this year.
Sora’s story
When OpenAI publicly unveiled Sora for the first time in early 2024, they described it as “an AI model that can create realistic and imaginative scenes from text instructions”. This was no hyperbole as the demoed clips weren’t far from being realistic.
However, this is not the first example of text-to-video models. Last year, New York-based Runway launched Gen-2, and Google announced its Lumiere model a few weeks before Sora’s reveal. What makes Sora stand out is that it understands both the prompt and how elements in the output should interact in the physical world.
More technically, Sora is a diffusion model that builds on OpenAI’s former DALL·E and GPT models research. It is able to interpret the text input to output convincing videos of up to 60 seconds. It generates a clip by initiating a “noisy” clip with static elements and incrementally removing that noise.
OpenAI acknowledges that its model has some downsides. For example, it can struggle with complex scenes, precise descriptions, and spatial elements (such as mixing up left and right). The need to further fine-tune the model is likely one of the reasons why Sora is not publicly available. OpenAI has provided access to Sora to “red teamers”, who are stress-testing the tool to better understand and counter how it deals with nefarious users.
From text to tailored healthcare videos
While Sora is not yet publicly available and other text-to-video models have yet to be applied in the healthcare setting, it is worth considering the potential of such technologies in the healthcare context. Dr. Meskó discussed with The Medical Futurist community how it will be used in healthcare and the following potentials stand out.
1. Patient education and management
With the ability to make realistic videos, healthcare teams can aid patients in better understanding their condition. Videos could depict the progression of their condition and the impact of adequate medication and lifestyle changes. Such a storytelling tool could assist in the patients’ treatment adherence and improve their health literacy. Furthermore, patients and their physicians could generate videos to devise the appropriate treatment journey that fits the patients’ individual schedules.
For example, a new video could be created to instruct a patient on how to properly perform exercises recommended by a physiotherapist or how to use a personal health sensor at home.
2. Medical training materials
Generating videos could assist in training healthcare professionals. Through such visualisations, trainees could better understand rare conditions, visualise complex procedures, and even simulate challenging conditions.

These are only two examples, but the actual use cases that could appear are limited by one’s creativity. As such, other uses of AI-generated videos are likely and can range from visualising medical research findings to visual aids for medical students.
Challenges of AI-generated videos
While outputs from text-to-video generators have the potential to enhance healthcare practice and delivery, they will also undoubtedly lead to some challenges. Such AI tools rely on training data and to generate medically-relevant videos, they will have to be trained on similar content. This can raise well-founded concerns over patient privacy. Most people won’t take it lightly to be filmed during a medical consultation or surgery only to have that video used to train an AI.
Moreover, with the ease of creating realistic videos comes the increased risk of misinformation content. In the healthcare context, it can have irreparable consequences if a patient is provided wrongful information on their condition management. One solution could be to integrate watermarks, visible or embedded in metadata, in AI-generated videos.
Such concerns, while currently speculative, are well worth our attention. With OpenAI’s Sora expected to be available this year, similar tools will likely follow suit. The healthcare industry needs to be prepared to not only consider the opportunities of such tools but also to counter its challenges.
Written by Dr. Bertalan Meskó & Dr. Pranavsingh Dhunnoo



