NVIDIA says the detector can now identify text-to-video content with 99.3% accuracy and image-to-video content with 97.7% accuracy. Rather than positioning the result as a definitive verdict, the tool produces another signal that newsrooms, forensic teams and other organizations can use when reviewing questionable footage.
That technology is already moving into products used by media companies. Dalet is integrating SVD into a cloud-based verification system where newsroom staff can submit video and inspect the resulting scores and metadata without leaving the Dalet interface.
TwelveLabs is taking a similar approach with Compliance by TwelveLabs. Its platform can screen video against different compliance requirements while using NVIDIA’s detector to flag frames that may contain synthetic content.
Wowza plans to bring the detector directly into live streaming. Its Video Intelligence Framework will be able to examine video feeds in real time for objects, scenes and potential signs of AI generation. The system can run in the cloud, on premises, at the edge or in fully disconnected environments.
NVIDIA is also targeting one of the most visually demanding areas of television: live sports.
Its 3D Body Pose technology can turn video from a single camera into data describing the position and movement of a person’s joints. That could give sports organizations another way to analyze player movement, biomechanics, officiating decisions and safety without requiring athletes to wear motion-capture markers.
Vizrt is already putting the technology into virtual studios. A person’s movement can be tracked and used to control effects such as shadows, reflections and lighting inside a 3D environment.
NVIDIA’s Video Frame Generation technology tackles a different sports problem. Instead of requiring every frame of a slow-motion replay to come from a high-speed camera, generative AI can create new frames between the ones that were actually captured.
Ross Video is integrating the technology into its Rio Replay system, where it currently supports 6x slow-motion generation. Work is underway on 8x interpolation. That means a replay can appear considerably smoother even when the original camera did not capture every frame being shown.
NVIDIA is also using AI to improve video that has already been recorded. Video Super Resolution can upscale footage while cleaning up blur, noise and compression artifacts, while TrueHDR can convert standard-dynamic-range footage into HDR in real time at brightness levels of roughly 2,000 nits.
The three technologies — Video Super Resolution, Video Frame Generation and TrueHDR — can be combined in the same processing pipeline. That gives streaming services and broadcasters a way to enhance existing footage without treating each improvement as a separate workflow.
Language is another major focus of NVIDIA’s IBC announcements.
The company’s LipSync technology changes a person’s mouth movements to match translated or replacement audio while trying to preserve the rest of the performance, including blinking, head position and body movement. NVIDIA says the latest version also handles partially obscured faces more effectively and better preserves details around the mouth and teeth.
Active Speaker Detection is designed to help the system determine who is actually talking when several people appear on screen.
NDI is using NVIDIA’s media technology for real-time translation and lip-synced dubbing. Instead of producing entirely separate feeds for different languages, broadcasters can create multiple localized versions from the same underlying stream.
NVIDIA is bringing those localization tools into Holoscan for Media as well, combining translated speech, captions, synchronized video, dubbing and localized graphics in one reference workflow. AI-Media, CAMB.AI, Chyron and Panjaya are contributing technologies to different parts of that system.
NVIDIA is also introducing Sports Intelligence Playbooks for organizations that want AI models to understand more than generic video.
Sports leagues and media companies can use their own footage and annotations to fine-tune NVIDIA models around the rules, players, scoring and strategy of a particular sport. The playbooks cover the process from preparing the data through training, evaluation and deployment.
Early tests showed a large difference after specialization. NVIDIA said multiple-choice accuracy on previously unseen footage increased from about 53% to 94% when the questions resembled those used during training. Open-ended evaluation rose from roughly 5.7% to 66%.
Machina Sports is combining the playbooks with its own sports data and agent technology, while Wowza is using models including NVIDIA Cosmos 3 and Nemotron to identify sport-specific moments in live streams.
Underneath those individual AI features, NVIDIA is also trying to make the infrastructure connecting them less complicated.
Holoscan for Media is being integrated with Media Exchange Layer, or MXL, so different software applications can exchange live video, audio and data while running on shared infrastructure. That could allow a broadcaster to combine conventional production software with AI-powered applications without building a separate technical environment around every new tool.
Taken together, NVIDIA’s IBC announcements show what AI in media production is starting to look like beyond simply generating videos. The company is putting it behind the scenes — checking whether footage is real, filling in frames during replays, following athletes across the screen, cleaning up old video and translating broadcasts while they are happening.
This analysis is based on reporting from Nvidia.
Image courtesy of Nvidia.
This article was generated with AI assistance and reviewed for accuracy and quality.