The system then interprets the task, identifies the relevant objects and sequence of actions, and transfers that information into robot movements. Skild describes the method as in-context learning because the new task is learned from the prompt rather than through a separate training cycle.
“Learning by experience, and not preprogramming, is the step change that has happened in robotics,” Skild AI co-founder and CEO Deepak Pathak said.
Skild says S1 can handle previously unseen tasks lasting as long as 10 minutes. Examples include potting plants, preparing pancakes, brewing pour-over coffee and assembling kits, with some workflows requiring dozens of manipulation steps.
In one plant-potting test, the company recorded a demonstration and had a robot performing the task autonomously 11 minutes later. S1 can also respond when objects are moved, recover after mistakes and combine previously learned skills into new sequences.
Skild reported a 66% per-step success rate on unfamiliar multistep tasks, compared with 9% for a similar AI system. The company also estimates that a short video demonstration can provide roughly the same training value as 380 manually collected examples, which could otherwise require 50 to 100 hours of human work.
The technology is already being used in industrial settings. Skild, Nvidia and Foxconn are deploying the Skild Brain on dual-arm robots assembling Nvidia Blackwell systems.
In one demonstrated process, a robot installs a busbar and limit block, fastens 16 screws and adjusts when conditions change during the workflow. The task requires precise movement, physical contact control, sequence tracking and the ability to recover when the environment differs from what was expected.
Skild built S1 using Nvidia infrastructure across several stages of development. Nvidia Cosmos models and Cosmos Curator are used for training data and video processing, while Isaac Sim and Omniverse provide simulated environments for testing and data generation.
The company also uses Isaac Lab and the Newton physics engine for reinforcement learning and physical simulation. Nvidia Nsight is used to identify performance bottlenecks, while TensorRT helps optimize inference for real-time robot operation.
Skild and Nvidia are also jointly developing GPU-accelerated simulation solvers designed to model how robots grip, touch and manipulate solid objects. The companies say those tools will become available to developers through Newton.
Skild’s commercial deployments currently span manufacturing, logistics, inspection, security and food preparation. The company says data from customer deployments can also feed back into the broader model when customer agreements allow it.
The larger goal is to reduce the amount of engineering required each time an industrial robot encounters a new job. S1 still has to prove that its one-video approach can remain reliable across longer and more complex workflows, but Skild is betting that faster task adaptation can make general-purpose robots more practical in changing workplaces.
This analysis is based on reporting from Nvidia.
Image courtesy of Nvidia.
This article was generated with AI assistance and reviewed for accuracy and quality.