Modelspublished

Skild AI’s S1 Uses One Human Video to Prompt Robots Through 10-Minute Tasks

The model’s central bet is that a demonstration can become immediate instruction rather than another task-specific training run. The harder question is whether that capability can turn into dependable deployments.

By 3 min read
Skild AI’s S1 Uses One Human Video to Prompt Robots Through 10-Minute Tasks
Skild AI’s S1 Uses One Human Video to Prompt Robots Through 10-Minute Tasks

Listen to this story

The audio brief

About 1:34
0:001:34
Read transcript
Skild AI has introduced S1, a robot foundation model that treats one human task video as an immediate instruction. Show it someone repotting a plant, making coffee, or cooking pancakes, and the company says S1 can reproduce the sequence without a new, task-specific training run. That is the central bet: a demonstration becomes a prompt, instead of the beginning of another lengthy post-training cycle. Skild AI co-founder and CEO Deepak Pathak says the system can handle multi-step jobs lasting up to ten minutes. In technical terms, this is in-context learning—the video is supplied in the model’s immediate input, or context, rather than used to change the model itself. S1’s underlying training combines four data sources: teleoperation, human video, computer simulation, and glove-captured actions. Each fills a gap. Teleoperation is closer to the robot’s hardware, but slower and less diverse. Human video is abundant and varied, but harder to translate into robot controls. Simulation scales, though it differs from the real world, while glove data is more scalable than teleoperation but less directly applicable. S1 is also designed to be omni-bodied, spanning quadrupeds, humanoids, and fixed robotic arms. But the humanoid results cited came from an earlier model, and S1 is not ready for homes. The consequential test arrives in the coming weeks, when Skild expects production-oriented demonstrations and customer-acquisition efforts: can video prompting sustain reliable work outside controlled demos?

Story brief

3 key points

Skild AI’s S1 pairs a robot foundation model with in-context learning: a single human task video can guide execution without a new post-training run. The company says this supports multi-step jobs lasting up to 10 minutes, including repotting, coffee-making, and pancakes. S1 is trained across teleoperation, human video, simulation, and glove-captured data, and is designed for quadrupeds, humanoids, and arms. Its key...

  1. 01

    Skild expects production-oriented demonstrations and customer-acquisition efforts in the coming weeks.

  2. 02

    Teleoperation data is higher quality but slower and less diverse; human video is abundant but harder to map to robot controls.

  3. 03

    S1’s omni-bodied design targets quadrupeds, humanoids, and fixed arms rather than a single robot platform.

Skild AI has introduced S1, a robot foundation model built around an unusually direct teaching method: show a robot a person completing a task in one video, then use that video as part of the model’s prompt. The company says the approach can cover complex jobs lasting as long as 10 minutes, potentially reducing the task-by-task retraining that slows robotics development.

The demonstration is an instruction, not a new training run

S1 uses in-context learning, meaning the human task video is supplied in the model’s immediate input, or context. Skild AI co-founder and CEO Deepak Pathak said the robot can then follow the demonstration without first changing the model through a dedicated training process. That differs from the common workflow in which robot models are post-trained for a new task, a step Pathak described as lengthy and a constraint on scale.

The distinction is consequential because S1 is aimed at sequences rather than a single brief movement. Skild AI cited repotting a plant, making coffee and cooking pancakes among the tasks it says the model can handle. Pathak characterized these as long-horizon work: a label for actions that unfold through multiple steps over time.

Skild AI’s S1 demonstration presents a human video as context for robotic execution. Video via therobotreport.com.

S1’s promise is a change in the unit of robot instruction: a demonstration video becomes a prompt, rather than the start of a bespoke training cycle.

Four data sources feed the underlying model

The one-video interaction sits on top of broader pretraining. Skild AI said it trains S1 with four kinds of data, rather than relying on one collection method. The company’s rationale is that each source brings a different trade-off: robot-controlled examples are close to the hardware but slower to gather, while human video is plentiful and varied but harder to translate directly into robot behavior.

  • Teleoperation data records a person directly controlling a robot; Skild says it is high quality but slow to capture and less diverse.
  • Human videos offer abundant, diverse examples, but Skild says their actions are farther removed from a robot’s body and controls.
  • Simulation scales on computers but retains a gap from real-world operation; glove-captured actions are more scalable than teleoperation but less directly applicable to robots, Skild said.

One model, several bodies, with a notable gap

Skild describes S1 as omni-bodied: it is intended to work with quadrupeds, humanoids and fixed robotic arms rather than one hardware type. The company is also not targeting a single industry or task category. That breadth is part of its general-purpose pitch, but the humanoid portion is less mature: Skild plans to improve S1’s humanoid performance, and the humanoid adaptation results Pathak cited came from an earlier model version.

The next test is operating outside the demo

Pathak said S1 is not yet ready for rollout into people’s homes, even as he framed it as an early indication of where robot models may be heading. Skild expects to show the system helping in production and supporting customer acquisition in the coming weeks. Those demonstrations will provide the more consequential evidence: whether a model that follows a video prompt can sustain useful work in a deployment setting.

Sources

  1. therobotreport.comSkild AI unveils S1 flagship robot foundation model - The Robot Report