OpenAI Teases Astra With 1979 Voice-and-Gesture Demo, but Leaves Product Details Out

The clip suggests Astra may combine speech, visual context and action. But OpenAI has yet to show how that approach will work in a general-purpose product constrained by new cyber safeguards.

By 2 min read
OpenAI Teases Astra With 1979 Voice-and-Gesture Demo, but Leaves Product Details Out
OpenAI Teases Astra With 1979 Voice-and-Gesture Demo, but Leaves Product Details Out

Listen to this story

The audio brief

About 1:35
0:001:35
Read transcript
OpenAI’s Astra teaser points back to a 1979 MIT experiment, suggesting the model may combine speech, visual context, and gestures—but it still does not show what the product will actually be. On September third, OpenAI shared footage of Put That There, created at the MIT Media Laboratory by Chris Schmandt and Eric Hulteen. In the demonstration, a user could point at an object, say “put that there,” and point to a destination. Speech supplied the action; gestures resolved words like “that” and “there.” When something was ambiguous, the system asked a focused follow-up question. The reference is revealing, but it is not a specification. Put That There worked with a constrained vocabulary, a graphical environment, and dedicated tracking hardware. Astra would face messier screens, longer tasks, and less predictable instructions if OpenAI intends this interaction style for general-purpose computing. There is another constraint: OpenAI classifies Astra as its first model to reach the Critical cybersecurity threshold. The company says it can discover unknown vulnerabilities and develop exploits across hardened systems when given sufficient tools and access. Planned safeguards include refusal training, account-level restrictions, and monitoring, and OpenAI delayed development while testing them. Those controls could also interrupt legitimate work in ChatGPT, Codex, and the API. Astra was said to be coming soon two days before the teaser, but its specifications and release timing remain undisclosed. The key question is how the finished product will make rich interaction useful while operating within those limits.

Story brief

3 key points

OpenAI’s September 3 Astra teaser points to multimodal interaction, using MIT’s 1979 Put That There demo as a reference for combining speech, visual context and gestures. The post offered no specifications or release date, despite OpenAI saying two days earlier that Astra would arrive soon. The model is also classified at the Critical cybersecurity threshold, making its eventual launch a test of whether richer...

  1. 01

    The reference demo resolved “that” and “there” through pointing, and asked focused follow-up questions when commands were ambiguous.

  2. 02

    Put That There relied on a constrained vocabulary, graphical environment and dedicated tracking hardware—not general-purpose computing conditions.

  3. 03

    OpenAI says Astra can discover unknown vulnerabilities and develop exploits across hardened systems with sufficient tools and access.

OpenAI has teased its forthcoming Astra model with a 1979 MIT demonstration built around a simple idea: a computer can understand an incomplete spoken command when it also knows where a person is pointing. The archival reference signals an interest in joining voice, visual context and action, but it offers no product specifications or release date.

An old answer to an unfinished command

On September 3, OpenAI posted a two-post thread featuring MIT’s Put That There interface and credited the MIT Media Laboratory, Chris Schmandt and Eric Hulteen. The company supplied no accompanying description of Astra, two days after saying the delayed model would become available soon.

Put That There paired spoken commands with pointing gestures inside a shared visual scene. A user could point to an object, say “put that there,” then point to a destination: speech supplied the action, while gestures identified the object and place. When part of a command was unclear, the system could ask a focused follow-up question.

A broader job than the MIT demonstration

The comparison sets a demanding test. Put That There operated in a narrow graphical environment, with a constrained vocabulary and dedicated tracking hardware. Astra would need to handle messier screens, longer tasks and less predictable instructions if it is to turn the same basic idea into a general-purpose product.

The safety limits remain part of the picture

OpenAI has classified Astra as the first model to reach the Critical cybersecurity threshold in its Preparedness Framework. The company says Astra can find previously unknown vulnerabilities and develop exploits across hardened systems when given the needed tools and access.

OpenAI’s planned safeguards

  • Initially restrict Astra’s most advanced cybersecurity functions.
  • Use refusal training, account-level controls and monitoring against cyber misuse and unauthorized model behavior.

OpenAI delayed parts of Astra’s development while testing those stronger controls. It has also warned that the safeguards may interrupt legitimate work in ChatGPT, Codex and the API. The teaser supplies an interface ambition; a product demonstration will need to show how useful interaction and those limits coexist.

Loading discussion...