Versos AI Adds Plain-Language Curation for Licensed Video Training Data

The workflow is designed to turn a detailed footage brief into a dataset while keeping ownership, licensing and provenance information alongside the selected material.

By 2 min read
Versos AI Adds Plain-Language Curation for Licensed Video Training Data
Versos AI Adds Plain-Language Curation for Licensed Video Training Data

Listen to this story

The audio brief

About 1:31
0:001:31
Read transcript
Versos AI has launched a workflow that turns a plain-language video brief into a structured training dataset, while keeping ownership, licensing, and provenance information attached to the footage it selects. That matters because a request for training data is rarely just visual. A team might need footage showing specific subjects, actions, or environments, but also impose limits on duration, resolution, format, language, production quality, and permitted use. Versos says its agents translate those requirements into search and grading criteria, search video at both scene and frame level, assess candidate clips, and assemble the results into a dataset. The company’s pitch is that this could replace a chain of separate searches, manual review, and preparation work. The differentiator is the rights layer. The system is intended to surface ownership and licensing details alongside visual relevance and technical fit, so a promising clip is not treated as usable without its permission context. Versos says the workflow runs with NVIDIA’s CUDA Toolkit, uses Nemotron Ultra for the agents, and relies on LangChain to coordinate the process. It also describes the architecture as model-agnostic, meaning customers could use other open-weight or frontier models when cost, performance, or deployment needs differ. Versos plans a live demonstration at IBC2026 in Amsterdam, from September 11 through 14. The key constraint is still untested performance: no independent evidence has been provided on customer libraries or real licensing requirements.

Story brief

3 key points

Versos AI is adding agentic curation to its video-data workflow, combining scene- and frame-level search with technical filtering and rights metadata for controlled AI licensing. Buyers can specify requirements such as subjects, actions, duration, resolution, language, format, production quality, and permitted use, then receive a structured dataset rather than a raw search result. The approach could reduce manual...

  1. 01

    Versos’ agents translate natural-language briefs into search, grading, and dataset-assembly criteria.

  2. 02

    Search results include ownership, licensing, and provenance information—not just visual relevance.

  3. 03

    The stack uses NVIDIA CUDA Toolkit, Nemotron Ultra, and LangChain; Versos says it is model-agnostic.

Versos AI has launched a workflow that lets AI teams describe needed video training data in plain language, then turns that request into structured search and curation criteria. The system is designed to find, assess and assemble footage prepared for controlled AI licensing into a dataset.

One footage brief, many constraints

Video-data requests can combine what must appear on screen with technical and legal conditions. Versos says its workflow is designed to handle subjects, actions and environments alongside duration, resolution, format, language, production quality and permitted use. The company says meeting a detailed brief can otherwise require multiple searches, manual inspection and dataset preparation.

The workflow’s stated inputs

  • Content requirements, including subjects, actions and environments.
  • Technical requirements, including duration, resolution, format and language.
  • Production-quality and permitted-use requirements.

Rights information is part of the search

The distinguishing feature is not simply prompt-based video search. Versos says the workflow searches footage prepared for controlled AI licensing and maintains its associated ownership, licensing and provenance information. That places documented permissions alongside relevance and technical fit when the system selects material for a training dataset.

The initial technology stack and next demonstration

Versos says NVIDIA CUDA Toolkit supports inference in its Video Library Intelligence Platform, NVIDIA Nemotron Ultra runs the agents, and LangChain coordinates the workflow. The company says its architecture is model-agnostic, allowing other open-weight or frontier models when customer performance, cost or deployment requirements differ.

Versos plans to demonstrate the capability at IBC2026 in Amsterdam from September 11 to 14. The announcement describes the intended workflow, but the company has not presented independent evidence here about how well it performs across customer datasets or licensing requirements.

Sources

  1. markets.businessinsider.comVersos AI Launches Agents Built with NVIDIA NeMo to Speed Curation of Video Training Data for AI Models

Loading discussion...