Research investigation R0925 / claim audit

Google’s Ten-Minute AI Film Tests One System, Not Four

Google attributes the ten-minute film to A²RD. The four linked papers evaluate different artifacts and do not report a test of all four systems as one integrated pipeline.

Archived snapshotv1Sep 25, 2026
Verified observations
9

9 measured fields

Supported claims
9

8 material findings

Cited sources
5

5 primary or authoritative

Research score
81

Automated topic and evidence score

Interactive figureGoogle’s Ten-Minute AI Film Tests One System,...
CSV JSON
Data status4.03 verified records across 1 period

Snapshot only. There is not enough history to claim a trend yet.

Verified observationHover or focus any mark for exact valuesLast updated Sep 25, 2026

Version ledger

Frozen public editions

Each edition preserves the records, method, sources, and downloads available at publication time.

  1. v1 / latestSep 25, 20269 records / 5 sources

    Initial public snapshot with 9 records and 5 cited sources.

Coverage note

The packet states that the five-publication evidence check used September 25, 2026 snapshots at 15:35:07 UTC; this synthesis does not independently refresh that record. Superpower Daily has already published coverage of the announcement and its ten-minute demonstration. The candidate’s Project Suncatcher material concerns a different topic and is not used here.

Dataset ID
spd:google-s-ten-minute-ai-film-tests-one-system-not-four-7e07fdb1
Stable URL
/research/google-s-ten-minute-ai-film-tests-one-system-not-four-7e07fdb1
Version
v1
Coverage
2026-04-15
Records
9
Fields
7
Updated

Read the data

The records behind the figure

CSV JSON
Google’s Ten-Minute AI Film Tests One System, Not Four data records
EntityMetricValueUnitObservedSourceTransform
CANVAS keyframes + Veo-3.1-previewAppendix Table 17 average score for approximately 1.5-minute videos; Gemini-2.5-Flash judge4.03FilMaster average score as reported2026-04-15https://arxiv.org/abs/2604.13452—
Co-DirectorGenAd-Bench average score for rendered ads81.4; final videos have four shots totaling 12 secondspoints on 0–100 scale2026-04-27https://arxiv.org/abs/2604.24842—
A²RDMLLM-judge average character consistency on ten-minute scenarios90.5%percent as reported2026-05-07https://arxiv.org/abs/2605.06924—
A²RDMLLM-judge average environment consistency on ten-minute scenarios84.0%percent as reported2026-05-07https://arxiv.org/abs/2605.06924—
A²RDMLLM-judge average object consistency on ten-minute scenarios91.5%percent as reported2026-05-07https://arxiv.org/abs/2605.06924—
A²RDNumber of ten-minute LVBench-C scenarios evaluatedTen scenariosscenarios2026-05-07https://arxiv.org/abs/2605.06924—
VQQAReported absolute improvement over vanilla generation on T2V-CompBench+11.57%absolute percent improvement as reported2026-03-12https://arxiv.org/abs/2603.12310—
CANVASReported background-continuity improvement over best-performing baseline on storyboard evaluation21.6%percent improvement as reported2026-04-15https://arxiv.org/abs/2604.13452—
A²RDVBench-Long inter-shot character consistency for approximately one-minute generated videos0.7353reported metric score2026-05-07https://arxiv.org/abs/2605.06924—

Measurement technique

How to read this report

  1. 01Use the supplied September 25 evidence snapshots of Google Research’s September 24 announcement and its four linked papers; recovering this saved evidence is not a new collection or experiment.
  2. 02Build an evidence matrix with one row per system and columns for evaluated artifact, duration or sample, measured outcome, judge or evaluation method, and what the result cannot establish.
  3. 03Separate Co-Director’s 12-second rendered-ad evaluation from its static-storyboard test; distinguish CANVAS’s main storyboard results from its appendix video comparison using storyboard keyframes.
  4. 04Separate A²RD’s one-, three- and five-minute video evaluations from its ten-scenario, ten-minute MLLM-judged results; classify VQQA’s benchmark gains as prompt-refinement results rather than long-story consistency results.
Next report / 01AI Model Economics Index All research reports
YOUR READING SPACE

Notifications