Ai2 Releases an Open Model for Faster Scientific Reports With Traceable Sources
AstaBrief’s weights and training data let researchers adapt the report writer locally. Ai2’s claimed speedup comes from simplifying the whole workflow, not just swapping models.
Ai2 has made AstaBrief 8B and its training data available for institutions that want to adapt or run scientific report generation on their own infrastructure, including over local PDFs. The model powers Asta’s Fast mode, which generates a report in one pass rather than first summarizing and grouping retrieved sources. Ai2’s tests found faster end-to-end reports than its Claude-powered Thinking workflow, but most evaluation was completed in 2025, and its measures do not yet establish whether reports preserve the scope and strength of cited findings.
01
Ai2 measured average end-to-end times of 51.1 seconds for Fast mode and 178.5 seconds for Thinking mode; these are full-workflow comparisons, not model-only speeds.
02
Training retained 47,000 example reports and used about 6,000 preferred-report pairs; Ai2 reports 95% agreement between automated judges and human preferences.
03
Filtering out reports with too few cited statements produced the strongest gains in citation grounding; more elaborate filtering combinations added little.
Researchers can now download the model behind Asta’s faster scientific reports rather than rely solely on its hosted service. Ai2 released AstaBrief 8B on October 2, 2026, alongside its training data. It turns a research question and retrieved literature excerpts into a cited report, with an example workflow for adapting it to researchers’ own PDFs.
Ai2 wanted preliminary reports that scientists could get quickly and refine in later turns. Its existing Claude-powered workflow first summarized and grouped retrieved excerpts, then wrote the answer section by section. The team tested whether a smaller, specialized model could preserve report quality while cutting generation time and serving costs.
Starting with Qwen3-8B, Ai2 chose training from example reports and preferred answers instead of a reinforcement-learning approach. It described that choice as cheaper, easier to debug and more manageable. After filtering user logs for relevance, quality and privacy, the team had 90,000 research-focused queries to draw from.
Example-report training: Ai2 retained 47,000 reports after quality filtering. The reports came from a pipeline using Claude 3.5 Sonnet, Claude 3.7 Sonnet, o3, o4-mini and GPT-4.1.
Preference training: A separate query subset produced about 6,000 report pairs. Ai2 kept pairs only when GPT-4.1 and DeepSeek-R1 agreed on which report was better.
Ai2 reports 95% agreement between its automated judges and human preferences. Using multiple report generators and requiring both judges to agree was intended to reduce noise without treating any single model’s answers or judgments as ground truth.
Early training improved content quality but still fell short on relevance and citation grounding. Ai2’s strongest filtering gains came from removing training reports with too few cited statements. More aggressive filters and combinations did not add meaningful gains: consistent attribution in the examples mattered more than a complicated filtering recipe.
The other filters checked whether reports generated too much text from too little evidence, relied on lower-ranked papers, or drew on too few retrieved papers. Those tests separated citation coverage from the relevance and breadth of the sources used.
Ai2’s reported end-to-end report times
51.1 secondsAsta Fast mode
Ai2 reports this average across the full Fast-mode pipeline.
178.5 secondsClaude-powered Thinking mode
Ai2’s comparison measures complete workflows, not isolated model generation speed.
The speed figures need a time boundary. Ai2 says most training and evaluation finished in 2025, and it has not rerun the full evaluation against today’s frontier models. The results therefore describe the training and system design it tested, not a current competitive ranking.
AstaBrief now powers Fast mode in Asta’s Generate a report feature, alongside Claude-powered Thinking mode. Given the question and relevant excerpts, it writes the complete report in one pass. That skips the summarization and grouping stages of Thinking mode; literature retrieval still supplies the material the model synthesizes.
The open weights also give institutions a route to run report generation on their own infrastructure. Ai2 points to research questions that reveal sensitive or unpublished work as a reason for local deployment. Its example workflow provides a starting point for generating reports from local PDFs.
Ai2’s main development test used 200 user-written computer science questions. It separately measured necessary content, paragraph relevance, whether citations supported attached claims, and whether claims had citation support.
Secondary checks included DeepScholarBench, a 63-question research-synthesis benchmark built from recent ArXiv papers, and comparisons with Claude-powered reports. A small human study covered 14 questions from three scientific researchers, who ranked reports for preference, completeness, relevance, organization and citation accuracy.
But Ai2 identifies a harder problem: a report can cite the right study and still overstate its findings. It might generalize beyond the study’s sample or turn an observation into a recommendation. Its development metrics focused mainly on coverage, relevance and citation grounding. Ai2 says richer evaluation should also test whether reports preserve the scope and strength of source claims.
DeepScholarBench chart compares Asta Brief, DR Tulu, and Claude-powered Thinking mode on average score, organization, nugget coverage, citation precision, and claim coverage.Source: allenai.org.
Sources
allenai.orgOpen-sourcing AstaBrief, the fast report-generation model in Asta | Ai2
Reader comments
Newest comments first. Replies stay oldest first.