Ai2 has posted AstaBrief 8B on Hugging Face. The job is narrow: take a research question and retrieved literature excerpts, and write a cited report. It is available today in Asta’s report feature as Fast mode, next to a Claude-backed Thinking mode. The weights, the training data, and an example workflow for turning a researcher’s own PDFs into a report are all released. Open weights let an institution run the model on its own machines, including when the question would reveal unpublished work.

[1]
A single blue-black line along the top of thin paper, five blank tabs on the left edge, and empty paper below.
One line runs across the top and ends in a small ink blot. Five blank tabs sit on the left edge, and the paper below is empty. That is a report written in one pass, with the section tabs unused. An illustration, not a screenshot., AI-generated illustration, not a news photograph

The base is Qwen3-8B. The team considered reinforcement learning for long reports and chose supervised fine-tuning plus direct preference optimization. The post says reinforcement learning can be unstable and expensive. They wanted to see how far a cheaper setup, one that is easier to debug, could push report quality.

The supervised data comes from real user queries, not only synthetic prompts. After dropping test accounts, bot traffic, and queries that were too short, and after a model filter for non-English, non-scientific, and personal prompts, 90,000 research-focused queries remained. Target reports were produced by the multi-step ScholarQA pipeline, backed by a mix of Claude 3.5 Sonnet, Claude 3.7 Sonnet, o3, o4-mini, and GPT-4.1. Quality filtering left 47,000 usable examples.

Preference data used a separate set of queries that was not used for the supervised set. Each query had two reports. GPT-4.1 and DeepSeek-R1 each picked a winner. The post says the judges agreed with human preferences 95 percent of the time, and the team kept a pair only when both judges agreed. After filtering, the preference set was about 6,000 examples.

[1]

The speed change is in the pipeline. Given a question and the relevant snippets, the model writes the whole report in one pass. It does not do the snippet summarization and clustering that Thinking mode uses, and it does not write the answer section by section. The post says this was possible without giving up performance.

Two sentences in that passage should not be collapsed into one. One says report-generation time fell by nearly an order of magnitude compared with the proprietary models they tracked. The other says that across the full Asta pipeline, Fast mode averages 51.1 seconds per report and Thinking mode averages 178.5 seconds, about 3.5 times faster. The 3.5 times figure compares those two modes. It is not a multiple against every proprietary model.

[1]

The main development benchmark is SQABench-CS2, 200 user-written computer-science research questions. It separates coverage, whether each paragraph is on the question, whether each citation supports the claim it is attached to, and whether the report’s claims are covered by the citations. DeepScholarBench has 63 queries, and its metrics are not comparable with SQABench-CS2. The first supervised runs improved the writing, but answer precision and citation quality still lagged the Claude pipeline. Of four statistical filters, the largest gain came from dropping synthetic reports with low citation density. Stricter filters, combinations of filters, and learning-rate sweeps did not add a meaningful gain. The charts in the post compare systems, but the numeric scores are in the images, not in the text, so they are not repeated here.

In a 14-question human study, three researchers each contributed four or five questions and ranked reports from three systems, with ties allowed. On overall preference, DR-Tulu wins. Two of the three researchers prefer AstaBrief on citation accuracy. That is a three-person sample.

Early use of Fast mode: among 374 Asta users who tried it, 29.1 percent used it on two or more days, and the average was 3.67 report threads. Twenty-three percent of people who tried Fast mode never switched back to Thinking mode for later threads. Another 18 percent switched between the modes and used Fast for about 40 percent of their threads. Positive feedback was 84.2 percent, against 85.2 percent for Thinking mode. The post says the feedback is too sparse for a strong conclusion.

Most of the training and evaluation described here was completed in 2025. The proprietary models used to build the data and as comparison points reflect the frontier at that time. The team has not rerun the full evaluation against today’s frontier models. The numbers are evidence about those data and system choices, not a placement of this 8B model against the current frontier.

[1]

要点

  • AstaBrief 8B starts from Qwen3-8B and uses supervised fine-tuning plus direct preference optimization to write a cited report in one pass.
  • Filtering left about 90,000 queries, 47,000 supervised examples, and about 6,000 preference pairs. A pair was kept only when both judges agreed; agreement with humans was 95 percent.
  • Across the pipeline, Fast averages 51.1 seconds and Thinking 178.5. A separate sentence says nearly an order of magnitude. Those are not the same multiple.
  • Of 374 users, 29.1 percent used it on at least two days. Most training and evaluation was done in 2025 and was not rerun against today’s frontier.