How should we judge an AI scientist? For most of the field so far, the answer has been "exams" — isolated benchmark problems, single tasks, fixed answer keys. A new Perspective paper, "Measuring AI Scientists: From Exams to Discovery" (Du, Dillmann, Laurent, Jansen, Jia, Schmidt, White, Persson, Arnold & Duan, 2026), argues that this is the wrong yardstick — and the first practical data supporting the shift has just landed: an AI-for-science platform called MIRA tops an end-to-end research benchmark while being the cheapest system in the comparison.
The claim: test the whole episode, not the isolated problem
Real science is a loop: read the literature, form a hypothesis, run an experiment or simulation, take measurements, revise, repeat. An exam-style benchmark tests one frozen slice of that loop — pattern matching dressed up as reasoning. Discovery episodes test the loop itself: can the system take a vaguely specified goal and carry it to a validated result, with an evidence trail a human can audit?
This is not just a measurement quibble. Evaluation defines the incentive structure — whatever gets measured is what gets optimized. Keep awarding points for isolated problems and you get specialized parrots. Score discovery episodes and the industry is forced to build systems that can run a research project from end to end.
The data: best score, lowest cost, at the same time
The accompanying practical case is MIRA (Towards a General AI Scientist), an open AI-scientist platform from the Chinese AI-for-science team Deep Principle. Under external evaluation on ResearchClawBench — an end-to-end research benchmark spanning Chemistry, Energy and Materials — MIRA scores 17.51 against Codex's 14.60, ahead of seven other agents in the comparison.
The more striking number is the cost column. MIRA completes a research task at $0.67 per task — the lowest cost in the benchmark. The second-ranked system (ResearchClaw) costs $0.89 and still scores lower; comparable-cost rivals like Nanobot ($0.70) and OpenClaw ($0.72) finish far behind at 10.90 and 13.63. On the drug-discovery suite of SciAgentArena, MIRA reaches 81.1%. The pattern breaks a comfortable assumption: that the best scientific agent must be the most expensive to run.
The framework: why "cheap + best" is a structural signal
This is the same pattern the model layer already witnessed with DeepSeek — a lower-cost system matching or beating far more expensive frontier rivals, rewriting the assumption that capability is a linear function of compute budget. Now the pattern is repeating one layer up, among research agents, and the evaluation regime is changing to make it visible.
Two consequences follow:
- Cost becomes a first-class scientific metric. If discovery episodes are the unit of evaluation, then "validated results per dollar" becomes the productivity measure — and cost-efficient systems become the default choice, not the bargain option.
- Evidence traceability becomes a product feature. An episode-based evaluation cannot be scored without an auditable trail from hypothesis to measurement. The systems that make that trail explicit will be the ones that win the new yardstick — which is also the property scientific rigor actually requires.
This is also part of a broader 2026 pattern: Nature-documented systems (ERA, The AI Scientist, MIRA) have moved from reasoning about science to acting on it — generating hypotheses, running code, driving simulations, even interacting with records. The question is no longer whether AI can do science; it is how we tell good autonomous science from cheap theater. That is exactly the question the new evaluation paradigm is designed to answer.
What to do about it
- If you evaluate AI-science tools: stop benchmarking isolated problems in your field. Run a full discovery episode — a vague goal, open literature, a real experiment or simulation — and audit the evidence trail at the end.
- If you build or buy research agents: add a "cost per validated result" line to every comparison. The MIRA numbers suggest the cheapest system can be the best system; that flips procurement logic on its head.
- If you publish: demand episode-level methods sections — what the system hypothesized, measured, and discarded — not just final benchmark tables.
The exams are ending. The discovery episodes — and the cost columns — are where AI science will actually be won.
Sources: "Measuring AI Scientists: From Exams to Discovery" (Perspective, 2026, chemrxiv); MIRA: Towards a General AI Scientist (Deep Principle, mira.deepprinciple.com); external evaluations on ResearchClawBench and SciAgentArena, August 2026.