Original Arxiv Paper: https://arxiv.org/abs/2512.15567
Executive Summary
This document synthesizes findings from a comprehensive evaluation of Large Language Models (LLMs) in the context of scientific discovery. The analysis introduces a novel evaluation framework, the Scientific Discovery Evaluation (SDE), designed to address the critical shortcomings of existing science benchmarks. Current benchmarks primarily test decontextualized knowledge through Question & Answer (Q&A) formats, failing to capture the iterative reasoning, hypothesis generation, and evidence interpretation central to actual scientific research. The SDE framework provides a more realistic and rigorous assessment by grounding evaluations in expert-defined research projects and scenarios across biology, chemistry, materials science, and physics.
Key Takeaways:
1. Significant Performance Gap: State-of-the-art LLMs consistently score lower on the SDE benchmark's discovery-oriented questions compared to their performance on general-science Q&A benchmarks, revealing that proficiency in static knowledge recall does not translate directly to discovery readiness.
2. Diminishing Returns on Scaling: The performance gains from increasing model size and applying more test-time reasoning are plateauing for scientific discovery tasks. This suggests that simply scaling current methodologies may not be sufficient to bridge the gap to genuine scientific reasoning.
3. Shared Systematic Weaknesses: Top-tier models from different providers (e.g., OpenAI, Anthropic, Grok, DeepSeek) exhibit highly correlated performance profiles, frequently succeeding and failing on the same research scenarios and questions. This indicates that their limitations may stem from common pre-training data and objectives rather than unique architectural differences.
4. No Single "Superintelligent" Model: Model performance varies drastically across different scientific projects and scenarios. No single LLM consistently outperforms others, demonstrating that current models are far from a general scientific "superintelligence" and that the "best" model choice is context-dependent.
5. Serendipity Outweighs Rote Knowledge: In project-level evaluations, LLMs can achieve successful outcomes even in areas where their question-level accuracy is low. This highlights the critical role of guided exploration and serendipity, suggesting that an LLM's ability to navigate a hypothesis space can be more important than its explicit, granular knowledge.
The Problem with Existing LLM Science Benchmarks
Current popular science benchmarks for LLMs, such as GPQA and MMMU, have been instrumental in tracking model progress. However, they are fundamentally misaligned with the realities of scientific discovery. Their primary limitations include:
• Decontextualized Q&A: The benchmarks consist largely of isolated, quiz-style questions that test static knowledge, similar to academic coursework. As the source notes, "earning straight A’s in coursework does not indicate a great researcher."
• Neglect of Core Scientific Processes: They overlook the essential skills of scientific inquiry, such as iterative reasoning under imperfect evidence, hypothesis generation and refinement, experimental design, and the interpretation of observations.
• Data Quality Issues: Existing benchmarks are often susceptible to label noise and can contain questions that are either irrelevant to practical research or have incorrect ground-truth answers.
This disconnect means that high performance on these benchmarks is not a reliable predictor of an LLM's utility as a tool for actual scientific research and discovery workflows.
The Scientific Discovery Evaluation (SDE) Framework
To address these shortcomings, the SDE framework was developed as a systematic, scenario-grounded evaluation paradigm.
Core Principles and Structure
The SDE framework is built on a hierarchical structure that ensures every evaluation is tied to a realistic research context:
1. Projects: The foundation consists of concrete research projects of genuine interest to domain experts (e.g., discovering new pathways for artemisinin synthesis, optimizing transition metal complexes).
2. Scenarios: Each project is decomposed into modular, reusable research scenarios. A scenario is a self-contained scientific reasoning unit, such as "forward reaction prediction" or "structure elucidation from NMR spectra."
3. Questions: Within each scenario, expert-vetted questions are constructed. These questions, formatted for automated evaluation (multiple-choice or exact match), serve as measurable indicators of progress toward a specific discovery goal.
This "tight connection among questions, scenarios, and projects" ensures that the SDE benchmark faithfully assesses an LLM's capabilities in a discovery-relevant context. The dataset comprises 1,125 questions across 43 distinct scenarios in four domains.
Two-Level Assessment
The SDE framework evaluates LLMs at two distinct levels:
• Question-Level: Measures the accuracy of models on specific, scenario-tied questions, providing a fine-grained view of their strengths and weaknesses.
• Project-Level: Assesses a model's ability to participate in an end-to-end discovery loop. Using the sde-harness software, an LLM is tasked to autonomously propose testable hypotheses, run simulations via computational oracles, and interpret the results to refine its next set of hypotheses.
Key Findings from Question-Level Evaluation
Performance Gap Between Discovery and Quiz-Style Questions
A consistent gap exists between LLM performance on SDE and general-science benchmarks. For instance, gpt-5 achieves a score of 0.86 on GPQA-Diamond but only 0.75 on SDE-materials and 0.60 on SDE-physics. This demonstrates that discovery-oriented questions, even in a Q&A format, present a greater challenge than decontextualized trivia.
Plateauing Gains from Scaling and Reasoning
While both model size and test-time reasoning contribute to better performance, their benefits are showing diminishing returns on SDE tasks.
• Reasoning: Models with explicit reasoning capabilities consistently outperform non-reasoning counterparts (e.g., deepseek-R1 over deepseek-V3.1). However, for top models like gpt-5, increasing reasoning effort from "medium" to "high" yields negligible accuracy improvements (e.g., 0.74 vs. 0.75 in materials).
• Scaling: Performance improves as models scale (e.g., from gpt-5-nano to gpt-5), but the pace of improvement has slowed. The performance gain of gpt-5 over its predecessor o3 was marginal, indicating a potential convergence for pretrained foundation models.
Shared Failure Modes Across Frontier Models
Top-performing models from different providers—gpt-5 (OpenAI), grok-4 (xAI), deepseek-R1 (DeepSeek), and claude-sonnet-4.5 (Anthropic)—exhibit highly correlated performance.
• Correlated Accuracy: The models tend to perform well on the same scenarios and fail on the same ones. The pairwise Spearman's rank correlation is greater than 0.8 in chemistry and physics.
• Convergent Errors: On the most difficult questions, these models frequently converge on the exact same incorrect answer. This suggests their weaknesses are systematic, likely inherited from similar pre-training data and objectives, and cannot be easily overcome with simple ensembling methods.
To probe these limits, the SDE-hard subset was created, consisting of 86 questions where top models consistently fail. On this set, most models score below 0.12. Notably, gpt-5-pro shows a significant improvement, correctly answering questions that all other models failed, highlighting its advantage in tasks requiring extended reasoning.
Key Findings from Project-Level Evaluation
Simulating the Scientific Discovery Loop
The project-level evaluation moves beyond static Q&A to assess LLMs as active participants in the scientific discovery loop of Hypothesis -> Experiment -> Observation. Eight projects were established, including protein design, retrosynthesis, transition metal complex (TMC) optimization, and symbolic regression.
The Disconnect Between Question- and Project-Level Success
A model's accuracy on related questions does not always predict its success in a discovery project. This reveals that different capabilities are required for each evaluation level.
• Negative Correlation Example (Retrosynthesis): Top LLMs score highly on questions about retrosynthesis but struggle to generate valid multi-step synthesis routes in the project-level evaluation, failing to outperform traditional models.
• Positive Correlation Example (TMC Optimization): Conversely, while no model demonstrates high proficiency on TMC-related questions (e.g., predicting oxidation or spin states), several models (gpt-5, deepseek-R1, claude-sonnet-4.5) perform exceptionally well in the TMC optimization project. They rapidly identified optimal candidates from a search space of 1.37 million TMCs, often in fewer than 10 iterations.
The Role of Serendipity in LLM-Driven Discovery
The success in the TMC project, despite poor scores on related questions, suggests that "rigorous knowledge of explicit structure-property relationships is not a strict prerequisite for LLM-driven discovery." Instead, the capacity to discern optimization directions and facilitate serendipitous exploration appears more critical. LLMs can effectively navigate a hypothesis space even with imperfect granular knowledge.
Conclusions and Future Directions
The SDE framework demonstrates that evaluating LLMs for scientific discovery requires a paradigm shift away from static knowledge tests toward dynamic, context-aware assessments. The findings indicate that while current LLMs show promise, they are approaching a performance plateau under existing development strategies.
Based on these results, several key directions are identified for advancing LLMs in science:
1. Targeted Training: Shift focus from indiscriminate scaling to targeted training on scientific methodology, including problem formulation and hypothesis generation.
2. Data Diversification: Diversify pre-training data sources and explore novel inductive biases to mitigate the shared failure modes observed across current models.
3. Tool Integration: Deeply integrate robust tool use into fine-tuning, as many research scenarios require coupling linguistic reasoning with domain-specific computational libraries and simulators.
4. Scientific Reinforcement Learning: Develop reinforcement learning strategies specifically tailored for scientific reasoning, as enhancements optimized for mathematics and coding have shown limited transfer to discovery-oriented projects.