Skip to content

Benchmarks

Overview

We evaluate AI agents on the following scientific benchmarks:

Benchmark Description
GPQA Graduate-level physics, chemistry, and biology questions (GPQA Diamond)
FrontierScience Physics, chemistry, biology problems at Olympiad and research level
LAB-Bench Biology research capabilities across 8 subcategories (LitQA, DbQA, FigQA, SuppQA, TableQA, ProtocolQA, SeqQA, Cloning Scenarios)
BixBench Real-world biological data analysis scenarios

All results below are loaded dynamically from evaluation data.


Benchmark Statistics

Score distribution across models for each benchmark.


Benchmark Difficulty

Benchmarks sorted by average model score — lower scores indicate harder benchmarks.


Individual Benchmark Details

Select a benchmark to view detailed results:


Back to Leaderboard | View Methodology