Benchmarks
Overview
We evaluate AI agents on the following scientific benchmarks:
| Benchmark | Description |
|---|---|
| GPQA | Graduate-level physics, chemistry, and biology questions (GPQA Diamond) |
| FrontierScience | Physics, chemistry, biology problems at Olympiad and research level |
| LAB-Bench | Biology research capabilities across 8 subcategories (LitQA, DbQA, FigQA, SuppQA, TableQA, ProtocolQA, SeqQA, Cloning Scenarios) |
| BixBench | Real-world biological data analysis scenarios |
All results below are loaded dynamically from evaluation data.
Benchmark Statistics
Score distribution across models for each benchmark.
Benchmark Difficulty
Benchmarks sorted by average model score — lower scores indicate harder benchmarks.
Individual Benchmark Details
Select a benchmark to view detailed results: