Terminal-Bench-Science: Evaluating AI agents on scientific research workflows

Terminal-Bench · since 2026

Benchmark for evaluating AI agents on real scientific research workflows in terminal environments.

PricingFree
LevelAdvanced
CategoryScience & Biotech
Best forAI researchers and ML engineers

Tags: benchmark, ai agents, science, evaluation, research

Visit Terminal-Bench-Science: Evaluating AI agents on scientific research workflows

Alternatives to Terminal-Bench-Science: Evaluating AI agents on scientific research workflows

Build your personal AI stack on Toolnaut