projects

Terminal-Bench-Science

Benchmark for AI agents on real research workflows across scientific domains (lead).

Terminal-Bench

Benchmarking agents on hard, realistic tasks in command line interfaces.

Harbor

Framework for evaluating and optimizing agents and models in containers.

Harbor Index & Adapters

Curated meta-dataset and adapters for large-scale agentic evaluation.

Evo 2

Genome modelling and design across all domains of life.

OpenThoughts-Agent

Data recipes for training agentic models.