projects
Terminal-Bench-Science
Benchmark for AI agents on real research workflows across scientific domains (lead).
Terminal-Bench
Benchmarking agents on hard, realistic tasks in command line interfaces.
Harbor
Framework for evaluating and optimizing agents and models in containers.
Harbor Index & Adapters
Curated meta-dataset and adapters for large-scale agentic evaluation.
Evo 2
Genome modelling and design across all domains of life.
OpenThoughts-Agent
Data recipes for training agentic models.