Announcing Harbor-Index
A compact, diverse, and genuinely hard benchmark for agentic evaluation — now live.
We just launched Harbor-Index — a compact, diverse, challenging, and high-quality benchmark for agentic evaluation. It distills 82 carefully vetted tasks from thousands of candidates so you can get a trustworthy read on agent capability without running dozens of separate benchmarks. It’s built on Harbor, the open-source evaluation harness I help maintain.
Read the full announcement, browse the leaderboard, and dig into individual agent trajectories at harbor-index.org.