FrontierChallenge
A benchmark for whether agent systems can complete specified scientific workflows and deliver a checkable artifact bundle.
| Rank | Model | Agent Scaffold |
|---|
13 model–scaffold configurations · 97 tasks · Historical table: Pass Rate used native Score ≥ 99.9. The current evaluation policy requires completed task_score > 0.999 (strictly greater, without rounding); these historical numbers have not been recomputed.
Selected cases