Phipps Eval

PE Energy AI Benchmark
Cloud benchmark

Hosted model quality and economics

The 33-task PE benchmark compares cloud/API models on one artifact-backed board. Local PE runs use the older harness and stay available as a separate legacy view—not mixed with production qualification.

33 tasks · 3× Finals highlighted
Overall /100 50% Phipps behavior + 35% broad capability + 15% BFCL tool policy. Missing components are re-normalized and labeled with coverage; hard failures remain disqualifying outside the weighted score.
Speed vs Overall /100