Phipps Eval

Model evidence, coverage, and runtime
How to read this ↗
Measured, not assumed · One-run Screens

API model evaluations

Quality, speed, and the failures behind the score. These are unranked measurements, not production approvals. Compare within a cohort; inspect the configuration before drawing conclusions.

Loading Rapid measurements…
Full current receipts and explicitly linked attempt history
Loading published component evidence…
Cloud benchmark

Hosted model quality and economics

The 33-task PE benchmark compares cloud/API models on one artifact-backed board. Local PE runs use the older harness and stay available as a separate legacy view—not mixed with the cross-suite comparison.

33 tasks · 3× Finals highlighted
Source score /100 Full Overall contract: 50% Phipps behavior + 35% broad capability + 15% BFCL tool policy. Current legacy rows contain Phipps-only 50% coverage; their numeric source order is retained for historical comparison, not qualification, and missing broad/BFCL components receive no invented Overall. Source validity failures remain disqualifying.
Speed vs Overall /100

Model evidence