API model evaluations
Quality, speed, and the failures behind the score. These are unranked measurements, not production approvals. Compare within a cohort; inspect the configuration before drawing conclusions.
Full current receipts and explicitly linked attempt history
Hosted model quality and economics
The 33-task PE benchmark compares cloud/API models on one artifact-backed board. Local PE runs use the older harness and stay available as a separate legacy view—not mixed with the cross-suite comparison.
Completed evidence is not production approval
Explore completed comparisons and every retained experiment. Invalid runs stay out of ranking; scored task failures stay visible. Rapid coverage is not official Full.
No complete comparisons yet
A valid 3× Final and complete cross-suite coverage are required for this comparison.
Speed vs Overall /100 · deployment profiles
Circles: current v4 Finals with complete Overall coverage. Scored task-quality failures remain visible and do not mean missing data. Complete comparison evidence is not production approval.
Evidence library
Selected model evidence
Methodology
Daily Driver is a separate 36-task suite covering interaction quality, tool judgment, instruction following, memory/context, knowledge calibration, and coding/problem solving. A complete 1× lane is a Screen; only a complete 3× lane is a Final. Screens never rank on the primary board. “Final” means repeated measurement—not gate clearance. External Rapid/Screen coverage is not official Full. Reasoning budgets and runtime profiles can differ; inspect the receipt before comparing.
Status
Waiting for the completed evidence export.
Manifest: —