The 33-task PE benchmark compares cloud/API models on one artifact-backed board. Local PE runs use the older harness and stay available as a separate legacy view—not mixed with production qualification.
The primary board shows only current-manifest 3× Finals and keeps one best-evidenced Overall score per model family. Screens, reasoning alternatives, superseded router builds, and every valid historical lane remain available below.
The Daily Driver board will populate only from valid database receipts. Invalid or incomplete results never enter the ranking.
Current v4 Finals only. All points share the task harness and 3× run shape; token rates still depend on each provider/runtime, so use the chart as an operating trade-off—not a hardware benchmark.
Daily Driver is a separate 36-task suite covering interaction quality, tool judgment, instruction following, memory/context, knowledge calibration, and coding/problem solving. A complete 1× lane is a Screen; only a complete 3× lane is a Final. Screens never rank on the primary board.
Waiting for the signed-off suite export.
Manifest: —