Skip to content

The benchmark viewer

Which model actually serves your decisions best? The Benchmark tab answers that — not with someone else’s leaderboard, but with runs scored against Prevail’s canonical suite, in your own app.

Prevail Desktop benchmark viewer — a leaderboard ranking models and council configurations by score across the canonical question suite.
The benchmark viewer — every scored run ranked, so you can see which model or council serves your decisions best.
  • A leaderboard ranking every model and council configuration by score.
  • Per-question drill-down — click any row to see the prompt, each model’s reply, keyword hits, and the judge’s rationale.
  • Council vs single model — see what convening a panel actually buys you over one strong model.

Runs are read straight from <vault>/benchmark/runs/, so the viewer reflects whatever you’ve scored — including the runs already present in demo mode.

You can kick off runs from Settings → Benchmark (or from the CLI). Prevail asks each configured model the canonical question set, scores the answers, and writes the results back into your vault for the viewer to read.

The canonical benchmark · Running & scoring