SignalPilot LogoSolutions · Model & vendor bakeoff

Judge every model and vendor on your data.

We build the evals once from your stack, then run every candidate through them: Claude, GPT, Gemini, warehouse agents, BI agents. You get pass rates per question, evidence for every answer, and a decision you can defend.

Why a bakeoff needs evals

Public benchmarks do not run on your data.

Our public results, for reference. Every team we talk to asks the second question, and only your own graded questions answer it.

#1

on both public benchmarks

Spider 2.0-DBT and ADE-Bench

96.9%

ADE-Bench, tasks passed

Claude Code alone: 39.5%, same model

75%

Spider 2.0-DBT, tasks passed

next best 60.29%

ADE-Bench · % of tasks passed

Same model, with and without SignalPilot.

SignalPilot
96.9%
Snowflake CoCo
65.1%
dbt Wizard
58.1%
Claude Code alone
39.5%

Both runs: Claude Code with Sonnet 4.6 · every transcript public →

Spider 2.0-DBT · % of tasks passed

The hardest public data-engineering benchmark.

SignalPilot*
75%
Databao · JetBrains
60.29%
Shadowfax
41.18%
Spider Ext · GPT-5
39.71%

*75% submitted in August; official leaderboard shows 65.6% (May).

#1 on the public leaderboard ↗ · task-by-task results →

What you get

The same evals, every candidate.

  • Evals from your own dbt models and dashboards, including the failure cases.
  • Each candidate run in the same governed sandbox with the same context.
  • A report: pass and fail per question per candidate, with the evidence graph for every answer.
  • The context layer and evals stay with you afterward, whichever vendor you pick.

Model routing

The best model per task, not one model for everything.

The report gives you pass rates per task group and cost per task for every candidate. Route each group to the cheapest model that clears the bar and the expensive model only runs where it earns its keep.

SAME QUESTIONS, EVERY MODEL

Every candidate answers the same graded questions in the same sandbox with the same context, so the pass rates are comparable and the cost per task is real.

A BAR, NOT A WINNER

You set the pass rate a task group must clear. The cheapest model above it takes the slot. Failure cases usually stay with the strongest model; routine groups rarely need it.

RERUN WHEN MODELS SHIP

The evals and the context stay in your repo. When a new model lands, rerun the grid and move the slots that changed.

TIMELINE

About two weeks from access to report for a scoped set of questions.

WHAT STAYS

The evals and the context layer are yours. Rerun the bakeoff when the next model ships.

Run the bakeoff on your stack.

Open source · Apache 2.0 · GitHub · Slack