Be the agent orchestrator, not the bottleneck.
Any harness can run the inner loop. The outer loop, your context, your dbt models, your evals, is what keeps agents honest. That is where your judgment matters most. SignalPilot does the writing and the checking; you set the standard.
Two loops
The inner loop is easy. The outer loop is the job.
A harness answers in seconds. Whether it is right depends on everything around it.
Claude Cowork, Claude Code, Codex, whatever ships next. Plan, query, answer. Fast, commoditized, not the hard part.
Context, dbt models and tests, trusted dashboards, skills, evals. It feeds the harness and checks it. Get it right and the agent has almost no way to be wrong.
What each metric means and how it is calculated.
The marts, the grain, the tests. Changes arrive as pull requests.
The numbers leadership believes, used as gold values.
How the work gets done in your stack.
Graded questions, rerun on every change. What keeps it honest.
Running the outer loop
Be the center of gravity of your company's data estate.
You decide what a metric means and what a right answer looks like. SignalPilot authors the context from your stack, generates the evals, and gates every change against them.
Context authored for you, approved by you
The agent reads your schemas, dbt models and tests, trusted dashboards, and the queries behind them, then writes what each metric means and how it is calculated.
- Metric definitions: meaning, calculation, sources, exceptions.
- Required changes to dbt models, proposed as pull requests.
- Business context in Markdown and Apache Ossie, next to the models, readable by your BI tools.
Evals keep it honest on every change
On every pull request the affected evals rerun, anchored in a gold value and the approved definition. Nothing merges until they pass.
- Only the evals the change touches run, and the blast radius names the models and dashboards affected.
- Failures point to the exact questions and definitions that need attention.
- Thumbs-downs and new questions come back as proposed context and regression tests, reviewed before they join.
feat: refund reason on int_orders_with_refunds #212
signalpilotproposed · feat/refund-reason → main
- add refund_reason to int_orders_with_refunds
- fix: keep 1:1 grain on order_id · signalpilot
- ✓Dana R. approved these changes
The loop learns from every correction
A thumbs-down, a new question, or a model that drifted becomes a proposal in your queue: the context to change and the regression test to add. You approve, and the standard gets stronger.
- Every proposal carries its trace: the thread, the queries, and the eval result that motivated it.
- Nothing changes behind your back. Proposals merge as pull requests you review.
- Precision on the frozen suite, trailing 30 days, so you can see the loop getting better.
What changes in your week
Less translating. More approving.
Approve, don't rewrite
Business questions stop landing on your desk as SQL requests. You approve definitions and settle the conflicts the agent surfaces.
Read the eval report
Every pull request comes with the graded questions it touched and whether they passed. That replaces checking agent answers one by one.
See the blast radius first
A model change shows which dashboards and questions it affects before it ships, not after an exec notices.
Stays in your control
Your repo, your stack, your approval.
Two numbers on one dashboard: how many of your models have graded questions, and how often those questions passed in the last 30 days. Together they tell you whether the context can be trusted and what to cover next.
- Snowflake
- Databricks
- BigQuery
Postgres
DuckDB
- Redshift
- dbt
- GitHub
- Claude Cowork
Codex
Slack
One governed MCP gateway with an audit log.
Open source, Apache 2.0. Every benchmark transcript public. Self-host with Docker Compose. Definitions in Apache Ossie, so they leave with you.
FAQ
Questions data leads ask first.
Short answers. The long ones are a demo away.
Do we have to document everything first?
No. The agent starts from your dbt models, dashboards, and the queries people actually run, and drafts the definitions. You review what it wrote and settle the conflicts it surfaces, like two dashboards with different definitions of revenue.
We already have a semantic layer.
Keep it. We read the dbt Semantic Layer, Cube, or LookML as an input and write Apache Ossie beside it, so the definitions work in your BI tools and any other agent.
What happens when the agent does not know?
It says so. Questions outside coverage are refused, not guessed, and the gap lands with you as a proposed eval. Refusals are never billed.
How much of our time does this take?
One hour with you on day one of the integration week, and review time on day three. After that, the weekly review of proposals, or hand that to Managed.