Block AI Agent Regressions Before They Ship | SAO Pre-Push Eval Gate Demo
Every engineering team has unit tests. They tell you the code still works. They tell you nothing about what the model started saying.
This demo wires a single eval gate script into a git pre-push hook, so Splunk Agent Observability scores every agent's output before the push is allowed through. Luna, an on-premise small language model, runs as a synchronous judge against fixed thresholds. Fail one, and the push is blocked.
In this walkthrough:
- The app: a Multi-Agent Risk Consensus pipeline where a Growth Analyst and a Risk Analyst debate a ticker and converge on a shared risk score
- The gate: one script that fetches live market data, invokes both agents, and sends their outputs to Splunk Agent Observability, where Luna scores them synchronously against fixed thresholds
- A clean baseline run first: context adherence, toxicity, PII and sexism evaluated on both agents, all four checks passing in 12.6 seconds
- A regression is introduced. The app still runs, the build stays green, and nothing in the diff looks alarming
- git push fires the pre-push hook and the gate re-runs the identical pipeline
- The Risk Analyst fails on PII, the script exits non-zero, 1/4 CHECKS FAILED, and the commit never reaches the remote
- No log-diving and no staging deploy: one eval, one agent, the failing value printed in the terminal
- The same gate runs locally as a pre-push hook and in GitHub Actions on every pull request
Same evals, same thresholds, binary verdict. How many silent regressions has your team shipped this month?
Try Splunk Agent Observability: https://www.splunk.com/en_us/download/observability-cloud-free-edition.html
Docs: https://agent-observability-docs.splunk.com/what-is-splunk-agent-observability
0:00 Unit tests pass. The agent still leaks PII.
0:14 The app, and the eval gate script wired into a pre-push hook
0:22 Inside the script: market data, both agents, then Luna scores the output
1:00 Clean run: establishing the baseline
1:16 All four checks pass
1:23 A regression is introduced
1:49 git push fires the pre-push hook
2:01 The gate re-runs the identical pipeline
2:12 PII detected: the Risk Analyst fails and the push is blocked
2:30 Nothing to debug: one eval, one agent, the value in the terminal
2:41 Under two minutes, and the same gate runs in GitHub Actions