Define the scope
Create an alert for a project or narrow it by environment, agent, model, customer group, or release. Scoping should match the team that can respond.
Select a signal
Reliability alerts support error rate, latency percentiles, tool failure, run outcome, and throughput. Budget alerts support absolute spend, forecast spend, and cost per successful outcome.
Set the evaluation window
Use a short window for sharp production failures and a longer window for cost or outcome drift. Add a minimum traffic threshold so low-volume workloads do not produce noisy percentages.
Route the alert
Send notifications to Slack, PagerDuty, email, or a webhook. Every notification includes the active release, affected dimensions, and representative traces.
Review alert history after a deployment to confirm the policy reflects customer impact and remains actionable.