Traces and alerts
Investigate a single execution in detail and alert on changes that require attention.
Use traces for investigation and metrics for patterns. A trace should answer: what started the run, what it was allowed to see, what it decided, what it called, what policy applied, and why it ended.
Alert on customer-impacting failures, unexpected policy denials, sustained latency, unusual cost, and changes in evaluation quality. Route alerts to an owner with the run or trace identifier needed to investigate. Avoid alerting on every individual model variation; alert on the operational condition that requires a response.
2026年8月31日に更新