What to observe
- Runs and traces — the request, plan, tool calls, observations, approvals, failures, and produced artifacts.
- Usage and cost — model and provider usage, tokens, latency, and cost over time.
- Quality — user feedback, eval outcomes, and judge signals.
- Context — the agent, instructions, data scope, files, and tools that informed a run.
Diagnose and improve
When a result needs investigation, start from its trace. Determine whether the root cause is the data, the available context, an instruction, a tool, an access boundary, or the model’s behavior. Then create the appropriate fix and add or update an eval so the same regression is caught again.Governance
Control who can access agents, connections, models, and settings. Use approvals for sensitive tools; audit activity where required; and apply organization controls such as PII protection, identity integration, and service accounts.Runs, traces, and diagnostics
Follow one agent run from request to outcome.
Usage, cost, and quality
Understand adoption, spend, latency, and reliability trends.
