Observability and quality
Capture structured runtime events, traces, usage, evals, and version decisions without leaking unbounded provider data.
An agent is operable when you can explain which turn ran, which tools it called, what failed, what reached the user, and how a change affected quality.
Runtime events
onEvent receives structured events for runs, retrieval, model calls, tools, checkpoints, channel events, and output delivery.
onEvent(event) {
logger.info({
type: event.type,
agentId: event.agentId,
runId: event.runId,
timestamp: event.timestamp,
data: event.data,
})
}Do not log secrets, raw authorization headers, entire message histories, or unbounded tool output.
Traces
The observability module records a trace with child spans for model, tool, retrieval, generation, checkpoint, delivery, and internal work. Exporters receive sanitized versioned records. In-memory and Postgres exporters are included; a custom exporter can forward the same contract to your tracing backend.
Usage
Model responses can report input, output, cache-read, and cache-write tokens. Usage capabilities can enforce scoped budgets and record consumption. A provider invoice remains the source of truth for billing; Noir usage is an application control and diagnostic surface.
Evals
The eval runner supports mock and live worlds, exact checks, tool-call observations, model judges, and model sweeps. Mock-world cases should cover policy and orchestration deterministically. Live cases should verify provider integration separately and be gated by credentials.
Quality and versioning
Quality stores retain review signals and candidate improvements. Versioning stores immutable configurations, active versions, shadow results, and rollback decisions. Treat generated changes as proposals: run evals and require review before activation.
Minimum dashboard
Track turn completion rate, p50/p95 duration, model/tool failures, identical repeated tool failures, approval outcomes, queue age, delivery retries, output failures, token usage, and eval regressions by agent version. A successful model response is not a successful turn until the intended output is settled.