NOIR
Capabilities

Operations and quality

Observe, cancel, budget, evaluate, compare, review, and roll back agent behavior with explicit self-hosted state.

Operational capabilities should describe what happened without becoming another source of truth for the conversation. They consume structured runtime events and never get to change a successful outcome silently.

Observability

@noir-agent/agent/observability records versioned traces and spans for runs, models, tools, retrieval, and delivery. sanitizeObservation() removes or bounds sensitive fields before export. Exporter failure is isolated from the user turn.

Every started trace and span must settle on success, failure, cancellation, timeout, or supersession. Memory export is useful in tests; postgresTraceExporter() provides durable bounded queries.

Use @noir-agent/agent/models/prompt-cache to place provider cache controls at stable prompt boundaries. Treat cache configuration as a model optimization, not application state.

Run control

@noir-agent/agent/run-control persists owner-scoped cancellation requests. Runtimes and long operations check the current request through their abort signal. A canceled or superseded run cannot deliver user-visible output even if downstream work finishes later.

Usage budgets

@noir-agent/agent/usage atomically reserves run, tool-call, and model-token budgets before work. Successful model calls settle with provider-reported usage. When a provider omits usage, Noir records the configured reservation as reserved_fallback; it does not label an estimate as measured usage.

Repeated idempotency keys reuse the same decision. Failures release their lease. Actual usage above a reservation is recorded as overage, not discarded.

Output guards

@noir-agent/agent/guards detects repeated tool failures, leaked deliberation, repetitive output, raw tool-result echoes, unsupported promises of future work, and untrusted payment URLs. Compose guards explicitly and decide whether a finding blocks, retries, or reports the output.

Guards reduce known failure modes; they do not replace tool-result validation or product-specific review.

Evals and model sweeps

@noir-agent/agent/evals runs bounded eval suites and model sweeps. Checks cover completion, tool calls and order, required or forbidden text, code, and model-judged criteria. Mock worlds never call live tools. Live worlds use the same capability path as production.

Winner selection is explicit. Never loosen an assertion to make a model pass.

Quality proposals

@noir-agent/agent/quality stores reviewable proposals tied to failing traces or eval evidence. A proposal cannot activate itself. Your application supplies the review and cutover surface.

@noir-agent/agent/versioning stores immutable configuration versions, supports explicit rollback, and runs shadow comparisons. Shadow output is observed but never delivered. Cutover remains a deliberate host action.

Testing helpers

@noir-agent/agent/testing provides fixtures and deterministic helpers for channel, runtime, store, and capability tests. @noir-agent/agent/harness is for an external model runtime, not a replacement for integration tests.

Release gate

Before production, verify:

  • duplicate ingress causes one run
  • completed effects are not repeated after restart
  • every trace settles on all terminal paths
  • redaction happens before export
  • cancellation blocks final delivery
  • budget admission is atomic across workers
  • provider usage and reserved fallback are distinguishable
  • mock evals cannot call a live provider
  • shadow output cannot reach a channel
  • rollback restores an immutable reviewed version

Operations state belongs in your database and observability systems. Noir supplies the contracts; it does not require a hosted control plane.

On this page