NOIR
Guides

Testing agents

Test deterministic runtime behavior, provider boundaries, policy, retries, and model quality without making every test live.

Separate runtime correctness from model quality. Most reliability behavior is deterministic and should be tested without a real model or provider.

Unit boundary

Test schema parsers, authorization callbacks, output projection, schedule timing, path containment, and policy matching as ordinary functions. Reject malformed model and provider values explicitly.

Runtime integration

Use the in-memory store and fake adapters to test a complete turn:

const channel = fakeChannel()
const agent = defineAgent({
  id: 'test-agent',
  channels: { test: channel.adapter },
  async handle(message, noir) {
    return noir.run('answer', () => ({ text: message.text.toUpperCase() }))
  },
})

await agent.receive(messageEvent('hello'))
await agent.drain()
await agent.flushOutbox()
expect(channel.sent).toHaveLength(1)

Send the same event twice and assert one response. Simulate a process failure after the external write but before checkpoint settlement. Race two workers against the same queue. Expire a lease and confirm takeover. Deny an approval and confirm the tool never executes.

Contract tests

Custom channel and store adapters need reusable contract suites. Run them against the real implementation, not mocks of the implementation. Provider-specific tests can use recorded sanitized fixtures for webhook payloads and API errors.

Evals

Use @noir-agent/agent/evals for behavior that depends on a model: correct tool choice, argument quality, evidence use, refusal, and output usefulness. Keep mock-world cases deterministic. Mark live cases explicitly and run them only with credentials and budgets.

Release matrix

At minimum test:

  • happy path and no-tool response
  • malformed tool arguments
  • repeated identical tool failures
  • approval, denial, and expiration
  • duplicate webhook and duplicate outbox claim
  • provider 429 with retry timing
  • model timeout and user-visible failure
  • oversized tool result and output
  • restart with queued inbound and outbound work
  • scope isolation between two users and installations

A passing prompt snapshot is not a reliability test. Assert observable behavior and settled effects.

On this page