Testing agents
Test deterministic runtime behavior, provider boundaries, policy, retries, and model quality without making every test live.
Separate runtime correctness from model quality. Most reliability behavior is deterministic and should be tested without a real model or provider.
Unit boundary
Test schema parsers, authorization callbacks, output projection, schedule timing, path containment, and policy matching as ordinary functions. Reject malformed model and provider values explicitly.
Runtime integration
Use the in-memory store and fake adapters to test a complete turn:
const channel = fakeChannel()
const agent = defineAgent({
id: 'test-agent',
channels: { test: channel.adapter },
async handle(message, noir) {
return noir.run('answer', () => ({ text: message.text.toUpperCase() }))
},
})
await agent.receive(messageEvent('hello'))
await agent.drain()
await agent.flushOutbox()
expect(channel.sent).toHaveLength(1)Send the same event twice and assert one response. Simulate a process failure after the external write but before checkpoint settlement. Race two workers against the same queue. Expire a lease and confirm takeover. Deny an approval and confirm the tool never executes.
Contract tests
Custom channel and store adapters need reusable contract suites. Run them against the real implementation, not mocks of the implementation. Provider-specific tests can use recorded sanitized fixtures for webhook payloads and API errors.
Evals
Use @noir-agent/agent/evals for behavior that depends on a model: correct tool choice, argument quality, evidence use, refusal, and output usefulness. Keep mock-world cases deterministic. Mark live cases explicitly and run them only with credentials and budgets.
Release matrix
At minimum test:
- happy path and no-tool response
- malformed tool arguments
- repeated identical tool failures
- approval, denial, and expiration
- duplicate webhook and duplicate outbox claim
- provider 429 with retry timing
- model timeout and user-visible failure
- oversized tool result and output
- restart with queued inbound and outbound work
- scope isolation between two users and installations
A passing prompt snapshot is not a reliability test. Assert observable behavior and settled effects.