A payment agent is easy to start and hard to trust. The demo version can call an endpoint and move money. The version you would let near a real balance needs something the demo does not have: guardrails that hold even when the model is wrong, confused or manipulated. The useful pattern is to place those guardrails outside the agent.

The rule: policy below the model

A guardrailed agent does not decide its own limits. Limits, approved counterparties and approval thresholds are configured below the model, in the layer that executes operations. The agent can propose; the perimeter disposes. This is the difference between asking a model to be careful and making carelessness impossible.

Three guardrails carry most of the weight:

  • limits — a ceiling per operation and per period
  • approved counterparties — a closed list the agent may pay
  • approval thresholds — operations above a level become a human request

Identity and mandate

An agent that can transact must be identifiable and revocable. In the SUPA model an agent owns nothing and is never anonymous; it acts for an identified principal. The mandate is registered, so a counterparty can verify an agent's authority in real time rather than trusting a claim.

Making it auditable

Guardrails are only credible if they leave evidence. Every operation is recorded in a double-entry ledger attributed to the principal, the agent and the version of policy in force at the time. If a rule changes, the record shows which rule applied. That is what turns a guardrail from a setting into something you can reconstruct after the fact.

Practical builds usually start with one narrow mandate — a fixed counterparty list and a low ceiling — and widen it only once the evidence looks right. Because balances sit with licensed partners before a banking licence, an agent's reach is limited by what the partner rails permit, which makes early experiments easier to contain.

In short

  • Put policy below the model, where it cannot be argued with.
  • Use limits, approved counterparties and approval thresholds as core guardrails.
  • Make every agent identifiable, revocable and tied to a principal.
  • Record operations against the policy version in force.

Start from the developer surface and build one bounded agent: developers.