In a chat window, a prompt injection is embarrassing. In a payment system, it is a loss. The difference is not the technique but the consequence: an agent that can move funds is a target whose mistakes have a price. This is a threat-model piece for developers building agents that touch regulated money — not a guide to attacks, but a description of the assumptions a system must be designed to survive.

Why the money layer is different

A general assistant is judged on the usefulness of its output. An agent in finance is judged on the safety of its actions. That changes the goal. The question is not "can the model be made to say the wrong thing?" — it can — but "when the model is fooled, what stops the money?" The answer has to live outside the model, because the model is precisely the thing under attack.

Defence in depth

The design that holds up tends to combine a few principles:

  • policy applied below the model, where the agent cannot argue with it;
  • limits and approved counterparties enforced by the surrounding system;
  • over-limit operations routed to a human approval request rather than executed;
  • a ledger with evidence, recording principal, agent and policy version;
  • revocable, attributable agent identity that is never anonymous.

Each layer assumes the one above it can fail.

Designing for the assumption of compromise

The most useful posture is to assume the model will eventually be manipulated, and to ensure the perimeter is not. Intent can be wrong; authority can still be checked. This is the reason a single legitimacy validation step sits in front of the order book — reliance confirmation, sanctions screening, agent mandate, client-profile fit and network anomaly signals — evaluating every operation once, regardless of how confident the agent sounds.

In short

  • Injections become expensive when the agent can move funds.
  • A manipulated model must meet controls it cannot talk past.
  • Policy below the model and limits enforce the boundary.
  • Assume compromise; make the perimeter the safeguard.

Explore how operations are validated before settlement in the developer docs.