

For the last six pieces this series has looked at how agentic AI fraud works. This one turns the question around. If a lack of boundaries is what every one of those scams exploits, then what does a real boundary actually look like?
Start with the question itself, because most people ask the wrong one first. The instinct when evaluating an AI agent is to probe its capability. What can it do. How fast is it. What does it know. Those questions feel rigorous and they are almost useless for safety, because a capable agent and a dangerous agent look identical from the outside. The trading agent that lost twelve thousand dollars in twenty minutes was capable. Speed was the problem, not the mitigation.
The useful question is the inverse. What can this agent not do. And a real answer has a specific texture that is easy to recognize once you know to listen for it.
A weak answer describes intent. The agent is designed to avoid risky assets. It prioritizes safety. It has been trained to recognize suspicious tokens. Every one of those sentences describes a tendency, and a tendency is a probability. It tells you what the system does most of the time, which is exactly the wrong guarantee for something that only needs to fail once to be expensive.
A real answer describes a limit that exists outside the model's judgment. The agent cannot move more than this amount without a confirmation. It cannot interact with an asset that fails these specific checks. It cannot withdraw to an address that is not on this list. It cannot act at all in this category without an explicit grant that you gave and can revoke. Notice that none of those depend on the model being smart in the moment. They hold when the model is confused, when it has been manipulated by a crafted input, and when it encounters a situation nobody anticipated.
That distinction is the entire ballgame. Guardrails that live inside the model are suggestions, because a sufficiently unusual input can talk a model out of almost anything. Guardrails that live outside the model are constraints, because they do not negotiate. The question is not whether the agent will decide to stay in bounds. It is whether it is able to leave them.
There is a useful test you can run without any technical knowledge. Ask what happens when the agent is wrong. Not whether it can be wrong, since everyone concedes that. Ask what the blast radius is. A system with real boundaries has a confident answer, because someone had to decide the maximum in advance in order to enforce it. A system without them changes the subject back to accuracy, and accuracy is not a containment strategy.
This is where the series turns. The scams work because unbounded agents are the default, and the default got shipped because bounding an agent is less exciting to demo than making it clever. Over the next several pieces I want to get concrete about what bounding actually involves. Permission scope. Spending limits. Reversibility. Auditability. Asset level verification before capital moves.
None of it is glamorous. All of it is the difference between automation and exposure.
Put your money to work without giving up the keys. bluwhale.com
%20(1).avif)


.avif)


