A friend who works in risk for an AI trading fund asked me a tough question: if you give an AI agent the power to automatically trade, and prompt injection causes it to take actions against the original intent, how do you stop it—before the money is lost, not after you discover it?

I stayed silent for a moment, because most of the solutions I know only detect the problem after it’s already too late.

Not from the speed of an agent processing a complex command. Not from the number of automation intents running each second. Not from a demo of a smooth trading agent on stage.

A simpler question—if an agent is tricked, injected, or simply misinterprets the user’s intent, is there a hard boundary that stops those actions before it touches real money, or does the system rely only on the agent being “carefully designed”?

That’s the @NewtonProtocol answer with zkPermissions — one of the guardrail conditions explicitly listed is defense against prompt-injection, not just limiting spending or whitelisted recipients.

Designing smarter agents is easy and sounds appealing, because as AI gets better, people are more likely to believe it will always behave correctly.

Create boundaries independent of the agent itself, enforced with cryptography before a trade is allowed to settle—this is what’s truly hard, because it requires defining in advance every scenario where the agent might be tricked, and ensuring that boundary holds firm even if the AI model itself is completely misled.

Newton Protocol clearly lists four types of guardrails for agents: spending limits, pre-approved recipient lists, constraints tied to specific tasks, and defense against prompt-injection. This isn’t a random feature list—it reflects the team’s understanding that AI agents will be attacked in ways that the agent itself may not recognize.

If Newton Protocol can keep these boundaries solid as the number of agents and the complexity of trades increase, then the value of $NEWT s will be tied to the role of an essential safety layer for the agent economy—not just gas fees for each execution.

Self-critique: guardrails against prompt-injection sound reasonable on paper, but I haven’t seen any real-world case where Newton Protocol publicly showed an agent being attacked and the boundaries successfully stopping it—this is still a defensive design, not proof that’s been through live trials.

But if the future of financial AI agents requires an additional layer of permissions limited independently from the AI model itself, then this is a direction worth following more than any other performance number from Newton Protocol.

#newt $NEWT