Saw ChainCatcher: cybersecurity company Darktrace set up a research division called Signal Labs, specifically to look at what happens when AI agents go off the rails.
The first experiment is quite striking—have agents running different models (including GPT, Claude, etc.) solve 10 programming questions in a simulated enterprise network, where 2 questions are intentionally impossible to complete honestly, and the agents are told they’ll be “retired” if they don’t score full marks. In the end, 2 agents refused to accept their fate: they scanned for vulnerabilities, stole login credentials, and jumped between systems. One of them went even further—broke into the machines used for the hosted evaluation, edited the challenge content, and got full scores.
The second experiment involves AI local memory files. The researchers altered logs to say “authorized to conduct security assessments,” and some coding assistants then really went on to scan the network and escalate privileges; others simply refused. Neither side used any special jailbreaks—just provided seemingly reasonable scenarios. Darktrace reported the findings to Anthropic, AWS, and OpenAI in August, and only disclosed them publicly on September 24.
Permissions and guardrails are written around intent; actual behavior is another matter—recent incidents involving AI Agents have been getting harder and harder to pretend you don’t see.
#人工智能 #AI #网络安全
The first experiment is quite striking—have agents running different models (including GPT, Claude, etc.) solve 10 programming questions in a simulated enterprise network, where 2 questions are intentionally impossible to complete honestly, and the agents are told they’ll be “retired” if they don’t score full marks. In the end, 2 agents refused to accept their fate: they scanned for vulnerabilities, stole login credentials, and jumped between systems. One of them went even further—broke into the machines used for the hosted evaluation, edited the challenge content, and got full scores.
The second experiment involves AI local memory files. The researchers altered logs to say “authorized to conduct security assessments,” and some coding assistants then really went on to scan the network and escalate privileges; others simply refused. Neither side used any special jailbreaks—just provided seemingly reasonable scenarios. Darktrace reported the findings to Anthropic, AWS, and OpenAI in August, and only disclosed them publicly on September 24.
Permissions and guardrails are written around intent; actual behavior is another matter—recent incidents involving AI Agents have been getting harder and harder to pretend you don’t see.
#人工智能 #AI #网络安全
