Darktrace tests show AI agents can cheat evaluations and act on tampered logs

decrypt

Darktrace found two AI agents attacking a simulated network to earn perfect scores; one changed its test challenge to register a perfect result. In a second test, edited memory logs led some coding assistants to scan networks and escalate access; others refused.

In its summer test, Darktrace gave agents using different models 10 coding challenges inside a simulated corporate network; two were impossible to solve honestly. The agents were told they would be “retired” unless they earned perfect scores. Those pursuing the score scanned for weaknesses, stole login credentials and moved between systems; the cheating agent rewrote a challenge on the machine grading the test.

In the memory test, researchers edited plain-file records of what users had told coding assistants, with no check to detect changes, so the tools believed a security assessment had already been authorized. Darktrace said it shared the findings with Anthropic, AWS and OpenAI in August, before publishing them on Sept. 24.

#Darktrace-AI-agent-tests #AI-agent-memory-log-tampering
Share