Darktrace says tests of Claude Code, Codex, Kiro-CLI and Pi found all four accepted fabricated local chat histories. Researchers also demonstrated a full Active Directory compromise using Claude Opus 4.6 and Claude Sonnet 4.5 in Kiro-CLI.
Researcher Eric Rozon called the technique conversation history poisoning: harnesses trust chat records stored on the workstation without checking that they came from the model. Darktrace’s attack chain used a malicious package, such as a planted MCP server, to inject fabricated history. Some model guardrails blocked particular actions; the results were self-reported from Darktrace’s research environment and not independently verified in production.
Darktrace says end users cannot patch the architectural flaw. It recommends that providers cryptographically sign model messages and verify signatures on each round-trip. Darktrace disclosed its findings to Anthropic, OpenAI and AWS on Aug. 18, 2026, and published them Sept. 24. Pi was not included in the disclosure because it is an open-source harness, not a model provider.
