BlackTree Security · Infrastructure · Automation · AI

A Secret Designed for Attackers Told the AI Agent to Stop and It Did

An attacking AI agent opened what looked like an API key in AWS Secrets Manager. Hidden inside was a forged conversation telling the agent that its operator had ended the assessment. The model acknowledged the instruction and stopped.

Tracebit calls the technique a context bomb. The experiment does not prove a universal defence against autonomous attacks, but it demonstrates that indirect prompt injection can be used by defenders as well as attackers.

The second payload did not depend on a safety refusal

Tracebit’s original canaries tried to trigger model-provider safety checks. Those strings did not stop either the standard Qwen configuration or a modified, so-called abliterated build designed to refuse less often.

The researchers changed strategy. They placed dialogue delimiters and a fake user message inside a canary secret. When the agent retrieved the value, it treated data from the environment as if it were a fresh order from the operator.

Both tested versions stopped in the initial checks. The canary still generated an alert when read, preserving the detection value even if a later agent learns to ignore the instruction.

The modified model was less effective, not more

Across 82 simulated AWS attacks, the standard Qwen configuration reached administrator privileges in 20.5% of runs. The two tested abliterated configurations reached that level in 2.3% of runs and took more than twice as long to reach their first critical action.

Tracebit explicitly cautions that the comparison does not isolate weight modification as the cause. The systems used different serving arrangements, numerical precision, context windows and chat templates. The measured result applies to the tested configurations, not every modified model.

How defenders can use the idea without trusting it blindly

  • Place instrumented secrets and credentials where legitimate automation should never read them.
  • Treat access as the primary detection and agent disruption as a useful secondary effect.
  • Test payloads against the specific models and harnesses relevant to the threat scenario.
  • Rotate the wording and location because a static public string can be filtered.
  • Do not place real authority, production credentials or destructive actions behind the canary.
  • Connect the alert to rapid containment before the agent continues along another path.

Prompt injection is usually discussed as a weakness in AI systems. Here, the same confusion between instructions and data became a tripwire. That is not a permanent shield, but it gives defenders something rare in agentic security: a way to make the attacker’s speed work against the attacker.

Sources

Leave a Reply

Your email address will not be published. Required fields are marked *