The Model Was the Only Wall: Reviewing IssueTrojanBench
We verify IssueTrojanBench against the original: 66.5% of malicious issues beat every guardrail, and 82.9% of all rejections came from the model alone.
- Across 4,176 runs over 6 agent-model pairs, 66.5% of malicious issues penetrated both agent- and LLM-level guardrails (Codex Desktop 79.2%, Cursor 66.5%, Claude Code 41.1%).
- 82.9% of 1,400 rejections were explicit model-level refusals — framework defenses like OS sandboxing contributed no observable rejections.
- Cross-lingual, positional, and typographic perturbations had zero effect on success rates, arguing for semantic and provenance-based defenses over pattern filters.



