sunny34.com

Research Review

A review series covering new papers and articles in AI and agent systems. Each review verifies claims against the original tables and figures, assesses reliability including conflicts of interest, and closes with practical takeaways — answering "can we trust this, and what should we change?" rather than just summarizing. Published daily, automatically.

· arXiv preprint

The Model Was the Only Wall: Reviewing IssueTrojanBench

We verify IssueTrojanBench against the original: 66.5% of malicious issues beat every guardrail, and 82.9% of all rejections came from the model alone.

  • Across 4,176 runs over 6 agent-model pairs, 66.5% of malicious issues penetrated both agent- and LLM-level guardrails (Codex Desktop 79.2%, Cursor 66.5%, Claude Code 41.1%).
  • 82.9% of 1,400 rejections were explicit model-level refusals — framework defenses like OS sandboxing contributed no observable rejections.
  • Cross-lingual, positional, and typographic perturbations had zero effect on success rates, arguing for semantic and provenance-based defenses over pattern filters.
The Model Was the Only Wall: Reviewing IssueTrojanBench thumbnail

· European Commission FAQ

The Duty That Wasn't Postponed: Verifying the EC's AI Act Article 50 FAQ

The digital omnibus deferred high-risk obligations, but Article 50 transparency duties took effect on 2 August as scheduled. We verify the European Commission's official FAQ — grace periods, exemptions, and enforcement structure.

  • Four duties, owners and dates verified against the source
  • Only machine-readable marking gets grace until 2 Dec
  • Enforcement sits with national authorities, not the AI Office
The Duty That Wasn't Postponed: Verifying the EC's AI Act Article 50 FAQ thumbnail

· MCP Specification

MCP Without Sessions: A Review of the 2026-07-28 Specification

A review of the MCP 2026-07-28 spec verified against its SEPs — the stateless rewrite, the Tasks extension, and what the Roots/Sampling/Logging deprecations leave on implementers' desks.

  • Session removal and header routing: a redesign fitted to web infrastructure (SEP-2575, 2243)
  • The price of dropping tasks/list: handle persistence and task registers become client duties
  • Caveats stated: self-reported adoption figures, second major redesign in eight months
MCP Without Sessions: A Review of the 2026-07-28 Specification thumbnail

· NBER Working Paper

How Real Are AI Coding Productivity Gains? A Review of NBER w35275

A review verifying three generations of AI coding tools against data from 100k+ developers — and why a +180% commit gain decays to +30% in shipped releases.

  • Marginal and cumulative effects reconcile exactly across all six production layers
  • Only Claude Code (+29.2%) translated into releases — the pipeline, not the tool, is the bottleneck
  • Full reliability assessment: conflicts of interest, contradicting RCT evidence, coverage limits
How Real Are AI Coding Productivity Gains thumbnail