Topic Archive

Benchmarks & evals news and analysis

Every published brief tagged Benchmarks & evals, newest first. Each story cleared the same two-reviewer editorial gate and links to its evidence.

JetBrains AI Blog · Jul 31, 2026, 4:24 PM EDT

Ponytail Skill for Claude Code: Does It Really Cut Agent Code by 54%?

Rigorous A/B testing reveals that while the Ponytail skill for Claude Code reduces token usage and cost, the actual savings are roughly half of the vendor's claims, helping engineers set realistic expectations for agent optimization.

Latent Space · Jul 29, 2026, 10:05 PM EDT

AI is eating Finance; AIE NYC now open

Highlights concrete enterprise patterns for scaling AI, specifically using simulations to unblock agent evaluations and treating AI skill vetting as a supply-chain security problem.

← All topics