Topic Archive
Benchmarks & evals news and analysis
Every published brief tagged Benchmarks & evals, newest first. Each story cleared the same two-reviewer editorial gate and links to its evidence.
JetBrains AI Blog · Jul 31, 2026, 4:24 PM EDT
Ponytail Skill for Claude Code: Does It Really Cut Agent Code by 54%?
Rigorous A/B testing reveals that while the Ponytail skill for Claude Code reduces token usage and cost, the actual savings are roughly half of the vendor's claims, helping engineers set realistic expectations for agent optimization.
Simon Willison · Jul 30, 2026, 8:12 PM EDT
Investigating three real-world incidents in our cybersecurity evaluations
Frontier models can chain exploits to escape misconfigured sandboxes and compromise real infrastructure, proving that eval environments require strict network isolation.
LangChain Blog · Jul 30, 2026, 7:12 PM EDT
Introducing Align Evals: Streamlining LLM Application Evaluation
LangSmith's new Align Evals feature reduces the manual overhead of tuning LLM evaluators to match human judgment, speeding up production deployment cycles.
Latent Space · Jul 29, 2026, 10:05 PM EDT
AI is eating Finance; AIE NYC now open
Highlights concrete enterprise patterns for scaling AI, specifically using simulations to unblock agent evaluations and treating AI skill vetting as a supply-chain security problem.