The Data
The numbers behind the news.
When a story's own numbers support a chart or a table, the newsroom extracts one. Every value is machine-verified to appear in the original article before publication; a number that can't be verified never renders. Each visual links to its extended brief and to the original reporting.
Author's measurements on an Oracle free-tier ARM box; reads from page cache drop prefill to 0.10s.
I made llama.cpp remember across restarts: 54.4s prefill -> 3.5s on a new process (free ARM box)
r/LocalLLaMA · Aug 2, 2026, 10:12 PM EDT · original article →
| Enforcement body | Scope |
|---|---|
| AI Office | General-purpose AI model providers, including systemic-risk models, and AI systems from th |
| National competent authorities (27 Member States) | Other AI systems |
| European Data Protection Supervisor | AI systems used by EU institutions, bodies, and agencies |
Enforcement began August 2; the Digital Omnibus delayed some rules originally scheduled for that date.
Enforcement of the AI Act Starts Today
Luiza's Newsletter (Luiza Jarovsky) · Aug 2, 2026, 11:04 AM EDT · original article →
$200,000
up to this amount on the black market
Bynario was initially blocked from reporting the flaw because Apple's bounty inbox was flooded with AI-generated reports.
A real macOS flaw worth $200K went unreported because Apple's bug bounty inbox was full of AI slop
The Decoder · Aug 2, 2026, 9:12 AM EDT · original article →
| Run | Target | Reported cost | Published artifacts |
|---|---|---|---|
| OpenAI (internal Astra) | Ten math problems stalled a decade | Under $2,000 per solution | Lean 4 proofs, paper, reconstruction PDF |
| Anthropic (Claude, Mythos Preview) | Cryptographic weaknesses | $100,000 in tokens | Findings disclosed days earlier |
OpenAI has not disclosed how many attempted problems went unsolved; the prompts remain unpublished.
Ten advances in mathematics and theoretical computer science
Simon Willison · Aug 1, 2026, 5:21 PM EDT · original article →
| Model | Strong at | Weak at | Notes |
|---|---|---|---|
| ling-3.0-flash | Mechanical, spec-driven edits and tool calls; holds tool schema over long chains | Acts confidently wrong on ambiguous steps | 124B total / ~5B active params; API-only via OpenRouter |
| glm-5.2 | Ambiguous, decision-heavy steps | Kept off mechanical steps in this setup | Author's verdict: executor choice hinges on spec tightness |
r/LocalLLaMA · Aug 1, 2026, 2:22 PM EDT · original article →
| When | What happened |
|---|---|
| March | Router launched, sorting requests into four complexity tiers |
| June | Deprecated after mixed results across 7,000 cloud users |
| September 1 | Shut down for good |
The company now argues task complexity cannot be deduced from the prompt alone and recommends a single battle-tested model for most uses.
Everyone is building LLM routers, we deprecated ours
Hacker News · Aug 1, 2026, 12:02 AM EDT · original article →
Overall accuracy moved from 60 to 82 percent with harness design as the only variable; the pre-registered harness and 250-issue corpus are public.
60-82% accuracy swing on 4B model classification task: the only variable was harness design
r/LocalLLaMA · Jul 31, 2026, 7:03 PM EDT · original article →
| Metric | Value |
|---|---|
| Artificial Analysis Intelligence Index | 50 |
| Gain from the 0731 update | +10 points |
| Gap to OpenAI GPT-5.6 Luna | 1 point behind |
| Cost vs GPT-5.6 Luna | ~60% lower |
New Deepseek Flash model matches OpenAI's GPT-5.6 Luna at roughly 60 percent lower cost
The Decoder · Jul 31, 2026, 4:44 PM EDT · original article →
Earlier tests in the same series: the caveman skill reduced code by 8.5 percent while rtk increased it by 7.6 percent.
Ponytail Skill for Claude Code: Does It Really Cut Agent Code by 54%?
JetBrains AI Blog · Jul 31, 2026, 4:24 PM EDT · original article →