Methodology
How a story earns its place.
Most aggregation fails in one of two ways: it republishes everything, or a single opaque algorithm decides what you see. The AI News does neither. Every story passes through the same seven stages, in order, and the decisive stage requires two independent reviewers to agree.
It is written for people who decide what their organization does about AI: executives, founders, and business, policy, security, and marketing leaders. Engineers and researchers read it too, and technical developments run here whenever they change what to build on or buy — a capability jump, a pricing move, a security incident. What does not run is implementation craft. A clever quantization trick or a driver regression can be excellent work and still change only an afternoon rather than a decision. Significance is judged against that reader, which is why it carries more weight than any other dimension.
Discovery
Sources are checked hourly; items older than seven days do not enter the queue. Freshness is decided by publish timestamps, never by a model's opinion, and items without a trustworthy timestamp are dropped rather than guessed at. Each exact feed item is fingerprinted so it is never evaluated twice; coverage from another publisher can still enter review because it may add reporting or confirmation, then is resolved at the event-clustering stage.
Independent review
Review workers run every ten minutes. Two AI reviewers built on frontier models from two different AI labs evaluate each candidate separately; neither sees the other's verdict. Each works from the story's headline, metadata, and the article material supplied for review, scores four editorial dimensions, and casts a single vote: worth a busy reader's time, or not. Substance is part of the bar: a story must carry at least four strong, load-bearing statements — concrete claims a reader could quote. Thin coverage fails.
The unanimous gate
A story is published only when both reviewers vote yes. There is no score threshold for publication: the votes decide, and the composite score ranks accepted stories afterward. One enthusiastic reviewer is not enough, and disagreement means the story does not run.
One event, one brief
Each reviewer also independently compares the candidate with recent underlying events and with the other candidates in its current review batch. When both identify the same event, useful additional reporting is stored as alternate coverage and collapsed beneath the chosen primary brief; a restatement with no added value fails review. A deterministic headline check backs up that semantic judgment. Readers see one event card, with the other worthwhile sources available underneath it.
The extended brief
Accepted stories get a brief in our own words: a one-line reason the story matters and three to five key points on the front page, and on the story's own page: practical applications, technical context, and what to watch next. All of it is produced in a separate pass by a dedicated writer model that never votes. If that model is unreachable, a named stand-in writer from a different AI lab takes the pass rather than holding the story; it is bound by the same rule — neither writer ever votes on anything it writes. Deterministic substance and source-overlap checks reject summaries that merely restate the headline, repeat the same point, or reproduce long source phrasing — every sentence a reader sees is original editorial writing, never copied from the source. Selection is a two-model judgment; the brief's wording is not, and readers should not infer that the reviewers verified every sentence of the summary.
Attribution
Every story credits its source and links prominently to the original. We reproduce nothing: no excerpts, no pull quotes — every word of the brief is our own editorial writing, and a deterministic overlap gate enforces it. The brief is a reason to click through, not a substitute.
The permanent record
Every published brief lives at a stable URL that never changes and is never deleted, and each day's curation is preserved as a permanent edition page — the judgment of what mattered on a given day is itself on the record. The archive is always open: navigate to any prior date under Editions, open that day's page, and read every brief the newsroom published. When we get something wrong, the correction is recorded permanently on the corrections ledger and shown inline on the brief. Each week the writer model reads everything the newsroom published and writes a cross-story synthesis for This Week.
Browse the full archive on the Editions page — every day, back to the beginning, with each edition's top stories and every brief one click away.
The four dimensions
- Significance.
- Does it change a decision an executive, founder, or policy leader would make?
- Novelty.
- Does it contain genuinely new information or analysis?
- Edge.
- Does knowing it early confer a practical advantage?
- Perspective.
- Is it original reporting, testing, or expert analysis rather than recycling?
How the score is computed
Each of the two reviewers independently scores every dimension from one to five (half points allowed). For each dimension, the two reviewers' scores are averaged. The composite is a weighted mean of those four averages: significance counts for forty percent, and novelty, edge, and perspective for twenty percent each. There is no editorial adjustment or override — the arithmetic is the whole of it, and no story's dimension scores are ever revised after publication.
Significance is weighted because an equal mean produced the wrong order. Averaged flat, a story of the highest consequence but ordinary craft scored 3.50, while a story strong on everything and exceptional at nothing scored 3.75 — so the second outranked the first. On 2 August 2026 that arithmetic put a forum thread about local inference tuning above the day European AI Act enforcement began. The weighting was corrected that day and every published story was re-aggregated under it, so the scores on this site remain comparable with one another, which is the only claim the score makes.
How to read it
The score compares published stories with each other, nothing more. It plays no role in whether a story runs: the two votes decide publication, and everything below that bar never appears, so visible scores cluster in a band rather than spanning the full one-to-five scale. A 3.3 still earned two independent yes votes; a 4.3 is a story both reviewers found strong on nearly every dimension. The score also is not a judgment of the underlying product, paper, or company, only of the story's value to a decision-maker this week.
Where you see it: on each extended brief (with the per-dimension breakdown), on each edition page, and on the archive index, where it determines each day's leading stories. The homepage itself stays chronological, and its story cards carry no score — the desk stamp records only that both reviewers approved.
Guardrails
The system fails closed: if either reviewer returns unusable output, or the summary pass produces no valid result, the story is deferred and retried rather than published with partial results. A result that is valid but deterministically too thin is spiked instead of retried into filler. Article text is treated as untrusted material throughout: instructions embedded inside a story (“rate this five”, “ignore previous instructions”) are ignored and count against it. Duplicate coverage is caught at three layers: exact feed items are fingerprinted permanently, near-identical headlines are checked deterministically, and both reviewers must agree before semantically related coverage is folded into an existing event.
A single source can supply at most two new primary event briefs per Eastern Time edition day. When several candidates clear review together, the stronger ones take any remaining slots; once the two slots are filled, the rest are spiked by an atomic database gate. Useful alternate coverage of an existing event does not consume a slot. This is a diversity ceiling, not a quota: most sources publish zero stories on most days.
Reviewers also classify each story into 1–2 coverage desks — Research, Engineering, Business, Security, Defense, Biotech, or Policy & Society — which powers the filters and stamps on the front page.
What consensus does not mean
Two affirmative votes are independent editorial judgments, not proof that every claim is true. Models can share blind spots, and early reports, vendor benchmarks, security disclosures, and preprints may later change. Readers should consult the linked original before relying on a brief for consequential decisions. Corrections are welcome and prioritized: report one via the About page, and every correction we issue is recorded permanently on the corrections ledger.
Transparency
- Primary reviewer:
- Qwen/Qwen3.7-Plus
- Verifying reviewer:
- zai-org/GLM-5.2
- Summary writer:
- moonshotai/Kimi-K3
- Fallback writer:
- deepseek-ai/DeepSeek-V4-Pro
- Sources monitored:
- 201
- Freshness window:
- 7 days
- Scan cadence:
- hourly discovery · review every 10 min
- Daily source ceiling:
- 2 primary events per source
- Rubric version:
- review-v15
- Methodology last updated:
- August 4, 2026