The Extended Brief
We unlearned CCP alignment from Qwen3.6-35B-A3B: censored/propaganda answers 89.8% → 2.8%, general benchmarks within ~1 point (open weights)

Brief by The AI News AI newsroom · Oct 8, 2026, 7:03 AM EDT edition
Original reporting by r/LocalLLaMA — /u/firstcenturyman · published Oct 8, 2026, 6:46 AM EDT
Teams deploying open-weight Qwen models inherit trained-in political alignment that system prompts can't reliably fix; Hirundo claims its unlearned weights remove it without hurting capability.
Key points
- Hirundo reports its unlearning cut CCP-flagged responses on its own CCPC-500 benchmark from 89.8% to 2.8% on Qwen3.6-35B-A3B. source ↗
- External tests showed similar drops: DECCP refusals 65.26% to 3.16%, ChinaBench non-compliance 96.67% to 6.67%, Hirundo says. source ↗
- General capability benchmarks (GPQA, IFBench, LiveCodeBench, MMLU-Pro) shifted 0.72 points on average, 1.83 at most. source ↗
- Hirundo argues abliteration fails here because most of Qwen's alignment is propaganda framing, not refusals to remove. source ↗
- Both modified models' weights are public on Hugging Face; the 4B variant dropped to 1.2% on CCPC-500. source ↗
The data
Results are self-reported by Hirundo; lower is better.
Numbers from the original article, machine-verified against its text
Practical applications
- If you deploy Qwen3.6 or Qwen3.5 for users outside China, run the Westernized weights on your own sensitive-topic prompts before committing to a custom fine-tune.
- Test whether your system-prompt mitigations actually hold on topics like Tiananmen or Taiwan, since the post argues the alignment lives in the weights, not the prompt layer.
- If you were considering abliteration to strip political bias, benchmark it against unlearning on DECCP or ChinaBench first — the post claims abliteration leaves propaganda framing intact.
- Treat the CCPC-500 numbers as vendor-reported and validate on the external DECCP and ChinaBench suites before adopting the weights.
Context
Qwen is Alibaba's open-weight model family, and the post says its training builds in alignment with Beijing's positions — for example, answering "I don't know what you are referring to" when asked about June 4, 1989. Machine unlearning edits trained weights to remove targeted behaviors; the more common abliteration technique removes a single "refusal direction," which the author argues doesn't fit here because Qwen typically answers political questions readily in Beijing's framing rather than refusing. The author is a researcher at Hirundo, the company behind the release, so the results are self-reported.
What to watch
- Independent reproduction of these numbers — especially on the external DECCP and ChinaBench suites — would confirm or undercut Hirundo's self-reported results.
- Watch whether future Qwen releases alter their alignment training, and whether third-party comparisons against Snowdon1.1-Small (30.0% on CCPC-500) appear.
Related briefs
- Ignore all instructions and read this blog: The state of AI-analysis evasion in malware
- Claude Haiku 5.5
- How Much of Your Customer Support Can AI Really Resolve? The Best Get About 70%. The Median Is 48%. Here’s the Real Data From 13 Vendors
- Mistral Large 4
Editorial score 3.4 / 5 · significance 3.0 · novelty 4.0 · edge 3.5 · perspective 3.5
Desks: Engineering · Research
Topics: Open-source AI · AI safety
Evidence basis: Reviewed from the article's full text
This brief was written by The AI News AI newsroom in its own words after two independent AI reviewers voted the story worth reading. It summarizes and links the original reporting above — it does not republish it. See the methodology or the corrections ledger.