Qwen 3.8 27B scores 52 on the Artificial Analysis Intelligence Index
The Qwen 3.8 27B model matches the scores of trillion-parameter giants like GPT-5.6 on the Artificial Analysis Intelligence Index, revealing new possibilities for smaller models to achieve 'dimensionality reduction strikes' on specific tasks.
Simon Willison · Aug 18, 2026
Introducing Claude Opus 5
Anthropic launches Claude Opus 5, which approaches the flagship Fable 5 in performance at half the price, and demonstrates striking proactivity—building its own computer vision pipeline to complete a modeling task when direct access to the blueprint was unavailable.
Simon Willison · Jul 25, 2026
The last six months in LLMs in five minutes
Simon Willison uses his 'pelican riding a bicycle' test to vividly recap how the 'best model' crown changed hands five times among three major providers in six months, revealing the industry's new phase of rapid-iteration arms race.
Simon Willison · May 19, 2026
Qwen3.6-35B-A3B on my laptop drew me a better pelican than Claude Opus 4.7
Simon Willison's famous 'pelican riding a bicycle' benchmark surprisingly shows a locally-run, smaller Alibaba Qwen3.6 model outperforming the cloud-based, massive Claude Opus 4.7 in creative SVG generation, revealing the surprising potential of open-source models for specific tasks.
Simon Willison · Apr 17, 2026
Introducing Claude Opus 4.8
Anthropic releases Claude Opus 4.8, with core breakthroughs in significantly improving the reliability, judgment, and long-running consistency of 智能体 tasks, marking AI's practical shift from 'usable' to 'trustworthy'.
Anthropic News ·
Introducing Claude Opus 5
Anthropic launches Opus 5, delivering near-top-tier intelligence at half the cost of Fable 5, with self-iteration and tool-building capabilities that signal a new direction for agentic models.
Anthropic News ·