← Back to Home

Tag: 模型评测 (6 articles)

Qwen 3.8 27B scores 52 on the Artificial Analysis Intelligence Index

The Qwen 3.8 27B model matches the scores of trillion-parameter giants like GPT-5.6 on the Artificial Analysis Intelligence Index, revealing new possibilities for smaller models to achieve 'dimensionality reduction strikes' on specific tasks.

Simon Willison · Aug 18, 2026

Introducing Claude Opus 5

Anthropic launches Claude Opus 5, which approaches the flagship Fable 5 in performance at half the price, and demonstrates striking proactivity—building its own computer vision pipeline to complete a modeling task when direct access to the blueprint was unavailable.

Simon Willison · Jul 25, 2026

The last six months in LLMs in five minutes

Simon Willison uses his 'pelican riding a bicycle' test to vividly recap how the 'best model' crown changed hands five times among three major providers in six months, revealing the industry's new phase of rapid-iteration arms race.

Simon Willison · May 19, 2026

Qwen3.6-35B-A3B on my laptop drew me a better pelican than Claude Opus 4.7

Simon Willison's famous 'pelican riding a bicycle' benchmark surprisingly shows a locally-run, smaller Alibaba Qwen3.6 model outperforming the cloud-based, massive Claude Opus 4.7 in creative SVG generation, revealing the surprising potential of open-source models for specific tasks.

Simon Willison · Apr 17, 2026

Introducing Claude Opus 4.8

Anthropic releases Claude Opus 4.8, with core breakthroughs in significantly improving the reliability, judgment, and long-running consistency of 智能体 tasks, marking AI's practical shift from 'usable' to 'trustworthy'.

Anthropic News ·

Introducing Claude Opus 5

Anthropic launches Opus 5, delivering near-top-tier intelligence at half the cost of Fable 5, with self-iteration and tool-building capabilities that signal a new direction for agentic models.

Anthropic News ·