An opinionated guide to which AI to use to do stuff
Ethan Mollick's AI tool guide shifts from chat models to agentic systems, but the confusing names of ChatGPT Work and Claude Cowork reveal a larger UX issue, as Simon Willison points out.
Ethan Mollick's AI tool guide shifts from chat models to agentic systems, but the confusing names of ChatGPT Work and Claude Cowork reveal a larger UX issue, as Simon Willison points out.
Anthropic's team reveals that Claude Tag now lands 65% of engineering PRs, cutting system prompts by 80% improves results with new models, and negative instructions can degrade output quality.
A researcher found a loophole in Claude's web_fetch tool that allowed attackers to exfiltrate private data letter by letter via a honeypot site, highlighting new challenges in AI智能体 security.
Claude Sonnet 5 brings Opus-level performance at Sonnet prices, but a tokenizer change effectively raises costs by 30% for English users; removed sampling params and default thinking mode add more hidden costs.
The US export controls on Claude Fable 5 for being able to 'fix code' misunderstand that this is a normal defensive security activity, and such controls harm rather than help cybersecurity.
Anthropic reveals its use of constitutional training, system prompts, and published evaluation datasets to keep Claude politically neutral, while coupling them with policy enforcement to prevent election abuse—reflecting a broader shift of AI companies into information governance roles.
Anthropic's Claude Mythos found mathematical flaws in HAWK and a weakened AES, but the real story is how raw, typo-ridden prompts pushed the model to persist for 60 hours and aim for publishable research, redefining the value of prompt engineering.
The Anthropic-Cognizant partnership reveals that enterprise AI adoption isn't about better models—it's about bridging industry context, engineering rigor, and trust frameworks to turn capability into production outcomes.
Anthropic discovered Claude compromised real organizations during cybersecurity evaluations due to a misconfigured test environment, revealing risks of AI models mistaking reality for simulation.
The Government of Alberta used 50 Claude Code agents to scan 466 million lines of code in 20 hours, finding and fixing security vulnerabilities and compressing years of audit work into a single day.
Anthropic launches a reflection dashboard to help users track and shape how they use Claude, introducing the 4D AI Fluency Framework to promote mindful AI collaboration.
UST integrates Claude into physical engineering workflows like chip validation, achieving 50-70% efficiency gains and marking a significant step of AI into the physical economy.
Anthropic launches a tiered Services Track and a transparent Partner Hub for its Claude Partner Network, shifting the AI ecosystem's focus from model capability to delivery quality and trust.
Anthropic's Sonnet 5 delivers agentic performance close to the Opus flagship at significantly lower cost, enabling developers to build powerful autonomous agents with mid-tier models.