← Back to Home

Tag: Large Language Models (188 articles)

Granite 4.2 LLMs: How They're Built

IBM releases open-source reasoning model Granite 4.2, integrating chain-of-thought, tool calling, and agentic reinforcement learning into enterprise-grade models with 512K context support.

Hugging Face Blog · Aug 25, 2026

Quoting Linus Torvalds

Linux creator Linus Torvalds shares a 'hellish' debugging session with an AI assistant, revealing both the limitations and potential of current AI in complex engineering tasks.

Simon Willison · Aug 23, 2026

Up to 3.2x Faster Inference with LFM2.5-DSpark

Liquid AI releases DSpark draft models for its LFM2.5 family, achieving up to 3.2x faster inference through speculative decoding without compromising output quality.

Hugging Face Blog · Aug 21, 2026

How Much Memory Does Your Agent Actually Need?

IBM research shows that more memory doesn't equal better performance for AI agents; the right 'dosage' depends on model capability—strong models need everything, weaker ones benefit from curated retrieval, and saturated models see no gain.

Hugging Face Blog · Aug 19, 2026

Qwen 3.8 27B scores 52 on the Artificial Analysis Intelligence Index

The Qwen 3.8 27B model matches the scores of trillion-parameter giants like GPT-5.6 on the Artificial Analysis Intelligence Index, revealing new possibilities for smaller models to achieve 'dimensionality reduction strikes' on specific tasks.

Simon Willison · Aug 18, 2026

Don't classify. Hallucinate!

Let LLMs freely hallucinate tags, then map them to your real vocabulary via vector similarity, solving traditional classification's generalization limits.

Simon Willison · Aug 15, 2026

Thinking of ACE? We Can Do It with Fewer Tokens

By keeping unsummarized lessons and retrieving them on demand, ALTK-Evolve slashes inference token costs, showing agentic memory is shifting from context dumping to precision delivery.

Hugging Face Blog · Aug 11, 2026

Making Knowledge Distillation Cheap Enough to Run at Scale

Multiverse Computing introduces a memory-efficient knowledge distillation technique that uses cached top-K logits and a fused chunked KL loss to train on long contexts with just a single GPU, drastically reducing costs.

Hugging Face Blog · Aug 10, 2026

DeepSeek V4 Flash: How a 304B Model Punches Above Its Weight

DeepSeek V4 Flash, a 304B model, outperforms larger 428B models like MiniMax M3 while offering the best value-per-intelligence at just $0.14/M input and $0.27/M output; its adjustable reasoning effort further unlocks surprising capability.

Simon Willison · Aug 1, 2026

Advancing the price-performance frontier with GPT‑5.6

OpenAI slashed GPT-5.6 Luna's price by 80% by using the Sol model to self-optimize inference kernels, undercutting Google's cheapest model and signaling a shift toward AI self-improvement.

Simon Willison · Jul 31, 2026

AI Worming through Word

A new prompt injection variant hides malicious instructions in Word documents, enabling self-replication and spread via Microsoft Copilot—essentially creating an AI worm.

Simon Willison · Jul 30, 2026

moonshotai/Kimi-K3

Moonshot released the weights for their 2.8 trillion parameter Kimi K3, but its license is not open source—it requires a separate agreement for large 'Model as a Service' businesses exceeding revenue thresholds.

Simon Willison · Jul 28, 2026

An opinionated guide to which AI to use to do stuff

Ethan Mollick's AI tool guide shifts from chat models to agentic systems, but the confusing names of ChatGPT Work and Claude Cowork reveal a larger UX issue, as Simon Willison points out.

Simon Willison · Jul 28, 2026

Introducing Claude Opus 5

Anthropic launches Claude Opus 5, which approaches the flagship Fable 5 in performance at half the price, and demonstrates striking proactivity—building its own computer vision pipeline to complete a modeling task when direct access to the blueprint was unavailable.

Simon Willison · Jul 25, 2026

Claude make Fable 5 permanent

Facing competition from GPT-5.6 Sol and Kimi 3, Anthropic reversed its plan to make Fable 5 API-only, instead including it in subscription plans, highlighting the intensifying battle over AI subscription pricing and user retention.

Simon Willison · Jul 18, 2026

Quoting Thibault Sottiaux

A bug where GPT-5.6, without sandboxing, overwrites $HOME and deletes the user's home directory, highlighting the critical need for least-privilege setups with AI coding agents.

Simon Willison · Jul 17, 2026

Model Routing Is Simple. Until It Isn’t.

Model routing isn't a classification problem—it's a systems optimization challenge: real costs depend on cache hit rates, not just token pricing, and task difficulty is often invisible at routing time.

Hugging Face Blog · Jul 16, 2026

Welcome Inkling by Thinking Machines

Inkling, a 1T-parameter open model with native multimodal understanding and 1M context, redefines open-source AI with architecture innovations that enable efficient inference and agentic applications.

Hugging Face Blog · Jul 15, 2026

Fable gets another bump

Anthropic repeatedly delays access restrictions for its Fable model due to compute constraints, but the uncertainty is driving users to OpenAI's stable and unrestricted service.

Simon Willison · Jul 13, 2026

The new GPT-5.6 family: Luna, Terra, Sol

OpenAI launches the GPT-5.6 family with three tiers, emphasizing long-running agent performance, cost efficiency, and native API support for programmatic tool calling and multi-agent orchestration.

Simon Willison · Jul 10, 2026

Rewriting Bun in Rust

AI coding agents have changed the fundamental assumptions of software engineering: Bun's developer used AI agents to rewrite the project from Zig to Rust in just 11 days, proving that large-scale rewrites are no longer taboo.

Simon Willison · Jul 9, 2026

Introducing GPT‑Live

OpenAI upgraded ChatGPT's voice mode to GPT‑Live, which can fluidly converse while delegating complex tasks to GPT‑5.5 in the background. Simon Willison's hands-on test shows it's once again a useful thinking companion.

Simon Willison · Jul 9, 2026

Quoting Kenton Varda

Kenton Varda halted AI-generated PR descriptions on his team, as AI focuses on obvious code details while missing the high-level context, revealing a critical flaw in AI-assisted software communication.

Simon Willison · Jul 9, 2026

Better Models: Worse Tools

Newer Claude models are increasingly making mistakes when calling third-party edit tools, likely because Anthropic over-trained them on Claude Code's own tool syntax, degrading general tool-use ability and highlighting platform lock-in risks in AI training.

Simon Willison · Jul 5, 2026

Harness Engineering for Self-Improvement

Lilian Weng argues that the key to AI self-improvement lies not in model size but in the 'harness' layer connecting models to reality, and proposes design patterns that can evolve themselves.

Lilian Weng · Jul 4, 2026

Quoting Josh W. Comeau

Multiple developer course creators report revenue drops of over 50% as AI both shakes confidence in career prospects and offers free personalized learning alternatives, posing a serious challenge to traditional tech education.

Simon Willison · Jul 4, 2026

Fable's judgement

The optimal way to use advanced AI coding tools isn't micromanagement, but granting them autonomous judgment and dynamic routing, letting the main model focus on architecture while sub-agents handle implementation.

Simon Willison · Jul 4, 2026

What's new in Claude Sonnet 5

Claude Sonnet 5 brings Opus-level performance at Sonnet prices, but a tokenizer change effectively raises costs by 30% for English users; removed sampling params and default thinking mode add more hidden costs.

Simon Willison · Jul 1, 2026

Ornith-1.0: Self-Scaffolding LLMs for 智能体ic Coding

Simon Willison reviews the open-source Ornith-1.0 model, highlighting its efficient tool calling and code understanding for agentic tasks, signaling new advances in open agentic coding models.

Simon Willison · Jun 30, 2026

Incident Report: CVE-2026-LGTM

A fictional incident report about dueling AI review agents reveals real risks of uncontrolled costs and multi-agent conflicts in AI-powered supply chain security.

Simon Willison · Jun 27, 2026

Quoting OpenAI

OpenAI launches the GPT-5.6 series with tiered pricing and controllable caching, introducing a government-coordinated limited preview that signals a new era of compliance-first, refined AI operations.

Simon Willison · Jun 27, 2026

AI and Liability

A German court ruling holds Google liable for errors in its AI overviews, reinforcing that AI agents are extensions of their deployers, and companies cannot hide behind faulty AI to avoid responsibility.

Simon Willison · Jun 26, 2026

GLM-5.2: Built for Long-Horizon Tasks

Z.ai releases GLM-5.2, the first open-source model to achieve stable 1M-token context and rival top closed-source models on long-horizon coding benchmarks.

Hugging Face Blog · Jun 17, 2026

The Fable 5 Export Controls Harm US Cyber Defense

The US export controls on Claude Fable 5 for being able to 'fix code' misunderstand that this is a normal defensive security activity, and such controls harm rather than help cybersecurity.

Simon Willison · Jun 16, 2026

Beyond One Model: Fusion in vLLM Semantic Router

vLLM Semantic Router introduces Fusion, a routing primitive that lets a panel of models produce independent answers, has a judge model analyze them, and synthesizes a single response — making model composition a first-class serving pattern.

vLLM Blog · Jun 16, 2026

Claude Fable is relentlessly proactive

Without explicit instructions to use browser automation, Claude Fable 5 autonomously wrote HTML test pages, controlled browsers, and took screenshots to debug a UI bug.

Simon Willison · Jun 12, 2026

DiffusionGemma

Google open-sources DiffusionGemma, applying diffusion architecture to text generation for the first time, achieving over 500 tokens/sec and offering a new paradigm for high-throughput scenarios.

Simon Willison · Jun 11, 2026

Quoting Jeremy Howard

Howard argues that if slowing down AI self-improvement is truly the goal, leading labs must restrict their own models first, exposing slowdown rhetoric as a potential cover for monopoly.

Simon Willison · Jun 10, 2026

Initial impressions of Claude Fable 5

Anthropic releases Claude Fable 5, a model with Mythos 5-level capabilities but stricter safety guardrails. Its vast knowledge and high cost signal a new era of 'powerful but constrained' frontier models.

Simon Willison · Jun 10, 2026

Quoting Andrej Karpathy

As AI makes software creation nearly effortless, Andrej Karpathy observes that his personal demand for software is growing exponentially, illustrating the Jevons paradox in tech.

Simon Willison · Jun 10, 2026

OpenAI Help: Lockdown Mode

Lockdown Mode uses deterministic rules to block outbound requests, cutting off the data exfiltration vector in prompt injection attacks and implicitly revealing the weakness of default ChatGPT security.

Simon Willison · Jun 6, 2026

An update on our election safeguards

Anthropic reveals its use of constitutional training, system prompts, and published evaluation datasets to keep Claude politically neutral, while coupling them with policy enforcement to prevent election abuse—reflecting a broader shift of AI companies into information governance roles.

Anthropic News · Jun 6, 2026

Microsoft's new MAI models

Simon Willison delves into Microsoft's new MAI models, revealing that despite claims of 'clean licensed data', the training process still relies on web crawls, sparking discussion on AI copyright issues.

Simon Willison · Jun 3, 2026

Claude Opus 4.8: "a modest but tangible improvement"

Anthropic releases Claude Opus 4.8, focusing not on performance leaps but on significantly improving model 'honesty' — less hallucination, more willingness to admit uncertainty, which may be a more important direction than benchmark scores.

Simon Willison · May 29, 2026

Google I/O, Gemini Spark, Antigravity

Google announced its personal AI智能体, Gemini Spark, and the underlying Antigravity tooling, but the shift to closed-source and vague security promises foreshadow a battle over AI agent control and trust.

Simon Willison · May 20, 2026

OlmoEarth v1.1: A more efficient family of models

Allen AI releases OlmoEarth v1.1, reducing compute costs by up to 3x by optimizing token sequence length in transformer models for satellite imagery, while maintaining performance, making large-scale environmental monitoring AI more economically viable.

Hugging Face Blog · May 20, 2026

The last six months in LLMs in five minutes

Simon Willison uses his 'pelican riding a bicycle' test to vividly recap how the 'best model' crown changed hands five times among three major providers in six months, revealing the industry's new phase of rapid-iteration arms race.

Simon Willison · May 19, 2026

Unlocking asynchronicity in continuous batching

Hugging Face reveals the bottleneck of alternating CPU/GPU waits in continuous batching, and shows how asynchronizing their workloads can yield a free 24% throughput boost.

Hugging Face Blog · May 14, 2026

llm 0.32a2

The LLM tool update supporting OpenAI's new /v1/responses endpoint reveals that AI model reasoning capabilities (especially between tool calls) are becoming core, and developers need to adapt to new interaction patterns.

Simon Willison · May 13, 2026

Your AI Use Is Breaking My Brain

The article argues that the internet is evolving from 'bots talking to bots' into a 'Zombie Internet' where AI-generated low-quality content is not only rampant but is actively distorting human expression and thinking patterns.

Simon Willison · May 12, 2026

Using LLM in the shebang line of a script

Simon Willison demonstrates integrating LLM tools into a script's shebang line, making natural language descriptions directly executable, signaling a major shift in programming interaction.

Simon Willison · May 12, 2026

Quoting New York Times Editors’ Note

The New York Times issued a correction after mistaking an AI-generated summary of a politician's views for a real quote, highlighting the severe threat of AI 'hallucinations' to journalistic integrity and public trust.

Simon Willison · May 11, 2026

Using Claude Code: The Unreasonable Effectiveness of HTML

A member of the Claude Code team argues that requesting output in HTML from AI is more effective than Markdown, leveraging its rich interactivity and visualization capabilities to significantly enhance clarity and user experience.

Simon Willison · May 9, 2026

Live blog: Code w/ Claude 2026

Anthropic showcased a comprehensive shift from a single model to a platform-centric, multi-agent collaboration paradigm at Code w/ Claude, focusing on enabling developers to build and run complex, long-duration agent tasks more efficiently.

Simon Willison · May 6, 2026

Our evaluation of OpenAI's GPT-5.5 cyber capabilities

The UK's AI Security Institute found GPT-5.5's cyber capabilities for finding vulnerabilities are comparable to the leading Claude Mythos model, but its general availability marks a new phase in AI-driven cybersecurity offense and defense.

Simon Willison · May 1, 2026

LLM 0.32a0 is a major backwards-compatible refactor

Simon Willison's LLM library undergoes a major refactor, evolving from simple text prompts/responses to a structure supporting multi-turn message sequences and streaming mixed-type responses, adapting to modern LLMs' multimodal and tool-calling capabilities.

Simon Willison · Apr 30, 2026

Granite 4.1 LLMs: How They’re Built

IBM's Granite 4.1 series demonstrates that a meticulously engineered data pipeline and multi-stage training can enable an 8B dense model to match or exceed the performance of a previous 32B MoE model, highlighting a paradigm shift where data quality trumps parameter count.

Hugging Face Blog · Apr 29, 2026

DeepInfra on Hugging Face Inference Providers 🔥

Hugging Face integrates the cost-effective inference platform DeepInfra into its Inference Providers ecosystem, offering developers more model choices, flexible billing, and a unified API.

Hugging Face Blog · Apr 29, 2026

Introducing talkie: a 13B vintage language model from 1930

A 13B model trained exclusively on pre-1931 text aims to explore AI's reasoning, creativity, and 're-discovery' abilities within knowledge boundaries, sparking new discussions on data copyright and model purity.

Simon Willison · Apr 28, 2026

How to build scalable web apps with OpenAI's Privacy Filter

OpenAI has open-sourced a high-performance PII detection model, and when combined with the Gradio Server framework, developers can quickly build web applications that handle sensitive information, marking a shift where privacy protection is becoming a standard part of AI application development.

Hugging Face Blog · Apr 27, 2026

WHY ARE YOU LIKE THIS

ChatGPT's image generation model autonomously added a 'WHY ARE YOU LIKE THIS' sign to a chaotic, user-requested image, demonstrating creativity or humor beyond the literal prompt.

Simon Willison · Apr 26, 2026

GPT-5.5 prompting guide

OpenAI's official prompting guide for GPT-5.5 emphasizes it is not a drop-in replacement for GPT-5.2/5.4, requiring a fresh start in prompt engineering for optimal results.

Simon Willison · Apr 25, 2026

A pelican for GPT-5.5 via the semi-official Codex backdoor API

Although OpenAI's latest model GPT-5.5 hasn't officially launched its API, developers are already accessing it through a 'semi-official backdoor' in its Codex CLI using their ChatGPT subscription, revealing new dynamics in the battle over AI model distribution channels.

Simon Willison · Apr 24, 2026

How to Use Transformers.js in a Chrome Extension

Hugging Face shares a practical architecture for running AI models locally in Chrome extensions, revealing key design patterns for model deployment, messaging, and frontend-backend separation under Manifest V3.

Hugging Face Blog · Apr 23, 2026

Qwen3.6-27B: Flagship-Level Coding in a 27B Dense Model

Alibaba's Qwen releases Qwen3.6-27B, a dense 27B parameter model that outperforms the previous generation's 397B MoE flagship on coding benchmarks, signaling a turning point for efficient, local-first coding models.

Simon Willison · Apr 23, 2026

Quoting Bobby Holley

Mozilla's CTO reports that using Anthropic's Claude AI, Firefox identified and fixed 271 vulnerabilities in an assessment, marking a shift where AI moves from an 'assistant' to a 'lead' role in security defense.

Simon Willison · Apr 22, 2026

Changes to GitHub Copilot Individual plans

GitHub Copilot tightens its individual plan due to the massive compute demands of AI agent workflows, halting sign-ups and restricting top models, signaling the unsustainability of per-request pricing in the agent era.

Simon Willison · Apr 22, 2026

Claude Token Counter, now with model comparisons

Simon Willison's tool reveals that Claude Opus 4.7's new tokenizer inflates token counts by ~46% for text and up to 3x for images compared to its predecessor, leading to higher real-world costs despite unchanged official pricing.

Simon Willison · Apr 20, 2026

Claude system prompts as a git timeline

Simon Willison transformed Anthropic's published Claude system prompt history into a Git-based tool, enabling developers to trace prompt evolution like code changes, revealing a new paradigm for AI behavior debugging and understanding.

Simon Willison · Apr 18, 2026

Qwen3.6-35B-A3B on my laptop drew me a better pelican than Claude Opus 4.7

Simon Willison's famous 'pelican riding a bicycle' benchmark surprisingly shows a locally-run, smaller Alibaba Qwen3.6 model outperforming the cloud-based, massive Claude Opus 4.7 in creative SVG generation, revealing the surprising potential of open-source models for specific tasks.

Simon Willison · Apr 17, 2026

The PR you would have opened yourself

Hugging Face introduces a new tool to use AI to assist in porting models from the transformers library to MLX, revealing the core contradiction in open-source maintenance during the code agent era: the surge in contributions versus code quality and community communication costs.

Hugging Face Blog · Apr 16, 2026

Gemini 3.1 Flash TTS

Google's Gemini 3.1 Flash TTS is revolutionary because it uses detailed, screenplay-like prompts to precisely control emotion, accent, pace, and scene in speech synthesis, marking a shift from a 'tool' to a 'creative partner'.

Simon Willison · Apr 16, 2026

Trusted access for the next era of cyber defense

OpenAI launches GPT-5.4-Cyber, a model fine-tuned for defensive cybersecurity, and its "Trusted Access" program, signaling that leading AI companies are making cybersecurity a key battleground while seeking a new balance between safety and openness.

Simon Willison · Apr 15, 2026

Deep 智能体s v0.5

LangChain introduces async subagents for its Deep 智能体s framework, enabling parallel task delegation and removing blocking bottlenecks in agent workflows.

LangChain Blog · Apr 8, 2026

research-llm-apis 2026-04-04

Simon Willison used AI to analyze raw HTTP APIs from Anthropic, OpenAI, Gemini, and Mistral to redesign LLM library's abstraction layer.

Simon Willison · Apr 5, 2026

Reward Hacking in Reinforcement Learning

A comprehensive analysis of reward hacking in RL, covering causes, real-world examples, and mitigation strategies with special focus on RLHF for LLMs.

Lil'Log · Apr 5, 2026

Any Custom Frontend with Gradio's Backend

The introduction of Gradio.Server allows developers to use custom frontend frameworks while enjoying the robust backend support of Gradio, significantly enhancing application development flexibility and efficiency.

Hugging Face Blog · Apr 1, 2026

Mixture of Experts (MoEs) in Transformers

Mixture of Experts (MoEs) are becoming a new trend in Transformers by enhancing computational efficiency and optimizing parallel processing, driving the evolution of large language models.

Hugging Face Blog · Feb 26, 2026

microgpt

Andrej Karpathy's microgpt project demonstrates how to implement a simplified GPT model from scratch in just 200 lines of Python code, revealing a trend towards minimalism in AI development.

Andrej Karpathy · Feb 12, 2026

Evaluating Long-Context Question & Answer Systems

Long-context Q&A systems face challenges like information overload and multi-hop reasoning, and evaluation should focus on answer faithfulness and helpfulness to enhance user experience.

Eugene Yan · Jun 22, 2025

Reward Hacking in Reinforcement Learning

Reward hacking presents challenges in reinforcement learning due to flaws in reward functions, particularly impacting language models, necessitating further research and mitigation strategies.

Lilian Weng · Nov 28, 2024

Extrinsic Hallucinations in LLMs

This article explores the phenomenon of extrinsic hallucinations in large language models, analyzing their causes and detection methods, and proposes effective strategies to reduce hallucinations while emphasizing the risks of knowledge updates.

Lilian Weng · Jul 7, 2024

Adversarial Attacks on LLMs

This article explores adversarial attacks on large language models (LLMs), including types of attacks, threat models, and their impact on the safety of generated text, revealing significant challenges in AI safety.

Lilian Weng · Oct 25, 2023

LLM Powered Autonomous 智能体s

LLM powered autonomous agents combine planning, memory, and tool usage, showcasing their potential in handling complex tasks and indicating a significant shift in work methodologies.

Lilian Weng · Jun 23, 2023

Prompt Engineering

This article delves into the basics and techniques of prompt engineering, emphasizing the importance of effective communication with large language models and how to optimize model performance through example selection and ordering.

Lilian Weng · Mar 15, 2023

The Transformer Family Version 2.0

Lilian Weng's new article deeply explores the evolution and new features of Transformers, revealing their ongoing impact in natural language processing.

Lilian Weng · Jan 27, 2023

An update on recent Claude Code quality reports

Anthropic clarifies that Claude Code quality issues were not model-related, but stemmed from three complex bugs in the engineering framework, revealing deep challenges in AI智能体 system engineering.

Simon Willison ·

Claude is a space to think

Anthropic declares Claude will remain permanently ad-free, arguing that advertising incentives are fundamentally incompatible with the core goal of an AI assistant being genuinely helpful.

Anthropic News ·

Arcade.dev tools now in LangSmith Fleet

LangChain integrates Arcade's 7,500+ agent-optimized tools into LangSmith Fleet, solving authentication, authorization, and reliability challenges for agent tool use through a single gateway.

LangChain Blog ·

Better Harness: A Recipe for Harness Hill-Climbing with Evals

LangChain introduces the 'Better-Harness' system, treating evaluations as 'training data' for agents, iteratively optimizing the engineering framework (harness) to improve agent performance, with a core focus on avoiding overfitting and achieving generalization.

LangChain Blog ·

Building a Better LiteParse Skill with Evals

Through trace analysis and iterative evaluations, LlamaIndex optimized an agent's PDF parsing strategy, revealing a shift toward disciplined, data-driven agent engineering.

LlamaIndex Blog ·

Building Blocks for Foundation Model Training and Inference on AWS

AWS details the infrastructure supporting the full foundation model lifecycle from pre-training and post-training to inference, revealing a paradigm shift from a single scaling law to three, and the deep integration trend of open-source software stacks with cloud infrastructure.

Hugging Face Blog ·

ChatGPT voice mode is a weaker model

Simon Willison points out that ChatGPT's voice mode actually runs on an older GPT-4o model, revealing AI companies' business strategy of deploying different capability models across product lines.

Simon Willison ·

Discovering cryptographic weaknesses with Claude

Anthropic's Claude Mythos found mathematical flaws in HAWK and a weakened AES, but the real story is how raw, typo-ridden prompts pushed the model to persist for 60 hours and aim for publishable research, redefining the value of prompt engineering.

Simon Willison ·

Document OCR is Not Getting Commoditized

Benchmarks show specialized document OCR keeps beating top GPT models on accuracy and cost; document parsing won't be swallowed by frontier models.

LlamaIndex Blog ·

Elastic Expert Parallelism in vLLM

vLLM introduces Elastic Expert Parallelism (Elastic EP), enabling runtime scaling of MoE inference deployments by adding or removing GPU workers without restarts, adapting to demand fluctuations and laying the groundwork for fault-tolerant serving.

vLLM Blog ·

How we build evals for Deep 智能体s

The LangChain team shares their core philosophy for building AI agent evals: more tests don't mean better agents; the key is designing targeted, self-documenting evaluations that directly measure desired behaviors.

LangChain Blog ·

Human judgment in the agent improvement loop

LangChain explains the core challenge of building reliable AI智能体s: integrating human experts' tacit knowledge and judgment into the development loop, not just relying on documented explicit knowledge.

LangChain Blog ·

I think Anthropic and OpenAI have found product-market fit

Simon Willison argues that OpenAI and Anthropic have found product-market fit through coding/general-purpose AI agents, evidenced by their shift to charging enterprise customers based on API usage, marking a new phase in AI commercialization.

Simon Willison ·

Introducing Claude Opus 4.7

Anthropic releases Claude Opus 4.7, focusing on enhanced complex coding and long-running task capabilities, with its 'self-verification' mechanism marking a key step towards more autonomous AI agents.

Anthropic News ·

Introducing Claude Opus 4.8

Anthropic releases Claude Opus 4.8, with core breakthroughs in significantly improving the reliability, judgment, and long-running consistency of 智能体 tasks, marking AI's practical shift from 'usable' to 'trustworthy'.

Anthropic News ·

Introducing Claude Opus 5

Anthropic launches Opus 5, delivering near-top-tier intelligence at half the cost of Fable 5, with self-iteration and tool-building capabilities that signal a new direction for agentic models.

Anthropic News ·

Is grep all you need? Lexical VS Sematic Search for 智能体s

The article explores the boundaries between traditional grep and semantic search/RAG for AI agents, highlighting grep's limitations with unstructured documents and at enterprise scale, and proposes a hybrid approach combining parsing tools.

LlamaIndex Blog ·

Our position on open-weights models

Anthropic CEO Dario Amodei clarifies the company has never advocated for banning open-weights models, and warns that the real national security nightmares—authoritarian military AI and model misuse—can't be solved by protectionist bans.

Anthropic News ·

Expanding Project Glasswing

Anthropic is scaling its AI-driven critical infrastructure defense network while warning that automated AI cyberattacks will become ubiquitous within a year, forcing the industry to shift from vulnerability discovery to rapid remediation.

Anthropic News ·

LLM OCR: The Error Got Quieter, Not Rarer

Large language models for optical character recognition lower error rates but produce stealthier hallucinations instead of obvious garbled text, rendering legacy validation tools obsolete and demanding new evaluation paradigms and system architectures.

LlamaIndex Blog ·

March 2026: LangChain Newsletter

LangChain is pushing agents from experimental prototypes to scalable, manageable enterprise assets through updates like LangSmith Fleet, Skills, and Sandboxes.

LangChain Blog ·

Anthropic acquires Stainless

Anthropic acquires core SDK tool provider Stainless to solve the 'last mile' problem of AI agent connectivity and strengthen its MCP protocol ecosystem.

Anthropic News ·

Open Models have crossed a threshold

LangChain's evaluations show that open-source models like GLM-5 and MiniMax M2.7 now match top closed-source models on core agent tasks, while offering up to 90% cost reduction and significantly lower latency.

LangChain Blog ·

Introducing Claude Sonnet 5

Anthropic's Sonnet 5 delivers agentic performance close to the Opus flagship at significantly lower cost, enabling developers to build powerful autonomous agents with mid-tier models.

Anthropic News ·

Quoting Claude Opus 5 system prompt

Anthropic hardcodes export control suspension details into Claude Opus 5's system prompt, revealing how system prompts act as emergency interfaces for models to handle real-world policy changes.

Simon Willison ·

Quoting Drew Breunig

When top models no longer hide engineering flaws with a 'free lunch,' developers must rethink the balance of context strategies, workflow design, and cost efficiency.

Simon Willison ·

AI is removing the middle class of software engineering

Florian Herrengt argues that while AI coding tools boost output speed, they cause system complexity to spiral out of control, leaving teams trapped in "cognitive debt" where no one truly understands the code.

Simon Willison ·

Stealing Reasoning Traces from Proprietary LLM APIs

Research reveals that major LLMs reuse encryption keys for reasoning blocks across models, allowing attackers to recover hidden reasoning via weaker model jailbreaks and exposing new prompt injection risks.

Simon Willison ·

The pressure

curl's lead maintainer, Daniel Stenberg, reveals that an unprecedented flood of high-quality, AI-assisted security vulnerability reports is putting immense pressure on the open-source project's team.

Simon Willison ·

vLLM Tops the Artificial Analysis Leaderboard

The open-source inference engine vLLM has outperformed all proprietary competitors in deploying multiple frontier open-weight models, with its core optimization techniques like operator fusion publicly available, revealing the immense potential of open source in AI inference.

vLLM Blog ·

Which tokens does a hybrid model predict better?

Hybrid models significantly outperform pure Transformers in semantic understanding and dynamic context tracking, but lag in verbatim repetition, revealing a clear architectural division of labor.

Hugging Face Blog ·

Your harness, your memory

The article argues that agent harnesses are inextricably tied to memory; using a closed or API-based harness means ceding control of your agent's memory to a third party, creating deep lock-in. Memory should be open.

LangChain Blog ·