Deploy local agents everywhere with LFM2.5-2.6B
LiquidAI's LFM2.5-2.6B uses innovative Agentic RL to outperform models 4x larger on tool use and instruction following, enabling capable, privacy-preserving agents to run locally on everyday devices.
LiquidAI's LFM2.5-2.6B uses innovative Agentic RL to outperform models 4x larger on tool use and instruction following, enabling capable, privacy-preserving agents to run locally on everyday devices.
Simon Willison reignites his interest in MCP with the new stateless spec, highlighting how a single HTTP request replaces the old multi-step session, simplifying both client and server implementation.
Meta released the first API for a Spark model, Muse Spark 1.1, with major improvements in tool calling and computer use; Simon Willison quickly built a CLI plugin to simplify developer access.
Newer Claude models are increasingly making mistakes when calling third-party edit tools, likely because Anthropic over-trained them on Claude Code's own tool syntax, degrading general tool-use ability and highlighting platform lock-in risks in AI training.
The system prompt update for Claude Opus 4.7 reveals the evolution of AI assistants from passive responders to proactive tool-users, deep task executors, and more responsible safety frameworks.
IBM and HuggingFace introduce the VAKRA benchmark, revealing that current AI agents perform poorly on complex multi-step tasks, with key failure modes including tool-chain planning, parameter passing, and error recovery.
LLM powered autonomous agents combine planning, memory, and tool usage, showcasing their potential in handling complex tasks and indicating a significant shift in work methodologies.