Granite 4.2 LLMs: How They're Built
IBM releases open-source reasoning model Granite 4.2, integrating chain-of-thought, tool calling, and agentic reinforcement learning into enterprise-grade models with 512K context support.
IBM releases open-source reasoning model Granite 4.2, integrating chain-of-thought, tool calling, and agentic reinforcement learning into enterprise-grade models with 512K context support.
The Qwen 3.8 27B model matches the scores of trillion-parameter giants like GPT-5.6 on the Artificial Analysis Intelligence Index, revealing new possibilities for smaller models to achieve 'dimensionality reduction strikes' on specific tasks.
The powerful Qwen 3.8 27B model defaults to its highest reasoning effort level, making even simple tasks extremely time-consuming.
In the first half of 2026, Chinese labs rapidly 'skipped' into training trillion-parameter open models, while US efforts shifted towards hardware and infrastructure-centric open-sourcing, reshaping the open model ecosystem.
NVIDIA releases Nemotron 3.5 Lightning, a lightweight model designed specifically for AI agents with only 3B active parameters out of 30B total, balancing efficiency and capability, with immediate production support from vLLM.
Meta releases Muse Glimmer, a 30B open-source multimodal model optimized for local agentic tasks. It outperforms larger models like Gemma4 and Qwen3.6 on key agent benchmarks, signaling a shift toward private, on-device AI agents.
DeepMind's AI model WeatherNext, published in Nature, gains a full day of predictive accuracy in hurricane forecasting—equivalent to a decade of progress—and was already used to issue historic warnings during the 2025 hurricane season. The model is now open-sourced.
A trio of high-profile open letters aired the deep rift in AI: Microsoft rallied for open weights and distillation, Anthropic warned of catastrophic misuse, and independent developers urged nuanced regulation that won’t crush innovation.
Ben Thompson proposes US legislation to clarify that training data collection is fair use and to ban terms that forbid distillation, countering Chinese open-source model competition and revealing the hypocrisy in AI policies.
NVIDIA releases a 4B-parameter edge world model enabling real-time vision reasoning and robot action generation on memory-constrained devices, setting new benchmarks.
Kimi K3 debuts with 2.8T parameters and a premium price tag, challenging the stereotype that Chinese models must be cheap, while its SVG generation skills reveal a new dimension in AI evaluation.
Meta released the first API for a Spark model, Muse Spark 1.1, with major improvements in tool calling and computer use; Simon Willison quickly built a CLI plugin to simplify developer access.
Simon Willison reviews the open-source Ornith-1.0 model, highlighting its efficient tool calling and code understanding for agentic tasks, signaling new advances in open agentic coding models.
Z.ai releases GLM-5.2, the first open-source model to achieve stable 1M-token context and rival top closed-source models on long-horizon coding benchmarks.
Holo3.1 makes critical breakthroughs in environment robustness, local deployment, and real-time speed, signaling that general-purpose computer use agents are moving from capability demos to production-ready engineering.
Hugging Face has released six Ettin reranker models of varying sizes, designed to significantly improve the accuracy of search and RAG systems at low cost through a 'retrieve-then-rerank' two-stage architecture.
IBM releases two Apache 2.0 open-source multilingual embedding models, where the 97-million-parameter compact version outperforms all models of similar size on various benchmarks, demonstrating the huge potential of 'small but mighty' models for specific tasks.
IBM's Granite 4.1 series demonstrates that a meticulously engineered data pipeline and multi-stage training can enable an 8B dense model to match or exceed the performance of a previous 32B MoE model, highlighting a paradigm shift where data quality trumps parameter count.
Microsoft releases VibeVoice, an MIT-licensed Whisper-style speech model with built-in speaker diarization, capable of locally transcribing up to one hour of audio on a Mac.
DeepSeek's V4 series delivers near-frontier performance at a fraction of the cost (Pro at $1.74/M input, Flash at just $0.14/M), potentially reshaping the cost-effectiveness standard for open-weight models.
DeepSeek-V4 makes million-token context windows practically usable for long-running AI agents by dramatically cutting inference costs and memory usage through its novel hybrid attention architecture.
Alibaba's Qwen releases Qwen3.6-27B, a dense 27B parameter model that outperforms the previous generation's 397B MoE flagship on coding benchmarks, signaling a turning point for efficient, local-first coding models.
NVIDIA trained the Nemotron OCR v2 model on 12 million synthetic images, achieving high accuracy (NED as low as 0.035) and high speed (34.7 pages/second on a single A100 GPU) across six languages, demonstrating that synthetic data is a key solution to the multilingual data bottleneck in OCR.
Simon Willison's famous 'pelican riding a bicycle' benchmark surprisingly shows a locally-run, smaller Alibaba Qwen3.6 model outperforming the cloud-based, massive Claude Opus 4.7 in creative SVG generation, revealing the surprising potential of open-source models for specific tasks.
Anthropic CEO Dario Amodei clarifies the company has never advocated for banning open-weights models, and warns that the real national security nightmares—authoritarian military AI and model misuse—can't be solved by protectionist bans.
LangChain's evaluations show that open-source models like GLM-5 and MiniMax M2.7 now match top closed-source models on core agent tasks, while offering up to 90% cost reduction and significantly lower latency.