← Back to Home

Tag: Reinforcement Learning (15 articles)

Granite 4.2 LLMs: How They're Built

IBM releases open-source reasoning model Granite 4.2, integrating chain-of-thought, tool calling, and agentic reinforcement learning into enterprise-grade models with 512K context support.

Hugging Face Blog · Aug 25, 2026

Deploy local agents everywhere with LFM2.5-2.6B

LiquidAI's LFM2.5-2.6B uses innovative Agentic RL to outperform models 4x larger on tool use and instruction following, enabling capable, privacy-preserving agents to run locally on everyday devices.

Hugging Face Blog · Aug 4, 2026

The State of Simulation for Physical AI: An Overview

NVIDIA's team systematically explains how physical AI leverages high-performance simulation engines to overcome three core challenges: data scarcity, expensive training, and safety risks, marking robotics development's transition from 'trial-and-error' to the 'digital twin' era.

Hugging Face Blog · Jul 22, 2026

Native RL APIs in vLLM

vLLM introduces native Reinforcement Learning APIs to standardize weight synchronization and improve asynchronous training support, addressing key pain points of framework fragmentation and fragile deployments in online RL for large models.

vLLM Blog · May 28, 2026

Reward Hacking in Reinforcement Learning

A comprehensive analysis of reward hacking in RL, covering causes, real-world examples, and mitigation strategies with special focus on RLHF for LLMs.

Lil'Log · Apr 5, 2026

Reward Hacking in Reinforcement Learning

Reward hacking presents challenges in reinforcement learning due to flaws in reward functions, particularly impacting language models, necessitating further research and mitigation strategies.

Lilian Weng · Nov 28, 2024

The Transformer Family Version 2.0

Lilian Weng's new article deeply explores the evolution and new features of Transformers, revealing their ongoing impact in natural language processing.

Lilian Weng · Jan 27, 2023