MetaRoCE: A New RDMA Transport Built for AI-Scale Ethernet
Meta releases MetaRoCE, a new RDMA transport protocol that shifts intelligence from switches to NIC endpoints, solving the network bottleneck problem in million-GPU AI clusters.
Meta releases MetaRoCE, a new RDMA transport protocol that shifts intelligence from switches to NIC endpoints, solving the network bottleneck problem in million-GPU AI clusters.
Meta unveils MTIA 300, a custom training chip with integrated NICs and dedicated communication engines, eliminating the compute-communication resource contention that plagues GPU-based recommendation model training.
SkyRL introduces IsoExec, a cross-framework unified execution abstraction that eliminates floating-point mismatch between training and inference engines in RL workloads through a shared execution contract and bitwise-consistent kernels.
NVIDIA NeMo AutoModel seamlessly plugs into the HuggingFace ecosystem, boosting MoE fine-tuning throughput by 3.4x-3.7x and cutting VRAM usage by 30% with a single import line change.
vLLM introduces native Reinforcement Learning APIs to standardize weight synchronization and improve asynchronous training support, addressing key pain points of framework fragmentation and fragile deployments in online RL for large models.
AWS details the infrastructure supporting the full foundation model lifecycle from pre-training and post-training to inference, revealing a paradigm shift from a single scaling law to three, and the deep integration trend of open-source software stacks with cloud infrastructure.