Introducing Muse Spark 1.1
Meta released the first API for a Spark model, Muse Spark 1.1, with major improvements in tool calling and computer use; Simon Willison quickly built a CLI plugin to simplify developer access.
Meta released the first API for a Spark model, Muse Spark 1.1, with major improvements in tool calling and computer use; Simon Willison quickly built a CLI plugin to simplify developer access.
An end-to-end multimodal agent demo running on NVIDIA Jetson Orin Nano Super, showcasing how the model autonomously decides when to use the camera and answers questions with visual context, signaling the descent of powerful AI capabilities to edge devices.
Hugging Face releases a new tutorial demonstrating how fine-tuning multimodal embedding models can yield performance far surpassing general-purpose large models in specific domains (like visual document retrieval), even outperforming models with 4x its parameters.
Sentence Transformers v5.4 introduces native multimodal embedding support, enabling text, images, audio, and video to share a unified vector space for cross-modal retrieval.
Google DeepMind's Gemma 4 models innovate in parameter efficiency and support multi-modal inputs, marking a significant advancement in research on small effective models.
Gemma 4 introduces enhanced multimodal capabilities, supporting image, text, and audio inputs, significantly improving model intelligence and deployment flexibility across devices.
Granite 4.0 3B Vision is a multimodal model designed for enterprise documents, offering efficient information extraction and chart understanding capabilities, transforming document processing.
Holotron-12B optimizes inference efficiency and handles long contexts, becoming a powerful tool for high-performance computing agents, crucial for AI applications.
Google DeepMind introduces SL2T, a breakthrough sign-language-to-text model integrated into consumer products like Gboard, treating sign languages as true languages rather than visual codes.
LlamaParse leverages multimodal LLMs to not only extract text but also understand charts, images, and complex layouts within low-quality scans, fundamentally changing the capability boundaries of document parsing in legal discovery.
NVIDIA releases Nemotron 3 Nano Omni, a 30B-parameter MoE model that achieves extreme efficiency by activating only 3B parameters, offering a unified and cost-effective solution for multimodal AI agents.
NVIDIA releases Cosmos 3, the first open omni-model for physical AI that unifies world generation, physical reasoning, and action prediction in a single architecture.