← Back to Home

Tag: 多模态 (9 articles)

Introducing Cosmos 3 Edge

NVIDIA releases a 4B-parameter edge world model enabling real-time vision reasoning and robot action generation on memory-constrained devices, setting new benchmarks.

Hugging Face Blog · Jul 20, 2026

Welcome Inkling by Thinking Machines

Inkling, a 1T-parameter open model with native multimodal understanding and 1M context, redefines open-source AI with architecture innovations that enable efficient inference and agentic applications.

Hugging Face Blog · Jul 15, 2026

Introducing GPT‑Live

OpenAI upgraded ChatGPT's voice mode to GPT‑Live, which can fluidly converse while delegating complex tasks to GPT‑5.5 in the background. Simon Willison's hands-on test shows it's once again a useful thinking companion.

Simon Willison · Jul 9, 2026

LLM 0.32a0 is a major backwards-compatible refactor

Simon Willison's LLM library undergoes a major refactor, evolving from simple text prompts/responses to a structure supporting multi-turn message sequences and streaming mixed-type responses, adapting to modern LLMs' multimodal and tool-calling capabilities.

Simon Willison · Apr 30, 2026

LlamaIndex Newsletter 2026-04-14

LlamaIndex launches ParseBench, the first OCR benchmark for AI agents, and demonstrates breakthroughs in structured document understanding and multimodal reasoning, signaling a shift from text extraction to deep semantic comprehension.

LlamaIndex Blog ·