← Back to Home

Tag: 评估基准 (2 articles)

The Open 智能体 Leaderboard

Hugging Face and IBM launch the Open 智能体 Leaderboard, shifting evaluation from standalone models to full agent systems (including tools, planning, memory), while measuring both performance and cost.

Hugging Face Blog · May 18, 2026