This article introduces Qwen-Image-Bench, a creator-centric benchmark for evaluating text-to-image models in complex, real-world scenarios, along with its open-source judge model, Q-Judger.
This article introduces PawBench, a benchmark designed to evaluate the combined performance of AI models and agent harnesses through 4,050 test runs.
The article introduces Alibaba’s open-sourced CosyVoice 3, AgentScope upgrades for production-grade AI agent development, and Qoder Teams, a new enterprise AI coding plan.