×
Tongyi Lab

Qwen-Image-Bench: Beyond Basic Generation — Evaluating T2I Models in Complex Scenarios

This article introduces Qwen-Image-Bench, a creator-centric benchmark for evaluating text-to-image models in complex, real-world scenarios, along with its open-source judge model, Q-Judger.

What We Learned from Evaluating 4,050 Agent Runs

This article introduces PawBench, a benchmark designed to evaluate the combined performance of AI models and agent harnesses through 4,050 test runs.

Multilingual CosyVoice 3, Upgraded AgentScope for Production-Grade AI Agents, Enterprise-Ready AI Coding

The article introduces Alibaba’s open-sourced CosyVoice 3, AgentScope upgrades for production-grade AI agent development, and Qoder Teams, a new enterprise AI coding plan.